Disclosing OpenAI GPT-4's vision+text model, data, and cost to train (speculated)

Tutorials

Disclosing OpenAI GPT-4's vision+text model, data, and cost to train (speculated)

OpenAI famously refused to disclose GPT-4's architecture, parameter count, training data, or compute budget, which turned the tech report into a Rorschach test for the ML community.

OpenAI famously refused to disclose GPT-4's architecture, parameter count, training data, or compute budget, which turned the tech report into a Rorschach test for the ML community. This video takes the opposite approach: piece together the plausible answers from public leaks, prior GPT-3 and ChatGPT scaling laws, and the vision-plus-text hints scattered through the report. It is speculation, clearly labeled, but grounded in numbers that anyone working with large models can reason about.

Expect a walk through likely model size ranges, mixture-of-experts hypotheses, multimodal training data composition, GPU-hour estimates, and the eye-watering dollar figure attached to a training run of this scale. For voice-AI teams eyeing multimodal speech-plus-text systems or trying to justify their own compute budgets to a CFO, the framing here is useful even if the exact numbers shift. Treat it as a structured hypothesis rather than a leak, and see how close the guesses land compared to what has since trickled out. Worth 15 minutes if you want a mental model for the economics of frontier training.