Meta's Movie Gen vs. OpenAI's Sora: a Detailed Review
Meta finally showed its hand against Sora, and Movie Gen is more than a video generator — it's a cast of foundation models that produces 1080p HD video across …
Google Researcher's In-Depth Analysis on End-to-End Speech Recognition, Part 1: Overview & Modeling
A decade of deep learning cut ASR word error rates by more than 50% relative and, along the way, made classical HMM pipelines look like a relic.
Deduct OpenAI GPT-4o's Neural Network Architecture
OpenAI hasn't published a paper on GPT-4o, so this video does the next best thing: it reverse-engineers what the neural network architecture almost certainly l…
Google's Universal Speech Model for 100+ languages beats OpenAI's Whisper Model
Google's Universal Speech Model quietly did something remarkable: matched or beat OpenAI's Whisper on ASR across 100+ languages while using roughly one-seventh…
Disclosing OpenAI GPT-4's vision+text model, data, and cost to train (speculated)
OpenAI famously refused to disclose GPT-4's architecture, parameter count, training data, or compute budget, which turned the tech report into a Rorschach test…
A Review of GPT-4's Technical Report (GPT-4 in a Nutshell)
GPT-4's technical report is famously thin on architecture details and thick on capability claims, which makes a clear-headed review essential reading for anyon…
[Olewave's Review] AudioLM: a Language Modeling Approach to Audio Generation
AudioLM is the paper that convinced a lot of people audio generation should look more like language modeling than signal processing.
In-depth review of OpenAI's GPT-3 : Language Models are Few-Shot Learners (Part 3/3: Results&Rest)
The final installment of this three-part GPT-3 deep dive gets to the payoff: what actually happens when you throw 175 billion parameters at a benchmark suite a…
[10 mins] Explain Why OpenAI's Whisper API Isn't As Good As ChatGPT
Whisper landed with a bang, but it did not reshape ASR the way GPT-3 reshaped NLP, and this ten-minute review argues the reasons are baked into the paper itsel…
In-depth review of OpenAI's GPT-3 : Language Models are Few-Shot Learners (Part 2/3: Results)
Part two of this GPT-3 deep dive moves past the setup and into the results that actually made people rethink NLP.
In-depth review of OpenAI's GPT-3 : Language Models are Few-Shot Learners (Part 1/3: Intro&Approach)
The opening installment of this three-part GPT-3 review sets up the paper that turned scaling from a research bet into industry gospel.
ChatGPT/ChatGPT Plus/InstructGPT:Training language models to follow instructions with human feedback
Behind the ChatGPT product sits the InstructGPT paper, and this detailed review pulls apart the pipeline that turned a raw GPT-3 into something people actually…
Explain how ChatGPT/ChatGPT Plus works in 3 minutes
ChatGPT looks like magic from the outside, but the recipe is well-documented if you know where to look, and this three-minute explainer condenses the core idea…
[Olewave's Review] CLIP (3/3): Learning Transferable Visual Models From Natural Language Supervision
The final part of this CLIP review lands on the results section, which is where OpenAI's contrastive image-text model shifted from an interesting idea to a fou…
[Olewave's Review] CLIP (2/3): Learning Transferable Visual Models From Natural Language Supervision
OpenAI's CLIP flipped the script on computer vision by tossing out fixed label sets and instead training on 400 million (image, text) pairs scraped from the in…
[Olewave's Review] CLIP (1/3): Learning Transferable Visual Models From Natural Language Supervision
Before there was BLIP, LLaVA, or any speech-LLM worth its salt, there was CLIP, and this is where the story begins.
[Olewave's Review] OpenAI's Whisper ASR: Robust Speech Recognition via Large-Scale Weak Supervision
When OpenAI dropped Whisper, the ASR community had to reckon with a system trained on 680,000 hours of multilingual, multitask, web-scraped audio that just wor…
[Long Review] Cascaded Diffusion Models for High Fidelity Image Generation
Before Imagen and DALL-E 2 made diffusion the default for generative imagery, this Google Brain paper laid out the cascaded diffusion recipe that a lot of late…
[Short Review] Cascaded Diffusion Models for High Fidelity Image Generation
Cascaded diffusion is the two-stage-plus recipe underneath a lot of modern generative work: generate a small image with one diffusion model, then hand it to su…
