[Olewave's Review] CLIP (3/3): Learning Transferable Visual Models From Natural Language Supervision
The final part of this CLIP review lands on the results section, which is where OpenAI's contrastive image-text model shifted from an interesting idea to a fou…
[Olewave's Review] CLIP (2/3): Learning Transferable Visual Models From Natural Language Supervision
OpenAI's CLIP flipped the script on computer vision by tossing out fixed label sets and instead training on 400 million (image, text) pairs scraped from the in…
[Olewave's Review] CLIP (1/3): Learning Transferable Visual Models From Natural Language Supervision
Before there was BLIP, LLaVA, or any speech-LLM worth its salt, there was CLIP, and this is where the story begins.
[Long Review] Cascaded Diffusion Models for High Fidelity Image Generation
Before Imagen and DALL-E 2 made diffusion the default for generative imagery, this Google Brain paper laid out the cascaded diffusion recipe that a lot of late…
[Short Review] Cascaded Diffusion Models for High Fidelity Image Generation
Cascaded diffusion is the two-stage-plus recipe underneath a lot of modern generative work: generate a small image with one diffusion model, then hand it to su…
