Tutorials
Review DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation ~ Seed-TTS
The Seed-TTS team is back with DiTAR, and it's a clean answer to one of the more awkward tradeoffs in modern speech generation: how do you keep the flexibility…
Tutorials
[Long Review] Cascaded Diffusion Models for High Fidelity Image Generation
Before Imagen and DALL-E 2 made diffusion the default for generative imagery, this Google Brain paper laid out the cascaded diffusion recipe that a lot of late…
Tutorials
[Short Review] Cascaded Diffusion Models for High Fidelity Image Generation
Cascaded diffusion is the two-stage-plus recipe underneath a lot of modern generative work: generate a small image with one diffusion model, then hand it to su…
