Phased Array Radar on China's Aircraft Carrier Fujian 003 and Its Connection with Speech Beamforming
The phased array radar mounted on China's newest aircraft carrier, Fujian 003, and the microphone arrays that let your smart speaker hear you across a noisy ro…
How Does the All-New Dictation in iOS 16 Work? Reveal Apple's Secret Sauce by a Speech Researcher!
Apple's iOS 16 dictation runs fully on-device, and that single design choice has cascading consequences for latency, privacy, and what kinds of ASR architectur…
[Long Review] Conformer: Convolution-augmented Transformer for Speech Recognition
Conformer is the model that quietly took over end-to-end ASR, and its recipe, bolt a convolution module onto every Transformer block, turned out to be one of t…
[Short Review] Conformer: Convolution-augmented Transformer for Speech Recognition
Conformer took the ASR world by storm by doing something almost boring: gluing a convolution module into each Transformer block so the model captures both glob…
博士大叔使用计算机作弊降维打击2022高考数学压轴大题 Ph.D. uses computer cheating to solve 2022 college entrance math exam
What happens when you turn a research-grade computer algebra toolkit loose on the hardest problem of the 2022 Chinese gaokao math exam?
[Long Review] Cascaded Diffusion Models for High Fidelity Image Generation
Before Imagen and DALL-E 2 made diffusion the default for generative imagery, this Google Brain paper laid out the cascaded diffusion recipe that a lot of late…
[Short Review] Cascaded Diffusion Models for High Fidelity Image Generation
Cascaded diffusion is the two-stage-plus recipe underneath a lot of modern generative work: generate a small image with one diffusion model, then hand it to su…
