Tutorials
Improving Speech Recognition Accuracy of Local POI Using Geographical Models
Ask a voice assistant to navigate to a small local restaurant and you'll quickly find where ASR falls apart — long-tail POI names that never showed up in the L…
Tutorials
XLS-R: Self supervised Cross lingual Speech Representation Learning at Scale
How far can you push cross-lingual speech pretraining before the well runs dry?
Tutorials
Scaling Laws for Neural Language Models
Before GPT-3 made scaling feel inevitable, this OpenAI paper made it predictable.
Tutorials
UniSpeech-SAT : Universal Speech Representation Learning with Speaker Aware Pre-Training
Self-supervised speech models tend to learn phonetic content beautifully and speaker identity as an afterthought — which is fine for ASR and terrible for speak…
Tutorials
RefineGAN: Universally Generating Waveform Better than Ground Truth with Highly Accurate Pitch and
"Better than ground truth" is the kind of claim that either makes you roll your eyes or hit play.
Tutorials
SNRi Target Training for Joint Speech Enhancement and Recognition
Joint training of speech enhancement and ASR sounds obvious — just backprop the recognition loss through the denoiser — but in practice it's messy, because agg…
