W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self Supervise
Why choose between contrastive learning and masked language modeling for speech pretraining when you can wire both into the same network?
Exploring Wav2vec 2.0 fine-tuning for improved speech emotion recognition
Speech emotion recognition has long been stuck with tiny labeled datasets and hand-crafted acoustic features.
Joint Unsupervised and Supervised Training for Multilingual ASR
Pretrain-then-finetune has become the default recipe for multilingual ASR, but it leaves a lot on the table: the two stages don't share losses, and the supervi…
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
Most speech SSL models optimize for ASR and hope everything else works out.
Improving Speech Recognition Accuracy of Local POI Using Geographical Models
Ask a voice assistant to navigate to a small local restaurant and you'll quickly find where ASR falls apart — long-tail POI names that never showed up in the L…
XLS-R: Self supervised Cross lingual Speech Representation Learning at Scale
How far can you push cross-lingual speech pretraining before the well runs dry?
Scaling Laws for Neural Language Models
Before GPT-3 made scaling feel inevitable, this OpenAI paper made it predictable.
UniSpeech-SAT : Universal Speech Representation Learning with Speaker Aware Pre-Training
Self-supervised speech models tend to learn phonetic content beautifully and speaker identity as an afterthought — which is fine for ASR and terrible for speak…
RefineGAN: Universally Generating Waveform Better than Ground Truth with Highly Accurate Pitch and
"Better than ground truth" is the kind of claim that either makes you roll your eyes or hit play.
SNRi Target Training for Joint Speech Enhancement and Recognition
Joint training of speech enhancement and ASR sounds obvious — just backprop the recognition loss through the denoiser — but in practice it's messy, because agg…
BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recog
What happens when you take the semi-supervised ASR recipe that beat the state of the art at 100M parameters and just… keep scaling?
