RefineGAN: Universally Generating Waveform Better than Ground Truth with Highly Accurate Pitch and
"Better than ground truth" is the kind of claim that either makes you roll your eyes or hit play.
"Better than ground truth" is the kind of claim that either makes you roll your eyes or hit play. RefineGAN backs it up with a neural vocoder that reconstructs waveforms with pitch and intensity responses so accurate that listeners sometimes prefer the synthesized output to the original recordings โ a first in the GAN-vocoder space. The trick is a pitch-guided refinement architecture that starts from a rough template signal and iteratively sharpens it under multi-scale spectral and adversarial losses.
Beyond the impressive MOS numbers, RefineGAN generalizes across singers, speakers, and even unseen languages without retraining, which matters if you ship TTS or voice conversion at scale. The talk breaks down the template-based generator, why the pitch conditioning is what unlocks the fidelity gains, and where the model fits alongside HiFi-GAN and its descendants. If you're building anything that turns mel-spectrograms back into audio โ TTS, singing synthesis, voice cloning โ the design ideas here are worth borrowing.
