One-Edit-Distance FSA/Network-based (OEDN) in Mispronunciation Detection and Accented ASR
Non-native acoustic modeling has a chicken-and-egg problem: bad phone alignments produce bad models, and bad models produce bad alignments.
Non-native acoustic modeling has a chicken-and-egg problem: bad phone alignments produce bad models, and bad models produce bad alignments. This talk on the One-Edit-Distance FSA/Network-based (OEDN) approach breaks that loop by using a constrained phone decoder to recognize the most likely pronounced sequence within one edit distance of the canonical pronunciation, then iterating detection and recognition until no more mispronunciations show up.
On a 300-hour non-native spontaneous corpus, the resulting acoustic model shaved 6% off WER against a well-tuned context-dependent factorized-TDNN HMM baseline with matched neural topology, which is not a small delta at that scale. Beyond the accuracy win, the pipeline drops out a data-driven catalog of common mispronunciation patterns from non-native English learners, which is directly useful for speech assessment and pronunciation-scoring products. The talk walks through how goodness-of-pronunciation scores drive the initial detection, how the one-edit-distance constraint keeps the search tractable, and how repeated iteration keeps improving alignments until convergence. If you are working on accented ASR, computer-assisted pronunciation training, or any acoustic modeling task where the dictionary lies to you, hit play for a concrete, reproducible recipe.
