Skip to main content
Back to timeline
arXivSource publication:

Frozen CTC acoustic and language models plus a test-time adaptation module let ASR learn new words from unlabeled test data, cutting recurring-OOV character error rate by up to 14.97% relative on LibriSpeech

Synopsis

The work proposes giving automatic speech recognition the ability to learn the contextual representations and spellings of new words from unlabeled test data at test time: a frozen CTC acoustic model provides spellings, a frozen language model provides contextual evidence for out-of-vocabulary (OOV) word detection, and an adaptation module expands the vocabulary by learning lexical token representations with distributions over CTC-generated candidates, optimized by minimizing a Kullback-Leibler divergence (KLD) objective; the authors further argue that the CTC-weighted language model log likelihood ratio can be interpreted as the KLD between the unknown correct ASR and the unsupervised learned ASR, and that via a Pinsker bound the square root of KLD can be interpreted as an upper bound on

Source-provided article image: Learning New Words from Unlabeled Test Data in Automatic Speech Recognition
Figure 1 ·

Figure 1: Overview of proposed test-time adaptation framework. The CTC model generates N-best hypotheses for an input utterance. Word spans with strong CTC-LM preference disagreement are selected as potential OOV regions (one span here). Top-K CTC spellings are used to initialize a placeholder token with uniform spelling prior. For the first occurrence, the placeholder is added into the vocabulary and fine-tuned in context, while the utterance is buffered for potential final refinement. Otherwise, we retrieve the existing placeholder and re-transcribe the test utterance, possibly reducing errors. The retrieved placeholder is further updated using unsupervised model adaptation.

arXiv

Interpretation

It introduces a test-time ASR adaptation scheme that learns spellings and contextual representations of new words from unlabeled test data, with a frozen CTC acoustic model supplying spellings, a frozen language model supplying contextual evidence for OOV detection, and an adaptation module learning lexical token representations over CTC-generated candidates while expanding the vocabulary. Unlike supervised adaptation that relies on a pre-given vocabulary or labeled data, new-word learning here happens at test time from unlabeled test data, with the vocabulary expanded as part of the learning process. At the abstract level the text states the system components and the learning objective (minimizing KLD) and reports relative OOV character-error-rate reductions on LibriSpeech and dysarthric Speech Accessibility Project data, but the visible text gives no model sizes, data volumes, or ablation details.

It offers a theoretical interpretation: the CTC-weighted language model log likelihood ratio can be read as the KLD between the unknown correct ASR and the unsupervised learned ASR, and via a Pinsker bound the square root of KLD can be read as an upper bound on the total variation distance between the true and estimated spelling. It links the KLD objective being optimized to a bound on spelling-estimation error, giving the unsupervised adaptation objective an interpretable statistical meaning. This is a derivation the paper itself reports; the visible text states it as a conclusion and does not show the derivation steps or assumptions.

It reports relative OOV character-error-rate reductions on recurring OOV words: up to 14.97% on LibriSpeech and 6.67% on dysarthric Speech Accessibility Project data, measured against the corresponding rescoring system. It quantifies new-word learning on two corpora of different character, one of them dysarthric speech, broadening the described applicability of the method. These are relative figures reported in the abstract with an explicit baseline, the corresponding rescoring system; the visible text provides no absolute error rates, sample sizes, or significance information.

Perspective

The result targets ASR settings where unlabeled speech is available at test time and the new words recur in the test data; the method is realized as a frozen CTC acoustic model plus a frozen language model plus an adaptation module, suited to recognition systems using CTC and rescoring; the dysarthric-speech evaluation indicates it can extend to adaptation needs in atypical speech.

The visible text is the abstract and metadata; it provides no absolute error rates, test-set sizes, distribution of new-word occurrences, or significance, and does not show the assumptions behind the KLD derivation. Which conditions make the KLD and total-variation interpretation hold, and whether the method works equally well for words seen only once, remain open questions that require the full paper.

Sources