Skip to main content
Back to timeline
arXivSource publication:

ATR judges dubbing lip-sync with monotonic alignment plus LLM reasoning, reaching 0.920 mean AUC across seven languages

Synopsis

The work introduces Align Then Reason (ATR), in which a monotonic CTC alignment scorer between candidate-text phonetic units and frame-level lip representations supplies per-unit soft tokens and a calibrated global alignment score to an LLM reasoner that outputs a Yes/No judgment with a short explanation, raising Qwen3.5-9B mean AUC from 0.610 to 0.920 on a seven-language benchmark and improving on three unseen MuAViC languages and two real dubbing downstream tasks.

AI-generated editorial illustration: Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing

Interpretation

The paper identifies and measures a structural gap: existing autoregressive visual-speech and video-language models stay near chance when the candidate text is fixed and only video timing is perturbed. Prior work treated lip-sync as transcription or audio-visual synchronization, whereas this work frames it as reference-free candidate-text-to-silent-video matching and builds separate content and temporal negatives to measure it. Pooled AUC is reported over seven corruption types in seven languages; Qwen3.5 SFT baselines from 2B to 9B sit in the 0.510-0.520 AUC range on the reverse, shift, freeze, and swap temporal axes.

ATR decouples temporal tracking from semantic judgment: a candidate-conditioned CTC alignment scorer emits per-phonetic-unit soft tokens and a calibrated global scalar, and an LLM reasoner reasons over them to give a Yes/No judgment plus a short explanation naming the weakest unit's position. Rather than asking a video-language model to infer timing from video directly, the monotonic CTC structure encodes order in the scoring itself, and continuous prefix representations inject local alignment evidence into the language model. Ablations show mean AUC falls from 0.920 to 0.672 without the calibrated scalar, with reverse, shift, and swap near chance; to 0.872 without soft tokens; to 0.856 without the LLM reasoner; and to 0.898 without auxiliary phonetic supervision.

On the seven-language benchmark ATR keeps high mean AUC across Qwen3.5, LLaMA-3.1-8B, and Mistral-7B reasoners and outperforms the compared frontier models and lip-reading baselines. The gains do not depend on a single LLM family, and variance across random seeds is low, indicating the interface transfers across architectures. ATR (Qwen3.5-9B) reaches 0.920 mean AUC, LLaMA-3.1-8B 0.890, and Mistral-7B 0.894; in the paired-accuracy table all ATR variants exceed 0.9, while frontier models average between 0.656 and 0.737 on temporal corruptions.

The alignment signal transfers to unseen languages and real dubbing pipelines: on MuAViC German, Arabic, and Russian, refitting only two scalar normalization parameters restores temporal sensitivity, and ATR beats lip-reading baselines on dub-line reranking and script-to-clip assignment. Cross-dataset evaluation updates no model parameters and only recalibrates two scalars; the downstream tasks use real professional dub lines rather than synthetic corruptions. After recalibration ATR-9B reaches 0.875 mean AUC on MuAViC versus 0.514 for Qwen3.5-9B Base and 0.578 for Auto-AVSR; dub-line reranking Top-1 is 0.479 versus 0.316 for the strongest lip-reading baseline, and script-to-clip assignment reaches 0.746 per-clip and 0.522 Exact Block in-domain and 0.704 versus 0.598 for Auto-AVSR on MuAViC two-way.

Perspective

The result targets dubbing review: at inference only silent video and candidate text are available, with no reference audio or reference transcript, and the judge outputs the Yes/No logit difference as the lip-sync score. It fits production pipelines that need one judge reused across languages, for example filtering or reranking candidate lines before speech synthesis and matching script lines to clips. For cross-dataset use, the paper's recipe is to refit only two scalar normalization parameters on 300 genuine target-language samples without updating model parameters, so the interface can be deployed at low cost under domain shift.

The paper reports AUC and paired accuracy on a seven-language benchmark and MuAViC, but does not measure human review cost or the cost of misjudgments in a real production setting. Under domain shift the raw scalar distribution moves, so two-parameter recalibration on genuine target-language samples is needed, and how the method behaves with no such samples remains an open question. In the downstream tasks, dub-line reranking candidates are LLM-generated and length-constrained, and in-domain script-to-clip assignment is a five-way problem, so how these settings differ from larger-scale or longer-segment pipelines is worth watching. The reasoner is also trained with supervised fine-tuning, and the paper lists reinforcement learning as a natural next step.

Sources