VTR-Bench tests 11 video generators on 300 prompts and 1,202 text blocks, and the best model still shows an overall word error rate of 0.250
Synopsis
The authors built VTR-Bench, a benchmark that embeds prescribed text and its carriers into 300 prompts (1,202 text blocks) across advertising, science, user interfaces, culture, and daily life, scores text fidelity by carrier-specific transcription and word error rate (WER) and scene and motion fulfillment by a 20-question prompt-specific chain of query, and adds a Keyframe-Guided Agentic Framework in which a Director agent coordinates image generation, motion planning, and visual feedback; across 11 state-of-the-art models they find widespread difficulty in rendering scene text, with the best model reaching an overall WER of 0.250, while the framework cuts overall WER on Minimax H3 from 0.4468 to 0.3015 (a 32.5% relative reduction) and raises Video Score.
Interpretation
VTR-Bench isolates whether text is rendered correctly as a distinct capability of video generation rather than folding it into general quality evaluation. Earlier video generation benchmarks mainly assess visual quality, prompt alignment, compositionality, and physical plausibility; benchmarks that explicitly assess rendered text, including EvalCrafter, T2VTextBench, and AVGen-Bench, primarily cover short textual targets with limited coverage of longer passages, and T2VTextBench relies entirely on human evaluation; VTR-Bench covers multiple textual targets from short labels to extended passages and supplies reference strings and carrier annotations. 300 prompts are evenly distributed across five scenarios with 60 each, containing 1,202 text blocks, two to six blocks per prompt with 94.3% of prompts requiring three to five; total required text per video ranges from 58 to 496 tokens with a median of 102.5, and individual blocks range from 1 to 275 tokens with a median of 23.
The evaluation pipeline decouples text fidelity from scene and motion requirements, and both branches are validated against human annotations. The text branch has a vision-language model transcribe each specified carrier from its clearest occurrence and compares the result with reference text using WER; the video branch builds a 20-question chain of query per prompt spanning Scene Attributes, Motion Adherence, Spatial Relationship, Entity Presence, and Temporal Consistency, with Video Score as the proportion of satisfied requirements. On 100 sampled videos covering 2,000 chain-of-query judgments and 404 text blocks, the automated evaluator agrees with the three-annotator majority on 92.15% of queries; text-block WER shows Pearson 0.9541, Spearman 0.9370, Lin's CCC 0.9536, and MAE 0.0545, while video-level WER shows Lin's CCC 0.9835 and MAE 0.0407.
Across 11 state-of-the-art models, accurate scene text rendering is broadly difficult, and a high Video Score does not guarantee correct text. The results separate fulfilling scene and motion requirements from reproducing the intended text: Kling v3.0 and Seedance2.5 both score approximately 0.79 on video requirements while their WERs are 0.979 and 0.641, and Minimax H3 renders text more accurately than most proprietary models despite a lower Video Score. Among open-source models Minimax H3 reaches a WER of 0.447 while most others remain close to the metric's upper bound; among proprietary models Wan3.0 leads text fidelity in every scenario with an overall WER of 0.250 and Seedance2.5 reaches 0.641; with an alternative evaluator, Qwen3.6-27B, the overall rankings for WER and Video Score are identical across all 11 models.
The Keyframe-Guided Agentic Framework improves both text fidelity and fulfillment of video requirements through visual feedback at inference time. The Director agent first constructs an image prompt and requests candidate first frames, which a vision-language model inspects for text accuracy, legibility, carrier coverage, and scene consistency, after which the agent compares candidates, refines the prompt, or edits a candidate; it then builds a motion plan and generates video with the same model weights, using sampled-frame feedback on text stability, carrier persistence, and motion coherence to refine further. On Minimax H3, overall WER falls from 0.4468 to 0.3015, a 32.5% relative reduction, while Video Score rises from 0.7562 to 0.8315, with both metrics improving in all five scenarios; relative to first-frame conditioning alone (I2V), the full framework reduces overall WER by a further 9.4% relative and achieves the highest Video Score in every scenario.
Perspective
The work targets settings where text in a video must convey information accurately, such as advertisements, scientific demonstrations, and user interfaces, and it evaluates how accurately models render prescribed text on designated carriers under a given prompt. The benchmark provides 300 prompts, 1,202 text blocks, and reference strings, and its automated pipeline agrees closely with human judgments, so it can be used to compare models at scale, track progress, and serve as a testbed for inference-time control methods such as the Keyframe-Guided Agentic Framework. That framework is validated on Minimax H3 with Qwen3.7-plus as both Director agent and vision-language model and Qwen-Image-3.0 as the image generator, showing that explicit first-frame construction and visual feedback can improve text fidelity and scene and motion fulfillment under the same video generator weights.
The evaluation covers 11 models and five scenarios, but the space of text lengths, carrier types, and motion complexity is large, so which factors most affect text fidelity remains an open question. The failure analysis links unreadable text to insufficient character detail (small and blurry appear 3,595 and 3,196 times), while recoverable text still shows misspellings and garbled content (misspelled 602 times, garbled 480 times), indicating that transcribability and content correctness are separate things to track. Resolution and duration effects are compared only for Minimax H3 and LTX-2.3: higher resolution reduces WER for Minimax H3, whereas LTX-2.3 stays close to a WER of 1 across all four configurations and shortening duration gives no consistent improvement, so how far generation settings can substitute for underlying text-rendering capability needs more models to verify. The agentic framework's gains are measured on Minimax H3, and in UI scenes I2V achieves the lowest WER while the full framework achieves the highest Video Score, suggesting the best strategy may differ by scenario. This summary is based on the full paper text and the homepage evidence bundle and does not include implementation details from the code repository or every appendix table.
