Skip to main content
Back to timeline
arXivSource publication:

PlaylistEval tests video-language judges on ~100-hour playlists: best judge reaches only 75.4% pairwise accuracy against 93.0% human agreement

Synopsis

The work introduces PlaylistEval, an agentic framework that builds video-language judge benchmarks without human annotation by automatically generating questions whose evidence is scattered across two distant segments of ~100-hour playlists and whose wrong answers differ only visually, yielding PlaylistBench with 630 pairs across seven domains; on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781), while evaluating 17 omnimodal and multimodal models from eight families shows frontier judges reach only 75.4% pairwise accuracy, open-source judge models perform far behind, retrieval improves accuracy by up to 10.5 points yet the best retriever finds both relevant segments in the top 10 only 37.

AI-generated editorial illustration: PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?

Interpretation

PlaylistEval replaces human annotation with an automatic pipeline: Phase I generates questions whose evidence lies in two distant segments, with a gold answer citing its supporting spans, and Phase II generates four wrong answers through controlled visual degradation that are indistinguishable from the gold on the transcript alone; validity gates at both phases return rejection reasons to the generator as a self-correcting loop, at roughly $1 per question. Existing video judge benchmarks use videos of only a few minutes, include answer pairs separable from the transcript alone, and rely on costly human annotation; this work addresses length, grounding, and scalability through two generation phases plus a feedback loop, and can be rerun on any new playlist collection. The paper specifies the two phases, the gate sets (structural validity, video necessity, video sufficiency; and structural validity, textual undetectability, visual detectability), cross-family generator-verifier assignments, and the roughly $1 per question cost; prompts and per-stage inputs are in the appendices.

The resulting PlaylistBench contains 630 preference pairs across seven domains (Education, Drama, Life, Art, History, Documentary, Podcasts) spanning static and dynamic knowledge; on a stratified subset of 152 pairs the benchmark agrees with human judgments 93.0% of the time, with inter-annotator agreement of 0.781. The benchmark retains high agreement with human judgments without human annotation, and every preference pair demands retrieval across the collection rather than being decidable within a single short clip. Human agreement is measured on the stratified subset of 152 pairs, with IAA reported as 0.781; domain coverage and the roughly 100 hours per domain come from the manually curated playlist collection.

Evaluating 17 omnimodal and multimodal models from eight families shows frontier judges reach only 75.4% pairwise accuracy while open-source judge models perform far behind; retrieval improves accuracy by up to 10.5 points, yet the best retriever finds both relevant segments in the top 10 only 37.9% of the time; frames or transcript alone costs several points against using both, while more reasoning budget or higher visual resolution brings only limited gains. These failure modes become visible only at day scale and beyond, which prior short-clip benchmarks do not cover; the paper concludes the bottleneck is finding the evidence rather than seeing it. Results come from a unified evaluation of 17 models and four retrievers, with an accuracy curve as playlist length grows from about 1 hour to 100 hours; per-domain numbers appear in the paper's tables.

Video reward models trained on short clips perform around chance in this setting, suggesting short-video judging does not transfer well to day-scale video; when the two answers swap sides, weaker judges (e.g., Gemma-4-26B-A4B and Gemini-3.5-Flash-Lite) reverse their verdict on roughly half of pairs, whereas stronger judges (e.g., Qwen-3.8-Max and Gemini-3.7-Flash) stay largely consistent. This extends the reliability question from accuracy alone to answer-order sensitivity and training-distribution transfer, and shows order robustness tracks judge strength. The order-swap experiment measures reversal rates on preference pairs, and the short-video reward model result uses the same evaluation protocol; the paper names specific stronger and weaker judges as examples.

Perspective

The work targets judge evaluation for day-scale video that requires retrieval across a collection: it applies to preference judgments where the unit is a playlist, evidence is scattered across two distant segments, and wrong answers differ only visually, serving research and engineering teams building long-video evaluation and training video reward models. Because the pipeline is designed to be rerun on new playlists, it can also be used to extend or rebuild the benchmark; the paper releases the pipeline, benchmark, and evaluation code.

Open questions remain: human agreement was measured only on the 152-pair stratified subset, so the remaining preference pairs were not human-verified; how the retrieval ceiling relates to judge accuracy on even longer collections is not yet given; and the near-chance result for short-video reward models rests on the current evaluation protocol, leaving their behavior under different training data and scales open. In addition, the loaded paper text is truncated at the evaluation table, so complete per-domain and per-model numbers should be checked against the original tables.

Sources