Skip to main content
Back to timeline
arXivSource publication:

VQS computes answers with programs over structured image records, cutting self-evolving VLM pseudo-label errors from 24% to 6%

Synopsis

VQS has a model first parse an image into a structured record such as a scene graph, chart table, or diagram graph, then uses fixed program templates to write questions and compute their answers from that record while the model only confirms the atomic facts the program reads one at a time; human raters find 94.4% of VQS answers correct versus 76.4% for majority voting and 82.2% for a model judge, and across ten benchmarks it improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales, reaching 3.84 points at 2B after three training cycles.

AI-generated editorial illustration: Program-Verified Self-Evolution for Vision-Language Models

Interpretation

VQS changes the answer source in self-evolving QA from model outputs to program computation: a parser turns photographs, charts, diagrams, and infographics into structured records, fixed templates traverse the records to generate questions and compute answers, and the model acts only as a checker confirming the atomic facts the program reads. Prior self-evolving methods (VisPlay, R-Zero, EvoLMM, iReasoner, Vision-Zero, VISE) fall back on proxy labels such as majority vote, agreement among reasoning steps, consistency under an image transformation, or a model judge when no gold answers exist; VQS instead runs deterministic programs over parses kept well formed by constrained decoding. On 500 shared questions in a human evaluation, VQS answers are 94.4% correct versus 76.4% for majority voting and 82.2% for a model judge; VQS keeps 76% of questions while majority voting keeps 92.0% and the model judge 90.8%.

Claim-level verification is also used to select the parser's training targets without labels, making the parser itself more accurate. The paper compares three selection rules: a random parse, a parse accepted as a whole with one yes or no, and the parse with the highest fraction of verified claims, rating parse precision and answer accuracy by hand. On 150 samples, per-claim checking raises parse precision from 85.0 under whole-parse checking to 87.6 and answer accuracy from 94.0 to 96.1; F1 against GQA scene graphs, ChartQA tables, and AI2D-RST diagram annotations improves by 5.0, 6.6, and 7.4 points.

Across ten benchmarks and three model scales, VQS gives a larger average gain than the self-evolving baselines it compares against, and the gain keeps growing over training cycles. The paper reports VQS gains over base of 3.18, 2.32, and 2.43 points at 2B, 4B, and 8B, against 1.01, 1.04, and 0.66 for the strongest baseline VISE; after three cycles the 2B gain rises from 3.18 to 3.84 points. Results come from lmms-eval on ten benchmarks (GQA, OK-VQA, InfoVQA, ScienceQA, MMMU, MMBench, EmbSpatial, LogicVista, MMStar, SEEDBench), with training data limited to unlabeled images and no human annotation at any stage.

Ablations show the components contribute differently: removing parser training costs the most, the fact-checker is next, and the remaining components are each worth 0.3 to 0.6 points. The paper removes constrained decoding, the fact-checker, the blind gate, and the difficulty band one at a time, and compares template-family difficulty, generator type, curriculum order, and difficulty-band refresh. At 2B the full pipeline reaches 62.4 solver accuracy and 87.6 parser precision; removing parser training drops these to 58.0 and 80.8, and removing the fact-checker drops solver accuracy to 60.9; middle-difficulty families work best (62.4), and within-family hop ordering beats shuffled order (62.4 versus 61.3).

Perspective

The framework targets four image domains (photographs, charts, diagrams, infographics) and relies on a per-domain fixed parse schema plus a hand-written template library that is never trained; training needs only unlabeled images, so it fits settings where human annotation is unavailable but images can be parsed into structures. The paper reports gains on all ten benchmarks for Qwen3-VL at 2B, 4B, and 8B, and also on all ten for Gemma3-12B, InternVL3-8B, and Llama-3.2-11B-Vision, so the method is not tied to one model family. Three cycles keep generating new QA pairs from the same unlabeled images and the gain is still rising, suggesting room to extend cycles and template families.

Parsing errors still become wrong answers, and the paper notes diagram arrows and chart values are the claim types where the checker most needs improvement; the checker shares weights with the parser, and the appendix audits 1,417 claims with five different models, including 247 traps, of which the in-house checker rejects 244, though the authors treat other checkers' rejections as warning signs rather than error rates. The rewrite step is not covered by the checker or the blind gate; the paper hand-checks 437 questions and all four 8B errors swap two years in a yes/no comparison. Gains shrink with model size, from 3.18 points at 2B to 2.43 at 8B, and the share of benchmark answers that change also falls from 2B to 8B. These are scope and open questions rather than faults.

Sources