Skip to main content
Back to timeline
arXivSource publication:

onPanda cuts median annotation time by 51.5% with locate-correct-continue token-level annotation, and releases the Panda-CVL dataset and benchmark

Synopsis

The work presents onPanda: while reading a model response, the annotator locates the first inappropriate token, either picks a substitute from the model's candidate tokens or types the correct text, and the system truncates everything after that position and continues generation from the corrected prefix, repeating this loop; a small controlled study suggests a 51.5% reduction in median annotation time over manual post-editing, 97.0% of tokens in the resulting data are model-generated, and the data carries position-precise, naturally paired positive-negative token-level correction supervision, alongside the released Chinese vision-language Panda-CVL dataset and a token-level correction benchmark.

AI-generated editorial illustration: onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

Interpretation

onPanda compresses annotation into a locate-correct-continue loop: the annotator only finds the first inappropriate token, clicks a substitute among the model's top-20 candidates (default) or double-clicks to type free-form text, and the system truncates everything after that position and continues generation from the corrected prefix. Prior alignment-data pipelines either write from scratch or manually post-edit (costly and off-policy) or rank responses (coarse, response-level supervision limited to candidates the model can already sample); onPanda brings the correct-a-prefix-and-regenerate-the-suffix interaction, previously seen in interactive machine translation such as Predictive Translation Memory and INMT, into general-purpose alignment-data annotation and extends it to multimodal and structured agent settings. The paper gives system and implementation detail: streaming logprobs, recording each token's sampled probability and top-20 candidates, grouping the token stream into readable chunks along grapheme boundaries, keeping tokens before the correction point rather than regenerating them, and recomputing probabilities for free-form text via a single prompt_logprobs request. These are capability descriptions; the efficiency result comes from the controlled experiment in Section 3.1.

In a controlled experiment with 3 annotators and 21 image-description prompts under a Latin-square rotation of three methods, onPanda's median time is 330 s per prompt, 51.5% less than POTATO's 681 s and on par with Argilla's 336 s; its mean of 515.6 s is 27.5% and 24.7% lower than POTATO (711.1 s) and Argilla (684.5 s). The paper attributes part of the gain to continuation: a hallucination or error often recurs across a response, so post-editing must spot and fix every occurrence, whereas onPanda corrects only the first occurrence and the continuation stays consistent with the fix. On quality, onPanda attains the highest pairwise win rate (66.7%), and an anonymized human comparison found evaluators preferred onPanda over POTATO in 54.8% of pairs. The controlled experiment is modest (3 annotators x 21 prompts, a single rollout model Qwen3.5-35B-A3B at temperature 0.7 and top-p 0.8), and the paper notes the comparison evaluates complete workflows, making it difficult to separate interface, paradigm, or their combination; win rates come from GPT-5.5 judging each pair in both orders, and the human comparison is small.

The produced data keeps high on-policy fidelity: onPanda and Argilla reach perplexity 1.181 and 1.161, both within 1% of the 1.171 baseline and within re-sampling noise, while POTATO yields 1.596; onPanda and POTATO both reach 100% SFT coverage (21/21) versus Argilla's 52% (11/21), and onPanda derives 7.43 preference pairs per prompt versus Argilla's 6.00 and POTATO's 0.95. Preference ranking is confined to candidates the model can already sample and provides little corrective signal when the model cannot sample good responses; onPanda's free-form fallback lets annotators inject correct text the model cannot generate, deviating from the distribution only at the few positions beyond its capability while reaching 100% SFT coverage. PPL uses the mean of four independently sampled rollouts per prompt as baseline, and the paper reports per-prompt PPL fluctuating across rollouts at roughly the re-sampling noise level; coverage and preference-pair counts are direct tallies within this experiment.

The paper releases Panda-CVL, a predominantly Chinese vision-language dataset of 7,491 annotation sessions (6,839 train, 652 test) whose auxiliary rollouts were generated by the 32B dense VLM step-1o-turbo, plus a benchmark that decomposes one token-level correction into judging acceptability, locating the first error token, and supplying a replacement, expanding 652 test conversations into 2,126 evaluation instances (652 good, 1,474 not good). Evaluation uses a tokenizer-agnostic protocol: correction triples are computed with the Qwen3.6 tokenizer, and a prediction is correct only if the model-generated correction yields the same triples as the ground truth. Token-level correction remains challenging for all evaluated models, with the best F1 only 17.09% (GPT-5.5); GPT-6 obtains the highest Loc.-NG (24.46%) and Corr.-NG (15.83%), and more than nine in ten reasoning models exceed 90% Format while Corr.-NG stays below 16%. Benchmark results come from Table 3 in Appendix A, covering nine reasoning models and one instruct model; the multi-reference analysis in Appendix B reports human location agreement of 30.95% under exact matching (random baseline 0.20%) and 44.44% at four-token tolerance (about 1.84%), with replacement-token agreement of 69.44% conditional on identical positions.

Perspective

The tool targets annotation teams and data platforms building on-policy SFT, preference, and process-reward data: the lightweight web mode can be statically hosted without a database, and its components can be embedded into an existing data platform. It suits settings where a model response is largely locally sound and intervention is needed at only a few positions, covering plain-text LLMs, multimodal data, and structured agent trajectories; agent settings connect external tools and harnesses via MCP, and tool calls can be configured to await annotator approval before execution. The paper also provides reusable public resources: the Panda-CVL dataset and a token-level correction benchmark for community research on this new data form.

The controlled experiment is modest (3 annotators, 21 image-description prompts, a single rollout model), user-study participants all come from the in-house annotation team, and quality evaluation relies primarily on LLM-as-a-judge supplemented by a small human comparison; the paper notes the comparison evaluates complete workflows, making it difficult to separate interface, paradigm, or their combination. The efficiency and on-policy advantages of token-level correction rest on a sparse-error premise, and both diminish when the model falls far short of the target task and corrections become dense; the on-policy property is also relative to the rollout model at annotation time, fading when the data trains a different model or as weights are continually updated. The paper explicitly states that the downstream training benefit of token-level correction signals remains unverified by training experiments and that no controlled agent study was conducted, with agentic evidence limited to system capability and deployment statistics. In addition, Appendix B reports human location agreement of 30.95% under exact matching and 44.44% at four-token tolerance, with replacement-token agreement of 69.44%, indicating both consensus and diversity in the reference annotations and providing context for interpreting single-reference exact-match scores.

Sources