One added colon flips Jev to accept wrong final answers, but only under sorted request presentation
Synopsis
The study tests whether the presentation of a structured judging request changes the effect of a candidate-text edit: on 200 previously unused DROP and GSM8K source clusters, adding one colon after random letters raised Jev's false acceptance from 1.0% to 26.0% with three output labels and from 3.0% to 26.5% with the published four-label instruction under sorted JSON keys, while both candidate variants were rejected under insertion presentation; Jev met the control thresholds under both grading configurations and presentations, and GPT-6 Sol produced no observed cue-condition false acceptances.
(a) Jev, three-label grading.
arXivInterpretation
A fixed candidate edit, adding one colon after random letters, made Jev accept explicitly wrong final answers under sorted JSON-key presentation, raising false acceptance from 2/200 to 52/200 with three-label grading and from 6/200 to 53/200 with four-label grading. Prior short-string judge attacks focused on punctuation, shared phrases or control tokens themselves; this work pairs a candidate-side edit with an integration-side request presentation and certifies the error with numeric references. 200 previously unused DROP/GSM8K source clusters, sixteen calls per source/configuration including four clean/revision controls and byte-identical duplicates; the paired presentation interaction passed the prespecified null-centered workload-stratified cluster bootstrap with Holm correction at the simulation floor.
Jev passed every predefined control threshold under both presentations and both grading configurations, yet excess colon-edit false acceptance appeared only under sorted presentation. This shows that clean control performance alone does not expose presentation-dependent vulnerability, separating basic judging competence from robustness to candidate perturbations. Control thresholds required at least 190/200 correct for bare correct, bare wrong and wrong-final-only candidates and 180/200 for a valid revision; Jev's four-cell panels were complete and all gates passed.
The interaction concentrated in candidates assigned the larger numerical mutation: near mutations pooled to zero interaction under both configurations, while far mutations produced positive interactions. This supplies a construction boundary, showing the effect varies across source groups that also differ in magnitude, digit length and textual plausibility. Each workload contained fifty near and fifty far mutations; different questions received each mutation distance, so the comparison describes a boundary across source groups rather than isolating one factor.
GPT-6 Sol passed the control thresholds and produced no observed false acceptance in its four cue cells, with 196 complete interactions equal to zero and six unknown transport completions unresolved. This provides a generative comparator observation on the same fixed grammar and source panel, not an architectural ranking or population equivalence claim. Sol used default reasoning, Standard service tier and strict categorical JSON with no elicited confidence; four primary sources remain incomplete, and finite-panel bounds are reported separately from the population interval.
Perspective
The result applies to the reference-aware numeric final-answer grading construction, the balanced DROP/GSM8K and near/far mutation mixture, and the recorded Jev service window; it is most useful to evaluators and integrators of structured judging interfaces who need to record candidate perturbations and request presentation separately. It supports testing candidate perturbations under the presentations an integration actually uses and preserving the executed request representation rather than only decoded field values. The observed zero cue errors under insertion presentation identify a tested condition and do not constitute a general defense claim for other edits or inputs.
Recursive sorting changes several request objects together, so state position, other object orders and hidden rendering remain unseparated, and internal cue coefficients and attention mechanisms are unidentified; the effect concentrates in far-mutation groups where numerical magnitude, digit length and textual plausibility vary with the source. Six unknown GPT-6 Sol transport completions remain unreplayed, leaving four primary sources incomplete, so population equivalence is not established. Natural-output prevalence, future service behavior and downstream consequences remain unmeasured, and most open implementations failed competence gates with precision configurations changing results, bounding transfer evidence to open deployments.
