AVRI Cuts Vulnerability-Discovery Cost by 18.0% for Codex and 23.7% for OpenCode via a Persistent Bidirectional Evidence Trace
Synopsis
A multi-axis open-coding study of 200 CyberGym traces found that code localization and understanding (43.1%) plus vulnerability reasoning and trigger design (17.3%) account for 60.4% of tokens and that only 24.4% of matched comparisons preserve success at lower total cost; the authors then present AVRI, an interface built on a persistent Bidirectional Evidence Trace (BET), which on 20 evaluation tasks reduces total cost by 18.0% for Codex and 23.7% for OpenCode while preserving success and improving or maintaining recall.
Figure 1 . RQ2. Mean token consumption per task by activity stage, for four agents and five configurations, ten tasks each. Bars count input plus output with cached input counted once, and include auxiliary model usage. Solid segments cover activity through the first observed PoC execution feedback, hatching covers activity after it, and dots mark traces in which no PoC execution feedback was observed. Textures indicate feedback timing; stage labels describe request purpose rather than newly generated reasoning. Horizontal stacked bar chart of mean tokens per task by activity stage for twenty agent and method groups, with code localization and understanding as the largest segment in most groups.
arXivInterpretation
A multi-axis open-coding study of 200 CyberGym traces across four agents (Codex, OpenCode, Cybench, EnIGMA) under an unaided baseline and four efficiency methods quantifies where tokens go and where runs fail. Prior work proposed individual interventions for history compression, hypothesis ranking, execution feedback, or code navigation, but did not connect method effects to activity stages and failure bottlenecks on the same set of execution traces. 32,924 normalized requests were coded into nine activity stages; after freezing automatic labels, agreement with human consensus was 94% (188/200) on a proportionally sampled set of turns and 96% (96/100) on H/L boundary cases.
Code localization and understanding (L) accounts for 43.1% of tokens and vulnerability reasoning and trigger design (H) for 17.3%, together 60.4%; among 67 unsuccessful traces, H contributes 27.0 units of failure weight and L 19.0, together 68.7%. Token share and failure attribution are coded separately and rank differently: L consumes more tokens while H carries more failure weight, so reducing code-inspection cost and helping agents resolve trigger conditions are complementary objectives. The 200 traces consumed 776.60M tokens, including 54.53M (7.0%) from auxiliary models; failure bottlenecks assign one unit of weight per failed trace, divided equally among its supported categories, and describe observed blockage rather than causal contribution.
Among 160 enhanced–baseline pairs, only 39 (24.4%) both succeed with lower cost for the enhanced run; unsuitable signals (59 pairs) and access or use limitations (38 pairs) are the most frequent coded limitations, and seven pairs show overhead reversal. This supplies a concrete mechanism taxonomy for why existing efficiency methods yield limited benefit and shows that auxiliary-model billing can offset main-model savings, so savings must be assessed against the full model bill. Mechanism coding examined method outputs, subsequent agent actions, and validation feedback against the paired baseline; overhead reversal is determined from billing records, applying when main-model cost is below baseline but total cost including auxiliary calls exceeds it.
AVRI uses a persistent Bidirectional Evidence Trace (BET) to connect harness input propagation with unsafe-operation conditions, reducing total cost on 20 evaluation tasks from $132.65 to $108.75 (18.0%) for Codex and from $4.38 to $3.34 (23.7%) for OpenCode while preserving success and improving or maintaining recall. Among the evaluated configurations, none of the four existing methods reduces cost for both agents without lowering either effectiveness metric; AVRI is the only evaluated method that does so. The 20 tasks are disjoint from the 10 tasks used in the earlier empirical study; success and recall are derived from vulnerable- and fixed-build submission records rather than agent self-reports; time and cost include all 20 runs including failures, with Revelio preprocessing and AgentDiet auxiliary calls charged.
Perspective
The work targets vulnerability discovery where the harness entry point, target program, and brief input-format guidance are already supplied: the agent must still identify the unsafe operation and derive a triggering input, and the scope excludes attack-surface selection and harness construction. AVRI assumes a bounded analysis region; the authors state they would not expect the design to transfer unchanged to open-ended attack surface discovery, and memory-safety defects in C/C++ suit constraint atoms such as length and offset better than logic flaws in managed languages. The evaluation focuses on Codex and OpenCode because the earlier study found substantially lower success rates and generally higher costs for Cybench and EnIGMA. For a reader, the directly reusable elements are the activity-stage and failure-bottleneck coding framework, the accounting practice of including auxiliary-model billing in total cost, and the interface idea of returning relevant code together with supported relations between input values and unsafe-operation conditions.
Failure bottlenecks and method mechanisms were each coded by one author without independent review, so category assignments and frequencies may carry systematic interpretive bias. RQ3 and the Section 5 evaluation use one run per agent, configuration, and task, which cannot separate a method effect from run-to-run variance by replication; the authors also note the absence of component ablations prevents isolating each mechanism's contribution. Whether AVRI's gains are robust to run-to-run variation, and which part of BET produces the savings, remain open. In addition, AVRI is not a whole-program verifier or symbolic executor: analysis stops at unresolved aliases, unsupported side effects, indirect calls, or ambiguous macro behavior, unmapped values stay unresolved, and an uncertain call target keeps the result conditional on that target; a static derivation also does not establish that the condition is satisfiable or produces an observed failure, so execution must still test the candidate.
