Skip to main content
Back to timeline
arXivSource publication:

ACG lifts long-horizon agent success from 44.5% to 50.2%, with an 11.1-point gain on BrowseComp-Plus

Synopsis

The work introduces the Adaptive Consistency Graph (ACG), which incrementally organizes execution evidence and its provenance into a persistent graph and builds a temporary requirement-centered view for each decision under a bounded context budget, raising GPT-5.6-luna's equal-weight average success across three benchmarks from 44.5% with ReAct to 50.2% and BrowseComp-Plus from 62.4% to 73.5%, without replacing the base planner or tool executor.

AI-generated editorial illustration: Adaptive Consistency Graph for Long-Horizon Agents

Interpretation

ACG stores execution evidence as source-linked memory units, each retaining content, type, source references, observed occurrences, and structured fields; repeated content can share a unit while each occurrence is preserved. Unlike memory schemes that compress history into generated summaries, ACG does not overwrite the underlying record, so omitted evidence stays in a source-addressable execution record for later retrieval. The paper formalizes the relation between memory units and occurrences, and the appendix states that ingest preserves event identity, source references, and requirement-origin information without inferring missing attribution from semantic similarity.

At each decision, ACG seeds retrieval with both the original task and the latest interaction, selects unique seeds by round-robin interleaving, then expands along mutual-nearest-neighbor semantic edges and source-affinity edges with a fixed one-hop depth to form a requirement-centered temporary view. Compared with baselines that only retrieve or only compress, ACG designs relational organization and readout together, and the view is a read-only temporary projection that is not written back to the persistent graph. The configuration uses eight unique retrieval seeds, fixed one-hop expansion, up to eight historical candidates per requirement, and a 5,000-token memory view cap that shrinks further with remaining budget.

In the matched evaluation, ACG raises GPT-5.6-luna's equal-weight three-benchmark average success from 44.5% with ReAct to 50.2%, and BrowseComp-Plus from 62.4% to 73.5%; with DeepSeek-v4-flash the average is 45.9%. The paper stresses that gains depend on benchmark and model: Luna reaches only 12.4% on DeepPlanning and DeepSeek 11.7%, so the claim is conditional rather than uniform. The three fixed task sets contain 360, 830, and 300 tasks, with a nominal 100 external-action budget and a 4,096-token executor output limit; standard deviations come from 5,000 task-bootstrap resamples.

Diagnostics show ACG failures concentrate at the stage of delivered artifacts passing official evaluation rather than at stalling or repetition: on SWE-bench Lite with DeepSeek, repeated execution accounts for 8.2% of assessable ACG failures and missing deliverables 0.9%, versus 6.2% and 26.2% for ReAct. This separates avoiding stalled execution from producing a correct solution, showing that fewer repeated actions or more frequent delivery cannot substitute for official success rates. Diagnostics rest on deterministic observations of traces and final artifacts, with six overlapping categories each using its own assessable-task denominator and reported separately for successful and unsuccessful tasks.

Perspective

The result targets language-model agents that must stay consistent with their objectives across long sequences of dependent actions and tool calls, and applies to three task families: constrained planning, fixed-corpus retrieval, and repository-level code repair, under the premise that the base planner and tool executor are left unchanged, so it can serve as a context-construction layer above existing agent frameworks. For practitioners, ACG offers a traceable way to organize evidence: records omitted from the view remain in persistent memory and can be retrieved again, giving an operational starting point for keeping global objective cues and immediate interaction information available under a limited token budget.

The paper itself notes that provenance validates the origin of an association, not whether an observation is true or a requirement has been satisfied; missing or ambiguous action origins can reduce requirement-specific coverage, finite budgets can exclude relevant evidence, and graph construction and retrieval add computation beyond the base executor. The baseline analyses are observational: task structure and realized execution length are associated with outcomes but do not identify failure causes. In addition, the equal-weight average is only a summary of three benchmarks and should not be read as a pooled estimate, and the execution-length bins are conditional comparisons of different trajectory groups rather than matched-task estimates of benefit. Generalization beyond the evaluated models, tools, and task sets remains unestablished.

Sources