Skip to main content
Back to timeline
arXivSource publication:

Chronos retrieves repository-evolution experience through a PR relation graph, lifting SWE-Agent's mean resolution rate from 69.2% to 72.9%

Synopsis

Chronos is a test-time framework that distills merged pull requests into four-field experience cards and links them in a weighted directed multigraph of eleven relation types spanning code, developer intent, and organizational context; at task time semantic search locates entry cards and weighted multi-hop expansion retrieves connected changes, and the same memory guides both candidate patch generation and patch selection. On SWE-Bench Verified the full workflow improves SWE-Agent across all six evaluated LLM backbones, raising the mean resolution rate from 69.2% to 72.9% and reaching 79.8% with MiniMax M2.5; with the same backbone it raises resolution rates from 48.3% to 51.7% on SWE-Bench Pro and from 41.0% to 43.5% on FEA-Bench Lite.

Source-provided article image: Chronos Enables Code Agents to Reason over Software Evolution
Figure 1 ·

Figure 1. Resolution rates of representative agents across three benchmarks. Labels identify the backbone and scaffold of each configuration; Tables 4 and 5 give our backbone-specific comparisons. Each system submits one final patch per task. Colors distinguish Chronos, open-scaffold systems, and proprietary systems. Comparison of Chronos with open-scaffold and proprietary code agents on three software-engineering benchmarks. Colors identify the three system groups, and the bars display resolution rates with one final patch submitted per task.

arXiv

Interpretation

Repository history is organized into a retrievable relational structure: each merged PR is distilled into a four-field card (Summary, Signals, Motivation, Solution), and cards are connected by eleven typed edges grouped into code evolution, developer intent, and organizational/social tiers with traversal weights decreasing from 1.0–1.5 to 0.3–0.45. Prior card-based reuse of repair experience, such as MemGovern, ranks cards mainly by query-to-index embedding similarity; Chronos adds weighted expansion along typed PR relations beyond semantic entry points, so related changes whose descriptions emphasize different concerns can still enter the candidate pool. Relation types, detection criteria, node coverage, and weight ranges are itemized in Table 1; graph-construction rules and default weight formulas are held fixed across the three benchmarks without benchmark-specific retuning.

One memory serves both generation and selection: a patch-focused change agent queries around design intent and a validation-strategy agent queries around behavioral mismatches and regression risks, each producing a candidate patch, while an evolution steward independently consults the memory and current source code to choose between them after X/Y order randomization. Repository memory is extended from auxiliary context for a single generation stage to evidence for two decisions, how to implement a change and which proposed implementation best addresses the task, with the retrieval service separated from agent roles and no model fine-tuning. Ablations show the single-agent variants reach 78.4% and 77.8%, both above the 76.4% SWE-Agent baseline; combining both strategies with the steward reaches 79.8%, the highest among the evaluated configurations.

Task resolution is evaluated on three benchmarks: on SWE-Bench Verified all six backbones improve, with the mean rising from 69.2% to 72.9% and MiniMax M2.5 going from 76.4% to 79.8%; SWE-Bench Pro rises from 48.3% to 51.7% and FEA-Bench Lite from 41.0% to 43.5%. Gains appear on standard issue resolution, longer-horizon repository-level changes, and new feature implementation, indicating that the same memory-construction and retrieval procedure transfers across task types. SWE-Bench Verified covers 500 tasks, 12 repositories, and six backbones; the full MiniMax M2.5 configuration was repeated five times with a mean of 80.0%, a sample standard deviation of 0.37pp, and a lowest run of 79.6%, still 3.2pp above the reported 76.4% SWE-Agent reference.

Human evaluation shows relational retrieval improves experience usefulness at fixed depth: with ten cards retrieved per task, graph-grounded search raises the mean number of useful cards from 1.24 to 2.87 and the fraction of tasks receiving at least one useful card from 68% to 89%, versus 0.18 and 14% for random retrieval. Retrieval quality is separated into coverage breadth (Hit) and useful experience per task (Useful), and per-edge removal experiments further measure the retrieval value of individual relation types. Two research-team members with professional software-engineering experience independently annotated 100 SWE-Bench Verified tasks with disagreements resolved by discussion; removal losses are largest for the code-continuity relations precedes (0.60) and extends (0.44), followed by developer-intent relations, with organizational relations contributing 0.06–0.18.

Perspective

The framework targets repositories whose evolution is recorded as pull requests; retrieval is scoped to the target repository of each task instance and uses the task PR's creation time as a cutoff, excluding PRs whose recorded creation or merge time is at or after that cutoff, with the same scope applied to semantic entry selection and graph expansion. It applies to the three evaluated task types: issue resolution, longer-horizon repository-level changes, and new feature implementation, with the card schema, graph-construction rules, and default weight formulas shared across benchmarks while historical content comes from each target repository. For practitioners, the reusable unit is a development decision together with the conditions that motivated it: Motivation explains the behavioral gap or assumption, and Solution gives the implementation logic, so reading them together helps judge which parts transfer to current code. Because the retrieval service is separate from agent roles, the same memory supports single-candidate agents as well as a two-candidate-plus-selection workflow; the single-agent variants reach 78.4% and 77.8% under the same memory interface.

Distillation is performed by an LLM and can omit context or introduce inaccuracies; the human study measures task-conditioned usefulness rather than source-level factual accuracy, and the two annotators share a research context that can introduce evaluator bias. Relation weights are a hand-specified prior held fixed across all evaluated settings, and sensitivity to the numerical weight constants remains unmeasured; per-edge removal losses are also not additive shares of a fixed total, so their ordering reflects each relation type's retrieval value in the presence of the remaining graph structure. Repeated-run statistics describe only the full MiniMax M2.5 configuration, while the baseline and remaining configurations use single runs; the single-agent variants combine graph retrieval with role-specific prompting, so their gains reflect that combination, and removing a strategy also removes the second candidate and steward selection, so the full workflow's additional gain reflects their combined effect. Evaluation covers six backbones on SWE-Bench Verified and MiniMax M2.5 on the two transfer benchmarks, leaving behavior on other backbones and repositories to be observed.

Sources