Skip to main content
Back to timeline
arXivSource publication:

PatchHolmes lets an agent read all 100 candidate commits at once, lifting patch-retrieval Recall@1 from 32.63% to 59.95%

Synopsis

PatchHolmes is a two-phase patch retrieval system whose first stage fuses BM25 with time decay and Qwen3-Embedding-8B dense retrieval via RRF into a top-100 candidate set, and whose second stage runs a frozen open-weight LLM agent that views all 100 candidates listwise through four tools (list_candidates, read_commit, read_file_diff, submit_answer) and submits a single best commit; on 809 CVEs from GitHubAD it reaches 59.95% Recall@1, 25.34% above the pointwise classifier Favia and 31.40% above IRCoT, adds 27.32% Recall@1 over the retriever's top pick on identical candidates, and transfers unchanged to PatchFinder_top10 to lift Recall@1 from 24.28% to 39.86%.

AI-generated editorial illustration: PatchHolmes: Agentic Patch Retrieval via Listwise Selection

Interpretation

It replaces pointwise scoring in the second stage of patch retrieval with listwise selection: the agent sees the top-100 candidate manifest at once, selectively opens 3 to 10 commits, and submits exactly one answer. Favia previously issued ten independent calls per CVE, each emitting yes/no plus a confidence, so candidates carried no joint signal; PatchHolmes uses one conversation instead of ten and lets candidates be compared against each other. On 809 GitHubAD CVEs, Recall@1 is 59.95% for PatchHolmes versus 34.61% for Favia, 28.55% for IRCoT, and 27.32% for SPFinder; McNemar's exact test is significant at every cutoff, and the 95% bootstrap interval for Recall@1, [56.7, 63.5], does not overlap Favia's [31.3, 37.9].

The gain comes from the agent loop rather than the retriever or model scale: with the candidate set held identical, the agent raises Recall@1 from 32.63% to 59.95%. Once the retriever is removed as a variable, Favia on the same candidates rises only from 34.61% to 39.80% and IRCoT from 28.55% to 37.58%, leaving 20.15% and 22.37% attributable to the selection method. On identical candidates the agent alone contributes 27.32% Recall@1; within the Qwen family a backbone swap (Qwen3-235B to Qwen3-Coder-30B) moves Recall@1 by 0.86%, and a second family reaches 56.98% (gpt-oss-120B) and 49.81% (gpt-oss-20B), both far above the 32.63% no-agent floor.

The four tools form a minimal sufficient set, and listwise viewing of the candidate manifest is the single largest step. Ablation shows that giving the agent only list_candidates and submit_answer already lifts Recall@1 from 32.63% to 48.21% (+15.58%), adding read_commit as an overview adds 9.27%, and reading actual diffs adds a further 2.47%. A capability ladder on the same Phase 1 candidates (Table 3, rows G, F, E, MAIN); in the Phase 1 ablation the dense path alone reaches 57.11%, within 2.84% of the full fusion, while RRF still adds +2.84% Recall@1 and +5.32% Recall@10.

The system runs on a frozen open-weight model over a local Git repository with no fine-tuning and no paid search API, at roughly one-seventh the token cost of Favia. Favia needs ten pointwise conversations of about 67,000 tokens each per CVE, whereas PatchHolmes uses one conversation per CVE averaging about 96,118 input tokens, 8 tool calls, and 5 commits read. At the appendix rate this is about USD 0.007 per CVE and about USD 58 for a pass over the full 8,401-CVE corpus, against about USD 400 for Favia; Qwen3-Coder-30B trades 0.86% Recall@1 for roughly half the wall-clock time (39 versus 86 seconds).

Perspective

The result targets patch retrieval over a local Git repository with a frozen open-weight model, serving downstream consumers such as security advisories, CVSS scoring, affected-version trackers, and SBOM scanners; the authors open-source the code at https://github.com/Aizhouym/PatchHolmes. The method never modifies code and the candidate set is fixed by Phase 1, so the tool surface only needs a survey call, two granularities of commit reading, and a single submit. Cross-corpus transfer is validated on 1,252 CVEs in PatchFinder_top10, and the real-world check covers 35 CVEs from 11 repositories disclosed between 2013 and 2025. The authors explicitly recommend surfacing candidates for human review rather than auto-populating vulnerability databases.

Evaluation can only cover CVEs whose patch has already been found, that is the 37% to 40% outside the 60% to 63% with no patch link; hard cases where no one has ever located the fix are absent from every benchmark. The authors add a check on 35 real-world CVEs, but scoring still requires a known fix. PatchFinder_top10 caps everyone at 46.10% because only 577 of its 1,252 CVEs have the true fix in the candidate pool. The GitHubAD draw passes 6 of 9 goodness-of-fit tests, with the three failures concentrated in the commit-count and diff-token tails, traced to 143 very large repositories whose commit counts and gold-patch diff lengths are replaced by corpus-wide medians. Multi-fix CVEs account for 7.28% and record only some of the valid fixes, so a different-but-equally-valid fix that was never recorded scores as a miss. In addition, the rank-frequency hint in the prompt is a rough aggregate statistic that does not match the final retriever on the test subset, and the authors did not rerun experiments with a numbers-free version.

Sources