Sub-2B agents using semantic search for code localization beat grep on SWE-bench Lite and Multi-SWE-bench Flash while cutting CPU latency by 44.1%
Related research and updatesSynopsis
The authors introduce a two-stage training framework around colgrep, a local semantic search tool, combining supervised fine-tuning with turn-level credit based on retrieval outcomes and reinforcement learning on localization quality, and train three 1–1.7B-parameter models for file localization; on SWE-bench Lite and Multi-SWE-bench Flash, colgrep-equipped agents outperform matched grep agents at every training stage, with Qwen3-1.7B reaching 67.29 and 54.00 file F1 after weighted SFT plus RL, while using 29.1% fewer tokens and reducing end-to-end CPU latency by 44.1% and transferring better to seven languages unseen during fine-tuning.
Figure 2: File-localization performance and latency for Qwen3-1.7B after RL with the turn penalty. The latency breakdown separates LLM inference, search ( grep / colgrep calls), and other runtime overhead.
arXivInterpretation
A semantic search interface consistently outperforms a lexical one for compact localization agents. Prior work studied either lexical search agents (e.g., CodeScout, CodeGrep) or semantic retrievers, but did not compare the two interfaces symmetrically for policies below two billion parameters; this work regenerates teacher trajectories under each harness and applies the same training pipeline. Table 1 shows colgrep has higher file F1 at every training stage on both SWE-bench Lite and Multi-SWE-bench Flash; Table 5 shows the advantage holds for all three model families after weighted SFT, with MiniCPM5-1B rising from 18.84 to 40.77 on Multi-SWE-bench Flash.
Supervised fine-tuning that weights assistant turns by their retrieval outcomes improves file F1 at unchanged fine-tuning compute. Unlike whole-trajectory selection (unfiltered or an F1 threshold), this method assigns weights of 1.0 for discovering a new gold file, 0.6 for re-inspecting one, 0.3 for retrieving only non-relevant files, and 0.05 for empty results, errors, or repeated commands, with the submission turn weighted by its recall over gold files. Table 3 shows weighted SFT gives the highest file F1 for all three students on both benchmarks, e.g., Qwen3-1.7B reaches 55.09 on SWE-bench Lite versus 54.87 for unfiltered and 53.50 for F1 filtering.
Agentic multi-turn search beats single-shot direct retrieval, and the extra retrieval cost is small. The policy is compared against BM25 and colgrep direct-retrieval baselines that use the full issue text as one query and return an oracle top-10 subset selected with gold files, in both single-turn and multi-turn settings. In Table 2 the single-turn agent exceeds BM25 by 8.19 file-F1 points on SWE-bench Lite and 16.48 on Multi-SWE-bench Flash, and the colgrep baseline by 2.19 and 4.46 points; multi-turn reaches 67.34 and 53.37 with only 2.06 average retrieval calls versus 2.00 for single-turn.
Semantic search lowers inference cost and improves transfer to unseen programming languages. The authors measure end-to-end latency and token length under GPU and CPU deployment, and contrast Python-only SFT with multilingual SFT while holding the training sample count fixed to isolate the effect of excluding evaluation languages. Mean trajectory length falls from 7,857 to 5,574 tokens (29.1%), CPU latency from 36.08 to 20.18 seconds (44.1%), and GPU latency from 2.15 to 1.92 seconds (10.6%); under Python-only training colgrep drops 2.13 points on average versus 14.90 for grep, and colgrep leads by 28.82 points across the seven languages.
Perspective
The results target settings where the localization step must run locally or on consumer hardware: the authors scope the study to models below two billion parameters, motivated by fast, inexpensive local agents. Methodologically, it applies to retrieval-style localization that takes a natural-language issue description as input and returns files (optionally with line ranges or symbols), and it relies on pre-built repository indexes and a local colgrep binary. For downstream users this means localization can be delegated to a small semantic-search agent whose candidate files are then handed to a larger model for editing; the authors also report that multi-turn interaction beats single-turn and direct retrieval, so retaining a few interaction turns is worthwhile. Efficiency gains are largest in CPU-only deployment, where LLM inference accounts for 99.2% of grep latency.
Open questions the authors list include whether these benefits persist for larger models, which is not evaluated; all experiments use a single harness, so sensitivity to prompting strategies and tool interfaces is unknown; evaluation uses files modified by a reference patch as gold, which captures only one possible solution and may omit useful contextual files; and experiments are limited to code repositories, leaving extension to other retrieval-intensive domains such as document search as future work. The authors also note that each colgrep call currently launches a new process and reloads the retrieval model, and this avoidable overhead is included in the reported latency, so keeping the model resident could reduce it further. Readers porting this to their own repositories should note that indexes are built once per repository at the instance's base commit, and that training data were deduplicated by instance ID and repository/base-commit/patch hashes and decontaminated with long-n-gram matching.
