SWE-Serve turns 53 SGLang tasks into a measurement: the same patches pass 45.9% under the full verifier and 69.4% once live-serving tests are excluded
Synopsis
SWE-Serve converts 83 merged SGLang pull requests into 53 executable tasks spanning six inference-engineering families—speculative and advanced decoding, model and backend enablement, kernels and quantization, serving APIs, caching and runtime state, and distributed execution and scheduling—and judges AI coding agents' patches with hidden verifiers; on the 19 tasks that start a real server and load a model, the same 627 patches pass 45.9% of the time under the complete verifier versus 69.4% when live-serving tests are excluded, meaning 147 patches flip from fail to pass, so passing local checks does not imply the full serving path works.
Interpretation
SWE-Serve makes repository-scale changes to inference-serving software executable: 53 tasks derived from 83 merged SGLang pull requests, with 37 tasks from a single upstream pull request and 16 combining two to six related changes; the median reference solution modifies 553 lines across seven files, and a typical verifier has seven tests for new behavior plus 10 regression tests. Existing repository-level benchmarks target general software-engineering tasks and inference benchmarks often concentrate on kernel generation or performance optimization; SWE-Serve instead covers model enablement, decoding, caching, scheduling, serving APIs, and runtime performance across the inference-serving stack. Task qualification screened 786 potential sources, built 156 executable candidates, and admitted 53; every admitted task was tested on its declared hardware, the unmodified repository had to fail the new-behavior tests while continuing to pass regression tests, a reference patch had to pass the complete verifier, and verifiers were challenged with agent-created patches, with tasks repaired, narrowed, or excluded when a concrete problem was found.
Live-serving checks are what separate patches that actually work: 19 tasks start a real server and contain 276 live-serving tests (242 sourced or adapted from SGLang, 34 covering behavior introduced by the corresponding merged changes), and the same 627 patches pass 45.9% under the complete verifier versus 69.4% without live-serving tests, with 147 patches flipping from fail to pass. This quantifies the gap where a patch passes tests yet fails when the server loads a real model and handles requests, and gives a concrete case: in the Gemma 4 MoE task, 16 of 33 patches passed every other check but failed at least one live-serving test covering model loading, expert routing, text and image serving, and batched generation with correct ordering and log probabilities. The 19 tasks load the required model on declared hardware and test the patch through a live serving interface; three tasks enforce a calibrated performance gate on an H100; twelve tasks run on CPU and 41 use a single NVIDIA H100.
How many runtime domains a task spans tracks its difficulty: dividing the request-to-output path into request handling and I/O, scheduling and request lifecycle, model execution, and KV-cache and runtime-resource management, the 26 tasks confined to one domain have a 69.0% pass rate versus 47.7% for the 27 tasks spanning more than one domain, a 21.3 percentage-point difference, with every model setting showing the same direction. This offers a measurable structural dimension for why some inference-engineering changes are harder for agents, beyond an overall score. The comparison is based on task groupings under each of the 11 models' best settings, and the direction of the difference holds across every model setting.
Under mini-swe-agent (a minimal software-engineering agent using only Bash) and closed-book conditions, 11 models and 31 model-effort configurations each ran the complete 53-task benchmark three times, with sessions capped at 210 minutes and 350 steps; mean pass@1 at each model's best configuration ranges from 34.6% to 75.5%, Claude Opus 5 and GPT-5.6 Sol both reach 75%, yet costs and runtimes differ sharply and no model leads all six engineering families. Results report pass rate alongside mean cost per task and mean wall time, showing cost does not map cleanly to performance: among the four models tied at 64%, mean cost ranges from $0.95 to $7.24 per task and mean wall time from 25.5 to 99.9 minutes; native harnesses did not improve the two leaders, with GPT-5.6 Sol at 73.6% in Codex and Claude Opus 5 at 69.8% in Claude Code versus 75.5% each with mini-swe-agent. Each configuration ran the complete benchmark three times, a task counts as solved only when the patch passes the complete verifier on the declared hardware, and all 1,749 trials behind the leaderboard were audited, with 196 prohibited retrieval attempts blocked and none succeeding.
Perspective
The work targets evaluating AI coding agents on repository-scale changes to inference-serving software, applies to running SGLang tasks on declared hardware, and releases task environments, verifiers, and baseline configurations so others can evaluate their own coding agents. This first release does not evaluate other inference engines, multi-GPU execution, or multi-node serving; twelve tasks run on CPU and 41 use a single NVIDIA H100. A SWE-Serve pass means only that the patch satisfies the benchmark verifier; it does not establish that an agent patch or benchmark reference solution is deployable, ready to merge, or endorsed by SGLang maintainers.
Of the live-serving tests, 34 were written to cover behavior introduced by the corresponding merged changes, so how equivalent their coverage is to upstream tests remains an open question; tasks come from a single project, SGLang, and 37 tasks derive from a single upstream pull request, limiting independence across tasks; the pass-rate difference between multi-domain and single-domain tasks holds in the same direction across every model setting, but the text does not decompose whether domain count itself drives it or whether it co-varies with factors such as change size; the closed-book setting allows Hugging Face access for model weights, and the effect of weight access on results is not separately quantified.
