Skip to main content
Back to timeline
Microsoft ResearchSource publication:

Microsoft Research Asia open-sources Agent Lightning v1.0: Harnessed Agentic RL in about 3,500 lines lifts Qwen3.5-9B from 41.8% to 56.4% on SWE-bench Verified

Related research and updates

Synopsis

Microsoft Research Asia introduces the Harnessed Agentic RL paradigm and open-sources a rebuilt Agent Lightning v1.0, in which the same agent harness used in deployment participates directly in reinforcement learning through an LLM proxy; the framework is about 3,500 lines of code, runs agents as standard Kubernetes jobs, and on an end-to-end coding agent pipeline built on SWE-smith, mini-SWE-agent, and Qwen3.5-9B, roughly 6,000 training samples raise Pass@1 on SWE-bench Verified from 41.8% to 56.4%, an absolute gain of 14.6 percentage points.

AI-generated editorial illustration: Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Interpretation

It proposes and formalizes the Harnessed Agentic RL paradigm, in which the agent harness used in deployment takes part directly in training, removing the need to reimplement the agent inside the training framework. Traditional agentic RL systems such as verl, AReaL, and slime assume the training framework owns the environment interaction loop and require rebuilding the agent loop inside the RL framework; Agent Lightning instead places an LLM proxy between the agent and the model, so pointing the endpoint that previously called the model API at Agent Lightning lets the framework observe and record model calls without changing existing harness code. The text presents this mainly as a paradigm definition and system design description with a schematic in Figure 1; no controlled comparison against rebuild-based methods is reported.

It implements a complete agent RL control plane in about 3,500 lines of code, with three core components: the API Gateway, the Rollout Controller, and the Customized Trainer. Compared with the original Agent Lightning, v1.0 emphasizes staying lightweight, integrating with real harnesses, and providing a complete, reproducible training pipeline; the Customized Trainer is built on verl and assembles training samples through a sample adapter. The evidence is the stated code size and component architecture, with Figure 2; no quantitative code-size comparison against other frameworks is provided.

It introduces Collocated Async RL, letting rollout and model updates share the same set of GPUs, achieving about a 2x end-to-end speedup over synchronous RL in experiments while using fewer GPUs than conventional asynchronous RL. Synchronous RL waits for the slowest agent in a batch and leaves GPUs idle, while fully asynchronous RL raises utilization but needs separate GPU pools for rollout and training; this approach pauses acceptance of new requests once enough rollouts are collected, waits for in-progress requests to finish before updating, and keeps the state transition transparent to the external harness. The evidence is the reported end-to-end speedup and GPU usage comparison, with Figure 3; specific GPU models, counts, or variance across repeated runs are not reported.

It builds an end-to-end coding agent pipeline on SWE-smith, mini-SWE-agent, and Qwen3.5-9B that uses only about 6,000 training samples to raise Pass@1 on SWE-bench Verified from 41.8% to 56.4%, and validates rollout-level advantage and normalization over sample-level handling. The pipeline covers data cleaning, environment construction, reward-hacking safeguards, and RL training, and needs no large-scale compute; rollout-level advantage combined with rollout-level normalization achieves a higher validation reward and keeps policy entropy more stable during training than sample-level handling. The evidence is the before-and-after Pass@1 comparison on SWE-bench Verified plus validation reward and policy entropy curves, with Figure 5; confidence intervals or statistical significance tests across multiple runs are not reported.

Perspective

The work targets agent developers and researchers who want to train with reinforcement learning using the real deployment harness, especially teams working on coding agents and general-purpose agent systems. The intended setting is: an existing runnable agent harness, where pointing the model endpoint at the OpenAI-compatible Agent Lightning proxy is usually enough; execution can be local processes or standard Kubernetes jobs on self-managed clusters, cloud Kubernetes, or local infrastructure, without paid commercial sandbox services. The end-to-end coding agent example uses SWE-smith, mini-SWE-agent, and Qwen3.5-9B, with a training set of about 6,000 samples and no need for large-scale compute. Natural next steps include extending the paradigm to harnesses beyond coding, reusing the same control plane across more models and tasks, and scaling rollouts at lower cost on existing clusters.

Several open questions remain for a careful reader. First, the 41.8% to 56.4% gain on SWE-bench Verified comes from a single reported comparison, with no variance across repeated runs or statistical tests, so the stability of the gain awaits further runs. Second, the about 2x end-to-end speedup and lower GPU usage are not accompanied by specific GPU models, counts, or parallel configurations, so behavior on different hardware needs independent verification. Third, the conclusion that rollout-level advantage and normalization outperform sample-level handling is presented through validation reward and policy entropy curves, without multi-seed or cross-harness replication. Fourth, the impact of retokenization-induced token boundary shifts and of sample splitting caused by subagents and context summarization under longer contexts or more complex harnesses is described mainly as a challenge, with systematic characterization of boundary conditions still lacking. Fifth, this is a full-text parse, but figures appear as references, so exact curve values and experimental configuration details require consulting the original figures.

Sources