Public articles linked to the same research event.
Microsoft Research Microsoft Research Asia introduces the Harnessed Agentic RL paradigm and open-sources a rebuilt Agent Lightning v1.0, in which the same agent harness used in deployment participates directly in reinforcement learning through an LLM proxy; the framework is about 3,500 lines of code, runs agents as standard Kubernetes jobs, and on an end-to-end coding agent pipeline built on SWE-smith, mini-SWE-agent, and Qwen3.5-9B, roughly 6,000 training samples raise Pass@1 on SWE-bench Verified from 41.8% to 56.4%, an absolute gain of 14.6 percentage points.
Microsoft Research Asia introduces the Harnessed Agentic RL paradigm and open-sources a rebuilt Agent Lightning v1.0, in which the same agent harness used in deployment participates directly in reinforcement learning through an LLM proxy; the framework is about 3,500 lines of code, runs agents as standard Kubernetes jobs, and on an end-to-end coding agent pipeline built on SWE-smith, mini-SWE-agent, and Qwen3.5-9B, roughly 6,000 training samples raise Pass@1 on SWE-bench Verified from 41.8% to 56.4%, an absolute gain of 14.6 percentage points.
Microsoft Research Asia introduces the Harnessed Agentic RL paradigm and open-sources a rebuilt Agent Lightning v1.0, in which the same agent harness used in deployment participates directly in reinforcement learning through an LLM proxy; the framework is about 3,500 lines of code, runs agents as standard Kubernetes jobs, and on an end-to-end coding agent pipeline built on SWE-smith, mini-SWE-agent, and Qwen3.5-9B, roughly 6,000 training samples raise Pass@1 on SWE-bench Verified from 41.8% to 56.4%, an absolute gain of 14.6 percentage points.
Microsoft Research Asia introduces the Harnessed Agentic RL paradigm and open-sources a rebuilt Agent Lightning v1.0, in which the same agent harness used in deployment participates directly in reinforcement learning through an LLM proxy; the framework is about 3,500 lines of code, runs agents as standard Kubernetes jobs, and on an end-to-end coding agent pipeline built on SWE-smith, mini-SWE-agent, and Qwen3.5-9B, roughly 6,000 training samples raise Pass@1 on SWE-bench Verified from 41.8% to 56.4%, an absolute gain of 14.6 percentage points.