Microsoft debuts Agent Lightning v1.0 to train AI agents
Microsoft Research Asia released Agent Lightning v1.0, a lightweight framework that trains AI agents using their real deployment harnesses to streamline reinforcement learning.

Microsoft Research Asia has open-sourced Agent Lightning v1.0, a lightweight reinforcement learning framework designed to train AI agents directly within their production environments. Comprising only about 3,500 lines of code, the system introduces a paradigm called Harnessed Agentic RL. Instead of forcing developers to rebuild complex agent loops inside a training framework, Agent Lightning places an API gateway proxy between the agent and the model. This allows the training system to observe and record model calls while keeping the deployment harness completely unchanged.
The framework is built around three main components: an API Gateway, a Rollout Controller, and a Customized Trainer built on verl. To run rollouts at scale, the Rollout Controller provides native Kubernetes support, executing agents as standard jobs on local, cloud, or self-managed clusters. This removes the need for expensive commercial sandbox services like Modal Sandbox or E2B. Additionally, Agent Lightning introduces Collocated Async RL, a scheduling method where rollout and model updates share the same GPUs. This approach achieves a twofold end-to-end speedup compared to synchronous reinforcement learning while using fewer resources than traditional asynchronous setups.
To demonstrate its data efficiency, researchers trained the Qwen3.5-9B model using a pipeline built on SWE-smith and mini-SWE-agent. Using only about 6,000 training samples from an open-source dataset, the reinforcement learning process raised the model's Pass@1 score on the SWE-bench Verified benchmark from 41.8% to 56.4%, representing an absolute gain of 14.6 percentage points. For practitioners, this framework resolves the discrepancy between trained and deployed agents. It also addresses statistical challenges like advantage calculation and loss normalization at the rollout level, ensuring more stable policy entropy and higher validation rewards during training.
This is our own summary of reporting by Microsoft Research Blog



