ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Agent Lightning 1.0 Is Open Source: RL Training for Real Agent Harnesses in 3,500 Lines

Agent Lightning 1.0 Is Open Source: RL Training for Real Agent Harnesses in 3,500 Lines

AI information • Admin • • 6 views

Agent Lightning 1.0 was open-sourced on October 7, 2026 by Microsoft Research Asia, and the whole framework is only about 3,500 lines of code. It targets an awkward problem in agent reinforcement learning: to train an agent with RL, teams have traditionally had to reimplement the agent inside the training framework, so what gets trained is never quite the agent that ships. The new release formalizes a paradigm the team calls Harnessed Agentic RL, built on one claim: the harness you deploy with should be the harness that takes part in training.

What is wrong with the old way

A modern coding agent is more than a model. It carries its own context management, tool protocols, execution logic and dependencies — mini-SWE-agent, OpenHands, Claude Code and Codex each do it differently. Traditional agentic RL systems such as verl, AReaL and slime assume the training framework owns the interaction loop, so every new agent means rebuilding its loop inside the framework. That is expensive, and the rebuilt agent can quietly behave differently from the deployed one: training scores look good, production does not match. It is the same friction that kept surfacing as agent-scale development infrastructure grew over the past year.

The key move is a proxy in the middle

Agent Lightning places an OpenAI-compatible LLM proxy between the agent and the model. The agent's code stays untouched: point the endpoint it already calls at the proxy, and the training system records every prompt, response and log probability as RL training material. Version 1.0 has three components: an API gateway that stores rollouts and serves the proxy, a rollout controller that launches agent executions, and a customized trainer built on verl. Execution runs natively on Kubernetes, with agents as standard jobs rather than paid commercial sandbox services, which keeps large-scale rollouts affordable. A scheduling scheme called Collocated Async RL lets rollouts and model updates share the same GPUs, delivering roughly 2x end-to-end speedup over synchronous RL while using fewer GPUs than fully asynchronous setups.

Six thousand samples, up 14.6 points

The release ships a complete coding-agent training pipeline: SWE-smith data, mini-SWE-agent as the harness, Qwen3.5-9B as the base model, and only about 6,000 training samples, with no large-scale compute. RL training alone lifted the model's Pass@1 on SWE-bench Verified from 41.8% to 56.4%, a gain of 14.6 percentage points. One detail practitioners will care about: when a single rollout splits into several training samples, advantage calculation and loss normalization must happen at rollout level, not sample level. Otherwise rollouts that produce more samples get counted repeatedly, and both validation reward and policy entropy destabilize.

Who should try it now

If you already have a working agent harness and skipped RL post-training because reimplementation looked too costly, this is aimed at you: the integration point is one proxy address, and the codebase is small enough to read end to end. If your agent is not stable yet, or your reward signal is still vague, wait — the framework answers how a real harness joins training, not what "better" should mean for your agent. So far the only fully demonstrated pipeline is a coding agent; results on other agent types will have to come from the community.

Recommended Tools

More