A developer oversees a coding agent workflow, with an LLM gateway and reinforcement learning training jobs running on Kubernetes.
Microsoft Research has put a blog post behind Agent Lightning v1.0, an open-source rebuild of its agent-training framework. The pitch is simple: train an AI agent inside the same harness you will use in production, instead of rebuilding that agent inside a training framework. The post is dated October 7, but the code is older. Agent Lightning v1.0 was released on August 17, and the GitHub release history now lists later point releases. So this is a "here's what shipped and why it matters" story rather than a fresh drop.

What "harnessed agentic RL" means​

Reinforcement learning (RL) lets a model learn by trial and error, guided by rewards. Modern agents are more than a model. They are a model plus a harness, the code that handles tools, context, and control flow. Microsoft's argument is that most agent RL systems require developers to reimplement the agent inside the training framework. That is costly, and it means the trained agent is not quite the one that ships.

Traditional frameworks such as verl, AReaL, and slime required the agent loop to be implemented directly inside the training framework. That makes it hard to reuse independently maintained harnesses. The paper names mini-SWE-agent, OpenHands, OpenCode, Claude Code, Codex, OpenClaw, and Hermes as examples.

Agent Lightning takes the opposite route. It places an LLM proxy between the agent and the model, so the agent keeps running as before while the framework observes and records its model calls. The paper formalizes this as harnessed agentic RL. The deploy-time harness is directly involved in model post-training, which narrows the gap between training and actual use.

The paper also says this proxy approach has caught on. It credits verl Uni-Agent, AReaL 2.0, slime v0.3.0, and Polar with adopting it.

The catch: the trainer is half-blind​

Handing the loop to the harness has a cost. The harness owns the interaction loop, while the training engine sees only a sequence of LLM request-response pairs. One task attempt (a rollout) can therefore become a variable number of training samples. That creates four problems, plus a fifth around scheduling:

  • Retokenization and merging. Harnesses store context as text, but RL needs the exact token IDs the model sampled. Re-tokenizing text can shift token boundaries, so adjacent calls can't always be merged into one sample.
  • Advantage calculation. If subagents or context summarization split a rollout into several samples, computing baselines per sample counts that rollout repeatedly. This distorts the rollout-level statistics.
  • Loss normalization. Averaging by sample count gives extra weight to rollouts that happen to produce more samples. Sample count is often just an artifact of harness behavior.
  • Backend scheduling. Sample counts and lengths are only known after the harness finishes. GPU counts and parallelism settings are usually fixed in advance.

The paper's practical answer is to keep reward and training weight tied to the rollout rather than to the number of rows it produces. In the authors' coding-agent experiments, rollout-level advantage combined with rollout-level normalization gave higher validation reward and steadier policy entropy than sample-level handling. These are the authors' own results, with no independent replication in the material I reviewed.

Architecture: three parts, about 3,500 lines​

Microsoft describes three components:

  1. API Gateway. An OpenAI-compatible proxy. It ties every model call to a rollout and records prompts, responses, and log probabilities.
  2. Rollout Controller. It starts and manages agents, either as local processes or as standard Kubernetes jobs.
  3. Trainer. Built on verl and vLLM, it creates rollouts, collects samples, assembles final training samples through an adapter, and updates the policy.

For an existing harness, the blog says that pointing the model endpoint at the proxy is usually enough to connect to RL training. The repository goes further and claims "ZERO changes" to the harness. Treat that as the intended integration model, not a promise that every setup works out of the box.

The 3,500-line figure is also easy to over-read. It describes the framework itself. You still need the underlying training and inference stack, models, environments, and GPUs. The repository's sample install targets a CUDA 13.0 machine. It runs uv sync and then a verl setup script with version 0.8.0 and the cu130 build.

Collocated Async RL​

Rollout times vary a lot between agents. Synchronous RL waits for the slowest agent and leaves GPUs idle. Fully asynchronous RL improves utilization but needs separate GPU pools for rollout and training.

Agent Lightning's "Collocated Async RL" lets rollout and model updates share the same GPUs. Once enough rollouts are collected, the gateway pauses new requests and lets in-flight ones finish. The update then runs, and rollout resumes. The harness doesn't see any of this. Microsoft reports about a 2x end-to-end speedup over synchronous RL while using fewer GPUs than conventional asynchronous RL. That figure comes from the project's own experiments, not a general guarantee for other workloads or hardware.

Kubernetes instead of paid sandboxes​

Running many agents at once eats CPU, memory, and compute. The paper says other frameworks commonly rely on commercial sandbox services, such as Modal Sandbox, Volcano veFaas, and E2B. Agent Lightning instead runs entirely on a self-hosted Kubernetes cluster. The Microsoft post says the controller works with self-managed clusters, cloud Kubernetes, or local infrastructure.

For IT teams that already run Kubernetes, that is an attractive option. It is not free, though. You still pay for and administer the compute, storage, networking, and agent environments. The material I reviewed doesn't include a security or isolation assessment. Running untrusted, model-driven code as Kubernetes jobs is a hardening exercise in its own right, and Kubernetes alone doesn't settle it.

One detail matters for Windows readers. The v1.0.2 release notes include a fix to fail fast when the local runner is used on native Windows. I only have that heading from the release notes, not the full context. It suggests the local runner isn't meant for native Windows, so Windows users should expect to use Linux, WSL, or a cluster. Check the project documentation before planning around it.

The headline result​

The reference pipeline combines SWE-smith data, mini-SWE-agent, and Qwen3.5-9B. It covers data cleaning, environment construction, reward-hacking safeguards, and RL training. The reported outcome: using only 6K training examples and modest compute, RL improves Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%, a 14.6-point absolute gain.

Two cautions apply:

  • The 14.6 figure is percentage points. As a relative improvement, 41.8% to 56.4% is roughly 35%. Mixing the two is a common way to make results sound bigger or smaller than they are.
  • It is one model, one benchmark, one harness, and one data mix, all reported by the authors. It does not predict results for your harness or tasks.

The repository also lists a later September example. It trains Qwen3.5-35B-A3B and reports SWE-bench Verified rising from 47.8% to 61.6% with about 1.8K examples. That is a separate experiment from the one Microsoft's blog describes.

Release status​

The newest tagged release in the repository is v1.0.2, dated September 29. It adds MoE coding-agent training and multimodal image flow into training batches. It also adds a long list of gateway and runner fixes. v1.0.1, from August 24, added an Agent Lightning Skill for Claude Code, Codex, and GitHub Copilot. The skill is meant to help a coding agent improve another agent against a benchmark. The project is MIT-licensed.

Why this matters​

Agent builders have a recurring problem: the version of the agent that gets tuned isn't the version that ships. Agent Lightning narrows that gap by training the real harness and capturing the model traffic through a proxy. The honest trade-off is that you must still build and validate the reward design, the data pipeline, the rollout infrastructure, and the trainer. The framework handles the awkward bookkeeping that harnessed training creates. It does not do that work for you.

Fairness note: this is Microsoft's own research, and the claims about speed, cost, and accuracy come from its authors. They are credible and reproducible in principle, since the scripts are published. Until outside teams report their results, treat them as promising rather than proven.

 

References

  1. Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses - Microsoft Microsoft 2026-10-07T16:00:00+00:00
  2. GitHub - microsoft/agent-lightning: The absolute trainer to light up AI agents. · GitHub github.com
  3. Releases: microsoft/agent-lightning github.com