Microsoft’s Orchard Turns Agentic AI Training Into an Infrastructure Problem
Microsoft Research's Orchard framework treats AI agent sandboxes as shared, reusable infrastructure, training small open models to near-frontier performance.
Training an AI agent that can fix a bug in a real codebase, click through a live website, or manage someone’s inbox takes more than a capable language model. It takes thousands of disposable, isolated environments where that model can try a task, fail, and try again, safely separated from production systems and from each other. Building that infrastructure from scratch has been one of the quieter barriers keeping agentic AI research concentrated inside a handful of labs with proprietary sandboxes and closed training pipelines.
Table Of Content
- What Orchard Actually Ships
- Orchard Env: A Shared Substrate Instead of Embedded Infrastructure
- Three Recipes Built on the Same Foundation
- Orchard-SWE: A Small Coding Agent Near Frontier Performance
- Training Method
- Benchmark Results and Generalization
- Orchard-GUI: A Browser Agent Competitive With Proprietary Systems
- Orchard-Claw: Personal Assistants and the Harness-Training Effect
- The Infrastructure Numbers Behind the Claims
- Latency, Throughput, and Cost
- A Regression Test Against Plain Docker
- Why This Is a Research Program, Not a Single Release
- What It Means for Teams Building Agent Infrastructure
On August 3, 2026, Microsoft Research published a detailed walkthrough of Orchard, an open source framework the team has been building since May. Rather than leading with a new model, Orchard leads with an environment: a reusable, Kubernetes-native sandbox service called Orchard Env that the same infrastructure has now trained coding agents, browser agents, and personal-assistant agents on top of, without rebuilding the underlying plumbing for each one. The headline result is that an open coding agent built on Orchard, with roughly 3 billion active parameters, reaches 73.0 percent on the SWE-bench Verified benchmark, which Microsoft Research describes as matching frontier systems more than ten times its size.
What Orchard Actually Ships
Orchard is not a single artifact. It is a foundation layer, Orchard Env, plus three domain-specific “recipes” trained on top of it: Orchard-SWE for software engineering, Orchard-GUI for browser navigation, and Orchard-Claw for personal-assistant tasks. Alongside the code, Microsoft Research released the training data, evaluation methods, the underlying research paper on arXiv, and trajectory datasets on Hugging Face, all under an MIT license.
Orchard Env: A Shared Substrate Instead of Embedded Infrastructure
The central design decision is that the runtime environment should be its own standalone service, not something bolted onto a specific training framework. Orchard Env is a lightweight, Kubernetes-native sandbox service that spins up thousands of isolated containers on demand and drives multi-turn interaction between an agent and its sandbox, including command execution, file I/O, and git patches, over a REST API. It ships with synchronous and asynchronous Python SDKs, an in-pod agent that bypasses the Kubernetes API server for low-latency command execution, and support for any base Docker image, since a lightweight init container injects its own self-contained Python interpreter rather than requiring one already installed. Network isolation defaults to deny-egress through Calico NetworkPolicy, with per-sandbox CPU, memory, and timeout limits.
That substrate is also harness-agnostic. According to the project’s GitHub repository, every sandbox ships with several popular agent harnesses, including Codex, Claude Code, OpenClaw, and ZeroClaw among others, preinstalled and ready to run, so switching which harness a team trains or evaluates against becomes a configuration change instead of a new container image.
Three Recipes Built on the Same Foundation
Orchard-SWE: A Small Coding Agent Near Frontier Performance
Training Method
Orchard-SWE targets one of the hardest agentic tasks: multi-step software engineering, where an agent has to navigate a real codebase, reason across files, use tools, and recover from its own mistakes. The team distilled 107,000 agent interaction trajectories from two open-weight teacher models, MiniMax-M2.5 and Qwen3.5-397B, covering a broad range of real GitHub issues, then trained a Qwen3.5-35B-A3B model, a mixture-of-experts architecture with roughly 3 billion active parameters, on top of that data.
Two techniques do most of the work. Credit-assignment supervised fine-tuning keeps the productive portions of trajectories where the agent did not fully resolve the issue, instead of discarding failed attempts outright, which expands the usable training data. Reinforcement learning follows, using a method the team calls Balanced Adaptive Rollout to work around sparse rewards, since an agent training on real repositories usually only learns whether its final patch passed or failed the hidden test suite.
Benchmark Results and Generalization
The published figures: 69.7 percent on SWE-bench Verified with reinforcement learning alone, and 73.0 percent with value-model reranking added on top, both from a model with roughly 3 billion active parameters, a result Microsoft Research frames as matching frontier systems more than ten times larger.
The GitHub repository adds a generalization test the blog post does not emphasize as heavily. Orchard-SWE scores 51.0 on SWE-bench Multilingual, compared to 28.7 for a baseline the team calls OpenSWE-32B, and when evaluated under a coding harness it never saw during training, Kimi-CLI, it still reaches 45.0 on SWE-bench Verified and 20.1 on Terminal-Bench 2.0. The same baseline collapses to 3.6 and 0.0 on those two tasks under the unfamiliar harness. Training against a harness-agnostic environment, rather than one hardcoded harness, appears to be what makes that transfer possible.
Orchard-GUI: A Browser Agent Competitive With Proprietary Systems
Web navigation is a different kind of hard: an agent has to interpret a visual layout, handle a dynamic interface, and complete an open-ended task described only in natural language. Orchard-GUI trains a 4-billion-parameter vision-language model, Qwen3-VL-4B-Thinking, for that job, using only 400 distilled demonstrations plus 2,200 open-ended training tasks run against live websites rather than static snapshots.
On three web-navigation benchmarks, the resulting agent reaches 74.1 percent on WebVoyager, 67.0 percent on Online-Mind2Web, and 64.0 percent on DeepShop, for a 68.4 percent average. Microsoft Research’s own characterization is direct: the results make Orchard-GUI, in the blog’s words, “the strongest open-source GUI agent while staying on par with proprietary systems from OpenAI and Google.” The project’s GitHub page adds that Orchard-GUI beats its own 235-billion-parameter teacher model despite training on two orders of magnitude fewer tasks, a distillation result as much as a training-infrastructure one.
Orchard-Claw: Personal Assistants and the Harness-Training Effect
Orchard-Claw is the newest and smallest-scale of the three recipes, targeting everyday productivity work: reading and drafting email, managing a calendar, searching for information, and coordinating across tools. It trains a Qwen3-30B-A3B-Thinking model on just 200 synthetic tasks, evaluated on a benchmark Microsoft calls Claw-Eval.
On its own, Orchard-Claw completes 59.6 percent of Claw-Eval tasks within three attempts, a pass@3 measure. Paired with a stronger agent harness called ZeroClaw at evaluation time, that rises to 73.9 percent. The more interesting number sits underneath that result. Because Orchard Env can train an agent directly inside the same real deployment harnesses it will later run in, rather than a simplified training-only loop, the team trained Orchard-Claw across several harnesses, including ReACT, ZeroClaw, OpenClaw, and Codex. Under the Codex harness specifically, success climbs from 18.6 percent for an untrained model to 51.5 percent after Orchard training. That gap is the clearest illustration in the whole release of why the harness an agent trains inside, not just the data it trains on, changes what it can do.
The Infrastructure Numbers Behind the Claims
Microsoft Research’s blog post focuses on model results, but the project’s GitHub repository backs the infrastructure claims with benchmarks of its own, measured against comparable sandbox services.
Latency, Throughput, and Cost
Orchard Env’s average command-execution latency is 0.28 seconds, close to SkyPilot Code Sandbox’s 0.284 seconds and well ahead of E2B, at 0.747 seconds, roughly 2.7 times slower, and Modal, at 2.046 seconds, roughly 7.3 times slower. Launching 1,000 sandboxes in parallel finished with a 100 percent success rate in 26 seconds end to end, at roughly 154 commands per second, with an average per-sandbox creation time of 11.75 seconds.
Cost is the more concrete number for anyone budgeting a research project. Running 128 sandboxes, 2 vCPU and 8 GiB memory each, for 240 hours cost $3,362 on-demand, or $673 on spot instances, against a stated $7,078 to $10,305 for comparable managed sandbox services: roughly a tenfold reduction on the spot-instance path.
A Regression Test Against Plain Docker
The team also ran a direct substitution test on Terminal-Bench 2.0, swapping plain Docker for Orchard Env underneath three different models, GPT-4.1, MiniMax-M2.5, and Qwen3-8B-Thinking, and found no regression. Scores held steady or improved slightly across all three (34.1 to 35.1, 52.6 to 54.4, and 7.0 to 8.8, respectively), evidence that the environment layer is not quietly making evaluation easier or harder than a plain container.
Why This Is a Research Program, Not a Single Release
The August 3 blog post reads like a single announcement, but it is really a checkpoint in an ongoing project. Microsoft Research first posted the Orchard paper to arXiv in May 2026, alongside the Orchard Env code and the original SWE and GUI trajectory datasets. In June, a follow-on paper called OpenWebRL extended Orchard-GUI into full online, multi-turn reinforcement learning on live websites, which is where figures like the 67.0 percent Online-Mind2Web and 64.0 percent DeepShop results trace back to. In July, a second follow-on, OpenForge RL, extended the same environment to train agents inside their real deployment harnesses instead of simplified reimplementations: the work underpinning Orchard-Claw’s harness-transfer results. The arXiv paper itself has been revised as recently as July 30 to fold in that later work. Today’s post is the first time Microsoft has presented all three strands together as one coherent framework rather than three separate papers.
The team’s own roadmap suggests the current design is a deliberate first step rather than a finished system. Orchard Env is currently linear: a sandbox is created, driven through a task, and destroyed, and an entire multi-turn trajectory, an average of 47.5 turns for the SWE dataset, collapses down to a single pass-or-fail reward that has to be distributed back across every turn after the fact. Microsoft says it is adding pause-and-resume checkpointing and branching, so a rollout can fork multiple continuations from the same saved state, turning that credit-assignment problem into something the team can measure directly instead of estimate indirectly.
What It Means for Teams Building Agent Infrastructure
For a platform team outside Microsoft, the individual benchmark numbers matter less than the underlying claim: a shared, reusable sandbox layer is now a documented, MIT-licensed, comparatively cheap thing to stand up, rather than something every research or platform team quietly re-invents. Orchard Env deploys to Azure AKS with setup scripts that Microsoft says take about 20 minutes for provisioning, image building, and deployment, and the REST API is documented as a stable contract specifically so that projects in other languages can build against the same substrate without going through the Python SDK.
That framing matters beyond Microsoft’s own three recipes. Coding agents, browser agents, and computer-use agents each tend to accumulate their own bespoke sandboxing code inside whatever framework a team happens to be using. Orchard’s bet, backed by two follow-on papers and a public roadmap, is that the sandbox itself is closer to shared infrastructure, like a CI runner or an object store, than it is to a model-specific implementation detail. The code, trajectory datasets, and paper are public under an MIT license, so the numbers above are independently checkable rather than something readers have to take on faith.








No Comment! Be the first one.