Microsoft’s ThinkingBox Turns AI Agent Scores Into a Database Audit, Repeated 20 Times
Microsoft and Hugging Face’s ThinkingBox grades AI agents on the records they leave behind and repeats every task 20 times, and the paper’s own tables show how far a good single-attempt score sits...
A customer’s $745 kitchen appliance has been stuck at a courier exception for fifteen days past its delivery date. In the example that opens Microsoft’s new ThinkingBox post on Hugging Face, an AI agent makes nine tool calls, reads the refund policy correctly, opens a ticket, documents the timeline and then closes the ticket as resolved. The carrier exception is still open, so the required end state was a ticket on hold. “An AI grader checking tool calls would see nine well-formed ones,” the post says. “The database is what disagrees.”
Table Of Content
- What ThinkingBox Grades
- The Method Predates the Benchmark
- A Trajectory Is a Claim, and Most Failures Look Finished
- One Success Is Not Reliability
- Always, Sometimes and Never
- Two Definitions of the Same Number
- What Consistency Costs
- Where the Failures Come From
- The Label Is Broad
- The First Matching Rule Wins
- The Share Varies by Domain
- What the Numbers Cannot Tell You
- What to Take From It
ThinkingBox is a sandbox and a benchmark, “built by the Microsoft Copilot Studio team in partnership with Toloka,” according to the post, with collaborators from four universities who interned at Microsoft. It grades 507 stateful business workflows by the records an agent leaves behind, and it runs every task 20 times against each model. The post went up on October 3. The paper behind it, arXiv 2608.19741, is now in its fourth version.
This piece reads the post against that paper, the released task data and the 2024 benchmark whose method ThinkingBox builds on, and it recomputes the arithmetic. Most of the headline figures reconcile. The more useful findings sit in the gaps: two definitions of the same reliability number, a leaderboard row that appears only in the blog, and a failure label whose 79.9 percent share depends on a first-match rule.
What ThinkingBox Grades
Each task hands the agent a domain policy, a set of Model Context Protocol (MCP) tools, a starting backend state and a simulated user who holds private details and releases them only when asked. Every attempt gets an isolated session with freshly initialized state. When the conversation ends, a side-effect extractor works out what changed, and the post says deterministic judges compare that against the required end state, “accepting any trajectory that produces the right outcome while rejecting wrong, missing or extra effects.” The paper describes the primary check as a comparison of the final database state with the golden state using “deterministic, hash-based assertions.” Of the 507 tasks, 477 are graded on state alone and 30 add a response rubric.
We did not run the harness, which the post says was “Tested on Linux and WSL” with Docker. We read the released task file instead, through Hugging Face’s dataset server, and the counts match the paper.
| Domain | Tasks | Tasks with a response rubric | Median reference tool interactions |
|---|---|---|---|
| Retail and e-commerce | 98 | 0 | 4 |
| Travel and hospitality | 104 | 15 | 9 |
| Auto insurance | 100 | 0 | 3 |
| Neobank support | 104 | 15 | 6 |
| Consulting IT and HR | 101 | 0 | 6 |
| All five domains | 507 | 30 | 6 |
Source: the tasks configuration of microsoft/ThinkingBox-Bench (release thinkingbox-bench-v1.0), counted by us. Reference tool interactions are the entries in each task’s expected_tool_interactions field, which mixes lookups and writes.
The 507 tasks carry 3,102 reference tool interactions across 108 distinct tools, with a median of 6 per task and a range of 1 to 19. Only 6 tasks list a single interaction and 459 list three or more. The benchmark passes any trajectory that reaches the same terminal state, so these lists show how much work a reference solution does, not the only path the grader allows.
The task the post adapted its opening example from, sandbox_external_retail_group1.py:test_case_ST003_006, is in the released data. It ends with a reference interaction that sets a ticket’s status to hold, and its response-rubric list is empty. That matters. The post names two things wrong with the agent’s ending: the ticket status, and a customer who “never got a real answer.” The post says the check that fails is “a single field,” and the released data lists no rubric that would score the second problem.
The Method Predates the Benchmark
End-state grading and repeat-trial reliability were not introduced here. The 2024 τ-bench paper by Shunyu Yao and co-authors compared “the database state at the end of a conversation with the annotated goal state” and proposed pass^k, a metric “to evaluate the reliability of agent behavior over multiple trials,” defined as the chance that all k trials succeed. The ThinkingBox paper credits it for that metric and describes its own analysis as “following prior reliability-oriented agent evaluation.” The Hugging Face post mentions neither τ-bench nor pass^k.
What ThinkingBox contributes, by its own paper’s account, is breadth and plumbing: five domains and 507 tasks, isolated MCP-compatible tool sessions, hash-based assertions with a side-effect extractor, a cost metric, and an environment packaged for OpenEnv, the interface library we covered in June as a common socket for open-source agent training. That is a real contribution. It is also an extension of a method with a 2024 track record, and the post would be easier to place if it said so.
A Trajectory Is a Claim, and Most Failures Look Finished
The paper’s evidence for grading state rather than transcripts is a retrospective ablation over 121,680 trials, which is 12 models, 507 tasks and 20 trials each. Of those, 79,853 failed the executable checks. The authors then asked how many of the failures would pass three “intentionally weak completion proxies,” each stricter than the last.
| Weak proxy (each includes the one above) | Failed trials it would accept | Share of the failures | Share of all 121,680 trials |
|---|---|---|---|
| Clean termination: a non-empty final response with a done marker and no question | 67,763 | 84.86% | 55.69% |
| Clean termination plus at least one state-changing tool call | 64,586 | 80.88% | 53.08% |
| Both of those plus no explicit error in the final tool response | 53,697 | 67.24% | 44.13% |
Source: arXiv 2608.19741v4, Table 16, for the trial counts and shares of failures. The last column is our division by 121,680.
Read from the other side, the executable checks found wrong field values in 77.61 percent of the failures, unintended extra effects in 43.30 percent and missing required effects in 25.36 percent, and the post notes that these findings overlap. The state check passed 41,827 trials, 34.37 percent of the set. The strictest proxy alone would have accepted another 53,697 trials that failed it, a group larger than the successes.
The blind spot runs the other way too. Because the verdict on 477 tasks rests on terminal state alone, the paper’s limitations section concedes that “a trial that executes the correct state transition while misreporting it to the user might still be scored as a success.” ThinkingBox therefore measures correct, policy-compliant backend outcomes. It does not score the reply the customer reads, except on the 30 tasks with a rubric, 15 in travel and 15 in neobank support.
One Success Is Not Reliability
The post reports three numbers per model: pass@1, the share of all attempts that succeeded; pass@20, the share of tasks solved at least once; and “observed 20/20,” the tasks that passed every one of their 20 attempts. The paper’s Table 15 gives the counts behind them, and for the 14 models we checked the counts reconcile: each pass@20 equals the solved tasks over 507, and each success count in the paper’s cost table equals pass@1 times 10,140 attempts.
Always, Sometimes and Never
The table sorts each model’s 507 tasks into three buckets. The last column appears to be the quantity behind the post’s Figure 2, which shows how much of each model’s single-attempt score survives 20 repeats. Computed as the always-passed share divided by pass@1, it equals the fraction of a model’s successful attempts that come from tasks it never fails. It reproduces the post’s 78 percent for GPT-6 Astra, 71 percent for Claude Opus 5 and Claude Opus 5.5, and about 8 percent for GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro.
| Model | pass@1 (%) | Always (20/20) | Sometimes | Never (0/20) | Share of successful attempts from always-passed tasks |
|---|---|---|---|---|---|
| Claude Opus 5 | 66.50 | 241 | 160 | 106 | 71.5% |
| GPT-5.4 | 65.36 | 128 | 334 | 45 | 38.6% |
| GPT-5.6-sol | 61.91 | 82 | 358 | 67 | 26.1% |
| Claude Sonnet 4.6 | 59.19 | 102 | 347 | 58 | 34.0% |
| GPT-6 Astra | 58.31 | 231 | 129 | 147 | 78.1% |
| Kimi-K3 | 57.37 | 68 | 408 | 31 | 23.4% |
| Qwen3.8-27B | 51.70 | 38 | 415 | 54 | 14.5% |
| DeepSeek-V4-Pro | 43.26 | 18 | 411 | 78 | 8.2% |
| Kimi-K2.6 | 37.66 | 16 | 411 | 80 | 8.4% |
| GLM-5.1 | 33.19 | 14 | 340 | 153 | 8.3% |
| Claude Opus 4.6 | 32.09 | 70 | 215 | 222 | 43.0% |
Source: arXiv 2608.19741v4, Table 15, for pass@1, never-solved and always-passed counts. The Sometimes column and the last column are our arithmetic.
Two models with almost the same pass@1 can be different products. GPT-6 Astra (58.31 percent) puts 231 tasks in the always bucket and 129 in the middle. Kimi-K3 (57.37 percent) puts 68 in the always bucket and 408, or 80 percent of the benchmark, in the middle. The first mostly solves a task every time or never. The second gets most of the benchmark right only some of the time. A lower average can hide steadier behavior too: Claude Opus 4.6 has a lower pass@1 than Kimi-K2.6, 32.09 against 37.66 percent, yet passes 70 tasks every time against 16.
Our reading: if every attempt were an independent coin flip at a model’s average rate, passing one task 20 times running would be rare. At Claude Opus 5.5’s 67.16 percent it works out to 0.035 percent. The post reports 47.53 percent of the benchmark passed every time. So the reliability a model shows depends heavily on the task: some workflows are solved every time and others rarely, and sorting tasks into always, sometimes and never says more about where an agent can safely run than one averaged score does.
Two Definitions of the Same Number
The paper’s abstract says Kimi-K3 falls “from 57.37% pass@1 to 17.60% pass^20.” The post says only 68 of 507 tasks, 13.41 percent, “succeed in all 20 attempts.” Both come from the same runs. The paper’s metric is a plug-in estimate: for each task it takes the observed success rate, raises it to the 20th power and averages the results, so a task that passed 19 of 20 attempts contributes about 0.36 instead of zero. The estimator introduced with τ-bench, which the paper also writes down, is the chance that a random set of k attempts all succeed, and when k equals the 20 trials it reduces to the literal count: 1 if all 20 passed, 0 otherwise. The paper explains its choice by saying the unbiased version offers “little resolution across models.” The post uses the literal count and adds, “No estimator, no smoothing.”
| Model | Tasks passing 20/20 | Literal share (tasks / 507) | Paper’s pass^20 | Difference (points) |
|---|---|---|---|---|
| Claude Opus 5 | 241 | 47.53% | 47.53% | -0.00 |
| GPT-6 Astra | 231 | 45.56% | 46.89% | 1.33 |
| GPT-5.4 | 128 | 25.25% | 30.62% | 5.37 |
| Claude Sonnet 4.6 | 102 | 20.12% | 25.36% | 5.24 |
| GPT-5.6-sol | 82 | 16.17% | 22.00% | 5.83 |
| Kimi-K3 | 68 | 13.41% | 17.60% | 4.19 |
| Claude Opus 4.6 | 70 | 13.81% | 15.53% | 1.72 |
Source: arXiv 2608.19741v4, Table 15. The literal share is our division by 507.
The plug-in is never lower than the literal count. In this table it runs 4.19 to 5.83 points higher for GPT-5.4, GPT-5.6-sol, Claude Sonnet 4.6 and Kimi-K3. For Claude Opus 5 the two agree to two decimals, which implies almost none of its tasks sit just below the line. The two definitions even swap the order of Claude Opus 4.6 and Kimi-K3: 70 always-passed tasks against 68 by the count, but 15.53 percent against 17.60 percent by the paper’s metric. The post’s closing advice is “Report a repeat metric, and define it.” The distance between the post and the paper is that advice in action.
One more difference between the documents: the post ranks Claude Opus 5.5 first at 67.16 percent pass@1 and credits it with 241 always-passed tasks, the same count as Claude Opus 5. That row is not in the paper’s version 4 tables, which name Claude Opus 5 and GPT-5.4 as the highest overall, and the post’s footnote says its Opus 5.5 pricing comes from Anthropic’s site while the other models are priced from OpenRouter. The figures may be right, but they are the post’s alone. The post also lists 18 models, as the paper does, though not the same 18: it adds Claude Opus 5.5 and does not list MiniMax-M2.5, which scored 0.19 percent.
What Consistency Costs
The post prices whole campaigns instead of single calls. It takes each model’s recorded token usage across 507 tasks and 20 runs, prices it at list rates, and divides by successes (cost per successful task attempt) or by always-passed tasks (cost per dependable task). The paper’s Table 23 holds the same figures for all 18 of its models, defines both in an equation, and calls them “usage-based estimates, not actual cloud bills” that exclude negotiated rates, the simulator and judge, infrastructure and downstream failure.
| Model | Campaign cost, 20 runs | Tasks passing 20/20 | Cost per dependable task | Added tasks vs the row above | Added cost per added task |
|---|---|---|---|---|---|
| GPT-5.6-sol | $800.00 | 82 | $9.76 | n/a | n/a |
| GPT-5.4 | $869.80 | 128 | $6.80 | 46 | $1.52 |
| GPT-6 Astra | $1,720.60 | 231 | $7.45 | 103 | $8.26 |
| Claude Opus 5.5 (post only) | $1,880.77 | 241 | $7.80 | 10 | $16.02 |
Source: the post’s Table 3 for the first four columns. The last two columns are our arithmetic.
The cheapest way to get a right answer is not the cheapest way to get a dependable one. GPT-5.6-sol has the lowest cost per success, $0.127, and the highest cost per dependable task of the four rows above, $9.76. Climbing the table gets steeper: 46 more dependable tasks for $69.80 going from GPT-5.6-sol to GPT-5.4, 103 more for $850.80 going on to GPT-6 Astra, and the last 10 tasks that Opus 5.5 adds cost $16.02 each. The paper’s version of the top step uses Claude Opus 5, which also passes 241 tasks but costs $3,206.00 for the campaign, so its 10 extra tasks over GPT-6 Astra cost $148.54 each.
These are our divisions of estimated list-price costs from one date, so they are a way to compare, not a procurement number. The paper itself says cost per dependable task is “not an estimate of future completion cost” and that the comparisons “do not certify future reliability.”
Where the Failures Come From
The post’s practical headline is that “roughly four in five failures are tool handling, not reasoning.” The paper’s number is an average Tool Usage share of 79.9 percent across models, with Wrong State Update at 10.3, Incomplete User Resolution at 7.0 and No State-Changing Action at 2.9. Three details change how to read it.
The Label Is Broad
The paper assigns Tool Usage when a trace contains “an explicit tool error, a failed precondition, or an unsuccessful lookup from which the agent does not recover,” and notes that “this category extends beyond malformed calls.”
The First Matching Rule Wins
Tool Usage is checked first, and the paper says that when several signatures apply, “the earliest applicable rule determines the label.” A trace with an unrecovered lookup and a wrong update therefore counts as Tool Usage.
The Share Varies by Domain
The paper’s Table 17 puts Tool Usage at 84.0 percent of failures in travel and 81.7 percent in neobank IT, but 51.3 percent in auto insurance, where 29.9 percent of failures are wrong state updates.
The paper itself calls the labels “reproducible observable diagnostics rather than unique causal explanations,” and the post repeats the caveat. The post still concludes, “That is a retry and error-recovery problem before it is a model problem,” and suggests classifying tool errors so retries target the recoverable ones, cutting the tool surface, and requiring human approval for changes that are costly to reverse. Then it adds: “We have not measured the lift from any of these on this benchmark.” That makes error recovery a hypothesis the environment can test, not a result. It is a good hypothesis, because the most dependable models’ failures are almost all labeled this way, 97.0 percent for GPT-6 Astra and 96.4 percent for Claude Opus 5. It may also be partly a product of the precedence rule.
What the Numbers Cannot Tell You
- The tasks are synthetic. The paper says “Tasks are synthetic reconstructions from a non-public source collection and are not claimed to represent the distribution of enterprise work.” Each retained task also has a single golden terminal state, so workflows with several defensible resolutions are excluded by construction.
- One simulated user plays every customer. The same GPT-5.4-mini simulator is the user in every trajectory and the judge for the 30 rubric tasks, it stays cooperative, and it allows at most ten follow-up turns. The paper adds: “Because the simulator also shares a model family with one of the evaluated agents, interaction-style effects cannot be ruled out.”
- The simulator behaves differently with different agents. A separate audit of 19,390 generated user turns labeled 6.33 percent of them ungrounded, from 0.39 percent in Qwen3.6-27B conversations to 23.88 percent in GPT-5.2’s. Two LLM judges produced those labels, and a human check of 120 turns gave “substantially lower confirmation of ungrounded labels.”
- Infrastructure failures count against the model. The paper says “Decoding failures, test-execution failures, API errors, and other system errors are treated as unsuccessful trials in the reported metrics.” The post says the OpenEnv adapter writes operational failures to a separate file “so they can be rerun rather than silently mixed with model outcomes.”
- The results are the authors’ own, and rerunning them takes infrastructure. The ThinkingBox code is MIT licensed, the data is under CDLA-Permissive-2.0 according to the post and the dataset card, and the OpenEnv environment is BSD-3-Clause. A rerun still needs Docker on Linux or WSL plus model endpoints for the agent, the simulated user and the judge.
- Part of the release is still pending. The post lists the RL training repository as “coming soon,” while the paper says its training code is released and links that repository. GitHub returned a 404 for that repository, from both its web address and its API, to anonymous requests on October 4.
What to Take From It
- Put a state check next to every trace check. On this benchmark a clean exit, a state-changing call and no final tool error together described 53,697 failed trials. A grader that reads the agent’s own summary or its tool-call log is measuring a claim. Diff the records the agent could touch instead, and keep any rubric for what is not in the database. Our multi-stage agent pipeline tutorial shows a deterministic validation gate catching extraction and reasoning bugs at each stage, which is the same principle at small scale.
- Repeat the task set, then sort tasks into always, sometimes and never. Pick k from the use case, say whether the number is best-of-k or every-of-k, and say which estimator produced it, as the post advises. If a rubric relies on an LLM judge, check the judge against human ratings first; our Cohen’s kappa tutorial walks through that.
- Price the dependable unit. Divide campaign cost by the tasks that passed every run, and look at the marginal cost of the last few. A model that is cheap per success can be expensive per dependable task.
- Publish the rule that labels failures. If the first matching rule wins, say which rule is first. Treat tool-error recovery as the first hypothesis to test, since the post has not yet measured what it would change.
The post says “The useful part of this work is not our pass@1 leaderboard. It is the environment.” The tables above point the same way. The leaderboard is one team’s runs on synthetic tasks with a single simulated customer, and its definitions shift between the post and the paper. The habits it demonstrates are portable: ask the database what changed, run the task again, and report which of those two numbers you are quoting.
How we checked: we read the Hugging Face post, version 4 of the paper, the τ-bench paper, the repositories and the released task file. The tables built on the paper’s and the post’s numbers are our arithmetic, and we say so where it applies. We did not run the benchmark.








No Comment! Be the first one.