Skip to content

Your Agent Passed the Test. The Database Didn't.

#agent-evaluation #harness #llm-ci #benchmarking #reliability #mcp

A customer's $745 appliance has been stuck at a courier "exception" for fifteen days. Your agent makes nine careful tool calls: pull the order, check tracking, look up her profile, search the refund policy twice, confirm no ticket exists, open one, document the timeline. Then it closes the ticket as resolved and asks if there's anything else it can help with.

Two things are wrong. The carrier exception is still open, so the required end state was "hold", and the agent set it to "solved". And the customer never got an actual answer, because she doesn't qualify for compensation, but nothing in the reply says so.

A grader watching tool calls would call that a perfect run: nine well-formed calls, a clean exit, no errors. The database is what disagrees. That gap is why agent evaluation infrastructure is more interesting than any single model release right now.

Three recent releases show the same shift from different angles. Microsoft and Hugging Face published ThinkingBox, a benchmark that grades agents on terminal backend state and side effects instead of conversation traces. DeepSeek Harness shipped an experimental compatibility layer for Claude Code Mods, betting that the harness itself is a product. ByteDance published FastCI, which cuts GPU-intensive CI latency for LLM training frameworks by 77.5%. None of these are models. All of them are infrastructure, and infrastructure is where agent reliability actually gets decided.

A tool call is not an outcome ​

ThinkingBox is two things: an agent sandbox and a benchmark. The sandbox runs an agent against isolated MCP tool sessions; the benchmark grades what the agent leaves behind. Each of the 507 stateful business workflows defines a starting backend state, a user goal, the available MCP tools, a domain policy, and executable checks over the terminal state. A simulated user holds private context (a booking reference, a preference, a date of birth) and releases it only when asked.

Every attempt gets a freshly initialized backend. Two attempts of the same task never share a database row or cached tool state, which is what makes repeated trials meaningful in the first place.

At the end, a side-effect extractor computes what actually changed, and deterministic judges compare it with the required end state. 477 of the 507 tasks are graded on state alone; 30 add narrow binary rubric questions for requirements with no clean database value, like "did the agent disclose that this isn't guaranteed?" The model never sees the golden state, the assertions, or the grading internals.

Then the numbers. In a common-set ablation covering 121,680 valid trials across 12 models, 79,853 attempts failed the executable checks. Of those failures, 67.24% terminated cleanly, invoked a state-changing tool, and reported no final tool error. A grader watching tool calls would wave two-thirds of these failures through. Executable checks still found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%.

Key numbers

  • 121,680 valid trials across 12 LLM models
  • 67.24% of failed runs ended cleanly, with no tool error
  • 77.61% of those clean failures left wrong field values in the database
  • 79.9% of all failures traced to tool usage, not reasoning

A trajectory is a claim. Database state is the evidence. Repetition is the trust test.

One success is not reliability ​

Every task runs 20 times from an identical clean backend, and ThinkingBox reports three different numbers.

MetricWhat it measuresWhat it answers
pass@1Share of all attempts that succeededHow does it usually do?
pass@20Share of tasks solved at least once in 20 triesCan it ever do this? Breadth.
Observed 20/20Tasks that passed all 20 recorded attemptsCan it always be correct?

The leaderboard view reads like an ordinary capability ranking. Claude Opus 5.5 leads overall at 67.16% pass@1; Kimi-K3 is the strongest open-weight model at 57.37%, within a point of GPT-6 Astra. Domains matter as much as models: Claude Opus 4.6 scores 68.62% on retail but 8.30% on auto insurance.

The consistency view tells a different story. GPT-6 Astra retains 78% of its single-attempt rate across 20 runs; Claude Opus 5.5 and Claude Opus 5 each retain 71%. At the other end, GLM-5.1, Kimi-K2.6, and DeepSeek-V4-Pro each keep about 8%. The chart below shows how much of each model's single-attempt score survives 20 repeats.

The gap is the whole story. Kimi-K3 solves 93.89% of the benchmark at least once, the broadest coverage of any model tested, with only 31 tasks defeating it entirely. It also passes just 68 of 507 tasks on all 20 attempts, 13.41%. Claude Opus 5 inverts that: fewer tasks solved at least once (79.09%), but 47.53% of the benchmark completed on every single attempt.

A newer model doesn't fix this. Claude Opus 5.5 scores higher than Claude Opus 5 on the single-attempt average (67.16% against 66.50%) and solves more tasks at least once. It passes exactly the same number of tasks on all 20 attempts: 241. Half a point of headline accuracy bought zero additional dependability.

Quick Take: The models at the top of a leaderboard aren't the ones you'd trust with a production database, so pick the column that matches your deployment: pass@1 if you can retry, observed 20/20 if you can't.

What consistency costs ​

Capability comparisons usually stop at the score. Anyone deploying wants to know what a successful unit of work costs. ThinkingBox priced each model's recorded token usage from the full 507 × 20 campaign at undiscounted list rates from a single provider endpoint per model, then divided by the number of attempts that succeeded.

Cost per successful task attempt = estimated cost for 507 attempts ÷ (507 × pass@1)

GPT-5.4 works out to $43.49 for 507 attempts, a 65.36% pass@1, so $0.131 per successful attempt. Claude Opus 5 is the clearest case: at $0.475 per success and 66.50% pass@1, it is both more expensive and less accurate than Claude Opus 5.5 at $0.276 and 67.16%.

But cost per success rewards a model that is cheap and often right, not one that is right every time. So ThinkingBox also computes cost per dependable task: the full 20-run campaign cost divided by the number of tasks passed on all 20 attempts. GPT-6 Astra costs 20 × $86.03 = $1,720.60 for the campaign and passes 231 tasks on every attempt, so $7.45 per dependable task.

ModelTasks passing 20/20Est. cost, 20 runsCost per dependable task
GPT-5.4128 (25.25%)$869.80$6.80
GPT-6 Astra231 (45.56%)$1,720.60$7.45
Claude Opus 5.5241 (47.53%)$1,880.77$7.80
GPT-5.6 Sol82 (16.17%)$800.00$9.76
Claude Opus 5241 (47.53%)$3,206.00$13.30
Claude Sonnet 4.6102 (20.12%)$1,587.60$15.56
GPT-5.244 (8.68%)$878.00$19.95
Kimi-K368 (13.41%)$1,406.40$20.68
Qwen3.8-27B38 (7.50%)$925.80$24.36

Rank by consistency and the picture flips. GPT-5.4 is the cheapest at $6.80 per dependable task, though only 128 tasks meet the bar. GPT-6 Astra reaches 231 tasks at $7.45. Claude Opus 5.5 reaches the joint-highest 241 at $7.80. GPT-5.6 Sol, the cheapest per single success at $0.127, costs $9.76 per dependable task. The cheapest way to get a right answer is not the cheapest way to get a dependable one.

These are estimated dollars, not invoices, and they price single successes and full campaigns, not production serving. But as a comparative index, the message holds: consistency is a line item.

Failure signatures: it's the tools, not the model ​

Each failed trace gets one deterministic diagnostic signature, and the finding is concrete: roughly four in five failures are tool handling, not reasoning. The unweighted averages across models are lopsided. Tool usage accounts for 79.9% of failures, wrong state updates for 10.3%, incomplete user resolutions for 7.0%, and no state-changing action for 2.9%.

The practical pattern is simple: agents usually get far enough to attempt the workflow, then fail to recover from tool errors, failed preconditions, or empty lookups. That is a retry and error-recovery problem before it is a model problem. Domain difficulty is wide too: retail averages 59.52% pass@1 across the tested models while auto insurance averages 33.83%. If most failures are tool handling, the cheapest fixes are retry logic, error classification, and a smaller tool surface. Prompt engineering and bigger models are second-order.

You can run this yourself. The customer example above is sandbox_external_retail_group1.py:test_case_ST003_006, and the executable check that fails is a single field: the ticket status is "solved" where the required end state is "hold". ThinkingBox and ThinkingBox-Bench are on Hugging Face; the benchmark sits behind the OpenEnv interface, and each finished episode returns a binary pass/fail reward. The same interface can be reused in training workflows for non-benchmark scenarios.

The harness is the product now ​

While ThinkingBox grades agents, the harness layer is consolidating underneath them. On Hugging Face you can already find spaces like FineEnvs/multi-harness-rl that wrap multiple RL harnesses behind a single interface. And DeepSeek Harness released v0.2.1-alpha.1 over China's National Day break with a more aggressive bet: an experimental compatibility layer for Claude Code Mods.

Claude Code Mods are event handlers written in JavaScript or TypeScript that run inside Claude Code, intercepting tool calls, prompt submissions, and UI rendering. Mods let you change the tool itself: Token Weather renders context usage as a weather forecast (sunny when it's low, stormy when it's filling up), Blast Radius shows the impact before a high-risk command executes, Replay Theater lets you step through file modifications. You can even describe a mod to Claude and have it write and hot-load the plugin into the current session.

DeepSeek's answer is less a copy than a philosophical one-up. The Harness team's position: "Everything is a Plugin". Models, tools, skills, sessions, sandbox, file system, agent loop, task orchestration, UI, all of it composes as replaceable plugins. The lead was direct about the architectural difference: Claude opened part of its capabilities as mods, while DSH designed pluggability in from the start. The new compatibility layer is a bridge that lets mods written for Claude Code's interface run inside DeepSeek Harness's plugin system. The release notes are explicit about the intent: verify that the Claude Code Mods API is roughly a subset of what DSH plugins already do.

What the community is saying: the release made Zhihu's hot list, and most of the discussion circled the word "subset". Some read it as a polite flex; others read it as a roadmap signal that DSH wants to be the compatibility layer other harnesses plug into. The skeptical read has evidence. When I went through the documentation, the three demo mods are runnable, but Claude Code's built-ins (diff, agents-md, sec-default, telemetry) are marked non-runnable: some events, interfaces, and UI hooks aren't bridged yet. Loading also goes through DSH's plugin mechanism, so you can't copy Claude's mods over as-is. The bridge proves the concept; it isn't drop-in compatibility.

The other new feature is arguably more useful: a creation mode where the agent inspects the running system and writes plugins on the fly, plus a developer toolkit for viewing session logs and debugging. The direction is clear. If the harness is extensible by design, every harness improvement becomes a shared layer, and no single vendor owns the agent runtime.

The other bottleneck: CI for LLM frameworks ​

None of this matters if you can't iterate quickly on the frameworks themselves. That's the problem FastCI tackles.

CI for LLM training frameworks isn't normal CI. Tests are GPU-intensive and often involve complete model training runs or evaluations, so the pipeline stalls as frameworks evolve at a rapid pace. FastCI, built and deployed at ByteDance, does four things: use runtime evidence to select tests affected by a change, prune tests that execute changed code in equivalent contexts, prioritize high-risk tests so failures surface earlier, and trim test workloads along dimensions outside each test's intended validation scope.

Evaluated on ByteDance's LLM training framework CI workload, the numbers are stark. CI latency down 77.5%, GPU resource usage down 63.9%, modified code coverage retention up 3.2%. A 77.5% latency cut is the difference between a merge queue that stalls for half a day and one that lands in a couple of hours. The 63.9% GPU reduction is recurring spend: CI runs on every change, so the savings compound across every merge.

The chain is only as reliable as its links. FastCI keeps the training frameworks moving, ThinkingBox measures the agents those frameworks produce, and DeepSeek Harness runs those agents in production-shaped environments. All three are infrastructure problems dressed up as model problems, and the field is finally building the tools to match.

Common pitfalls ​

These are the mistakes I keep seeing in evaluation setups, from benchmark design to the deployment decision.

  1. Grading tool calls instead of state. Watch a conversation trace and an agent that writes "solved" when it should write "hold" looks perfect. The equivalent in production is trusting the model's summary of what it did instead of checking the terminal state of the system it touched. Check the end state before you commit.

  2. Reporting pass@1 as reliability. A 67.16% pass@1 model fails roughly one attempt in three. If a customer interaction gets one attempt, that's your real failure rate. Report the observed 20/20 rate and demand it from vendors.

  3. Choosing on per-attempt cost when you can't retry. GPT-5.6 Sol is cheapest per success at $0.127 and close to the most expensive per dependable task at $9.76. The right denominator depends on whether a failed attempt burns a customer or just burns tokens.

  4. Tuning the model when the failure is in the tools. With 79.9% of failures classified as tool usage, retry logic, error classification, and a smaller tool surface are the first lever. A prompt rewrite won't fix a failed precondition.

  5. Assuming the newest model is the most dependable. Claude Opus 5.5 and Claude Opus 5 pass the exact same 241 tasks on all 20 attempts. Newer and smarter is not more consistent. Re-run the 20-trial check on any model before it touches real records.

One thing to remember: the measurement tools now exist, and they grade state, not claims. ThinkingBox-Bench runs through the OpenEnv interface and returns