Skip to content

Model + Harness = Agent: The Scaffolding Is the Story

#ai-agents #agent-harness #coding-agents #reinforcement-learning #tool-orchestration

The same model, two harnesses, four extra wins ​

Here's the cleanest agent benchmark I've seen this year. Thirty hard agentic tool-use tasks against real apps, eight harnesses, every one wired to the same model: DeepSeek V4 Flash. Same model, same tasks, same tools, same 900-second ceiling per task. The only variable is the harness.

Claude Code passed 16 of 30 tasks. Pi passed 20. Cost per successful task: $0.195 for Claude Code, $0.028 for Pi. The model didn't change. The harness did, and it moved the pass rate by 13 points and the cost per win by 7x.

That's the story of the agentic wave right now. DeepSeek has a formula for it: Model + Harness = Agent. The past month has made that hard to argue with.

Model + Harness = Agent ​

DeepSeek published that formula alongside the open-source DeepSeek Harness developer preview on August 13, 2026, hours after a messy V4 Pro launch that saw the official announcement pulled, then quietly re-posted. The timing was deliberate. The model gets the attention, but the harness is the product.

The framing is simple. The model is the brain: it reasons, plans, understands intent. But a language model only outputs text. It can't touch a file, run a command, or call an API. The harness is the body and nervous system: tool calling, file I/O, sandboxing, permissions, task orchestration. No harness, no agent. Just a very expensive chat endpoint.

DeepSeek's version makes a specific design bet. "Everything is a plugin" is the core principle, and it goes further than most: the agent loop itself is replaceable, not just the tools. The time-space composability framework from PKU and DeepSeek splits components into two dimensions. Temporal composability uses reversible effects, so unloading a component undoes its side effects. Spatial composability uses declarative dependency management, so components declare what they need and the runtime wires it up. No privileged core. You compose an agent from a config file.

That's a different strategy from OpenAI and Anthropic. They integrate downward: build a polished harness, bind it to their own models, own the experience. DeepSeek spreads sideways: open-source a model-neutral harness and try to become the standard layer underneath everyone's agents.

The loop every agent runs looks like this:

The harness is everything wrapped around that loop: permissions, sandboxing, session state, the tools themselves, and increasingly the reward signal that trains the next model.

Two philosophies, one question ​

The harness debate was already running before DeepSeek shipped. Pi and Claude Code are the two poles.

Pi is the coding agent Mario Zechner built after getting fed up with Claude Code. He was a hardcore user: he wrote cchistory to track its system prompt changes and patched the binary to add features Anthropic hadn't shipped. His complaint wasn't that Claude Code was bad. It was that it kept changing underneath him. New releases shifted the system prompt and tool definitions, broke his workflows, changed model behavior. So he built the opposite: four tools (read, write, edit, bash), a system prompt under 1,000 tokens, no MCP, no permissions, MIT licensed. It now has 85k+ GitHub stars and powers OpenClaw.

Claude Code is the default the whole category gets measured against. Ten-plus tools, subagents, plan mode, MCP as client and server, skills, plugins, hooks, checkpoints, five permission modes, and it runs in the terminal, VS Code, JetBrains, desktop, web, and mobile. More than 80% of Anthropic's own engineers use it daily.

Both tools answer one question differently: how much harness does a frontier model actually need?

Anthropic's answer is drifting toward Pi's. In July 2026 they cut over 80% of Claude Code's system prompt for the Claude 5 generation models with no measurable loss on their coding evals. Their bet is that model and harness get trained together, so today's scaffolding becomes tomorrow's model behavior. Pi's bet is that the scaffolding was never load bearing in the first place. Anthropic quietly deleting most of its own system prompt is the strongest validation of Pi's thesis you could ask for. Zechner just got there a year early.

DimensionClaude CodePi
DesignBatteries included: 10+ tools, subagents, plan mode, MCP, skillsFour tools: read, write, edit, bash
System prompt~10K-14K tokens pre-Claude 5, cut by 80% for Claude 5 genUnder 1,000 tokens including tool definitions
ModelsClaude only, tuned end to end20+ providers, 300+ models, mid-session switching
PermissionsDeny by default, five modes, sandboxingNone. Full system access from the first prompt
Session modelLinear log with checkpoints and rewindBranchable session trees with fork and rewind
ExtensibilityShell hooks, MCP servers, skills, pluginsTypeScript extensions inside the agent process
Pricing$20/month Pro or API ratesFree, MIT. You bring API keys or run local models
Cost per success (same-model eval)$0.195$0.028

What the benchmark actually measures ​

The 30-task eval is the cleanest data point we have. Same model, same tasks, same tools, only the harness changes. Whatever gap shows up is the wrapper, not the model.

Claude Code posted the faster median time: 122.7 seconds against Pi's 132.2. It just burned 741,659 tokens per task getting there. Pi used 558,885. The overhead is the story. Claude Code's batteries-included harness spends a third more tokens to finish slightly faster, and you pay for those tokens.

The cost column is brutal. Claude Code: $3.12 total, $0.195 per success. Pi: $0.56 total, $0.028 per success. Codex landed between them: 16/30 for $1.29, which works out to about $0.081 per success. If you're running agents at scale, that's not a rounding error. That's a budget line.

Token discipline shows up in another place too. In the 36kr test of DeepSeek V4 Pro, the same model in DeepSeek Harness hit a 99% cache hit rate versus 98% when routed through Claude Code. One point sounds small. In long-context tasks, cache hit rate is the difference between paying for full re-processing and paying for a small delta, and it compounds across every turn.

The same test surfaced something stranger. The first attempt used a detailed prompt: 1,817 words of source text plus a five-item requirement list covering scene elements, camera moves, lighting, and proportions. It produced a rough, broken result. The second attempt cut the prompt to a single goal: read the source text, build the Shire birthday party in Three.js, follow the pacing of the original, save a runnable HTML file. It worked on the first try. Same model, same harness, different prompt.

The pattern is showing up across the field. Community reports around GPT-5.6 Sol's launch described carefully optimized prompt chains and flow skills sending the model into loops, while plain goal statements worked. Prompts are shifting from writing a script to setting a goal. For top models, the era of the 2,000-word prompt may be over.

Quick Take: The model decides what an agent can do, but the harness decides what it actually gets done.

RL data is what trains the harness ​

Harnesses don't just wrap models. They generate the data that trains the next models. NVIDIA's Nemotron-RL-Agentic-Terminal-Pivot-v1 dataset is a direct look at how that works.

The dataset holds 31,111 training samples from 630 unique ATCB seed tasks. Each seed task is a containerized Linux environment with a natural-language instruction, a hidden reference solution, and an automated verifier. The tasks are real operational work: building and repairing data pipelines, auditing telemetry stores for injected faults, diagnosing crashed services and CI builds, security and compliance workflows like access-log audits, SIEM triage, PII handling, and TLS fixes. Set in SCADA, firmware/OTA, cold-chain, satellite, HPC, healthcare, finance, and media archive scenarios.

Trajectories were generated by the Terminus-2 agent running in the Harbor execution harness, with GLM-5.1 as the teacher model. Only verifier-passing trajectories were kept, at most five per task. Each valid assistant turn became one next-action training sample. That's the key design choice: every sample is a single decision point, not a full trajectory. The prompt is the task instruction plus the interaction history up to that point. The expected answer is the reference next action: analysis, plan, command keystrokes, and a task_complete flag.

MetricValue
Training samples31,111
Unique seed tasks630 (ATCB)
Source trajectories2,716
Median samples per task45
Teacher modelGLM-5.1 driving Terminus-2 in the Harbor harness
Reward mechanismRLVR via terminus_judge in NeMo Gym
LicenseCC-BY-4.0

The dataset is built for RLVR: reinforcement learning from verifiable rewards. During training, a policy model generates a next action, and the terminus_judge resources server in NeMo Gym scores it against the reference action. That's a reward signal you can compute at scale, no human in the loop. NVIDIA used the same 630-task ATCB set for post-training both Nemotron Ultra and Nemotron 3.5 Lightning. A dataset like this isn't a side artifact. It's the training pipeline for the next generation of terminal agents.

Key numbers:

  • 31,111 decision-point samples in NVIDIA's terminal RL dataset, from 630 ATCB tasks
  • 20/30 vs 16/30 tasks passed by Pi vs Claude Code with the same model
  • $0.028 vs $0.195 cost per successful task, a 7x gap
  • 56.92 → 60.32 Biology-Instructions score from a 4B Memory Decoder with the 397B backbone frozen

Scientific agents push the same stack further ​

The same recipe scales to science. Intern-S2-Preview, the scientific agentic foundation model series from Shanghai AI Lab, applies the full stack: multimodal pretraining over rendered scientific documents, then a unified post-training pipeline of supervised fine-tuning, scalable multi-task RL, black- and white-box agentic RL, and on-policy distillation.

Partial rollout with off-policy correction cuts the cost of generating rollouts. Adaptive length regularization keeps long-horizon tasks from collapsing into short answers. Online speculative decoding speeds up rollout generation. Trace-aware experience assembly feeds the RL update coherent decision sequences instead of scattered turns.

TechniqueWhat it does
Partial rollout with off-policy correctionCuts rollout generation cost without biasing the policy update
Adaptive length regularizationKeeps long-horizon tasks from collapsing into short answers
Online speculative decodingSpeeds up rollout generation during RL training
Trace-aware experience assemblyFeeds RL coherent decision sequences from agent trajectories
Memory Decoder (4B)Specializes the frozen 397B backbone: Biology-Instructions 56.92 → 60.32

The 397B model extends time series modeling from long-sequence understanding to numerical forecasting. And the Memory Decoder is a separate 4B-parameter path for rapid scientific specialization. It improves the Biology-Instructions average from 56.92 to 60.32 without modifying the frozen 397B backbone. A 4B model bolted on the side delivers a +3.4 point gain on a specialized benchmark while the big model stays untouched. That's the modularity thesis applied to scientific agents.

The harness ecosystem is already here ​

If the harness is the product, the plugin ecosystem around it is the moat. That's where things are moving fast.

dot-skill, formerly colleague.skill, distills a person into an AI Skill: a colleague, a partner, a public figure, even a fictional character. It builds a two-layer structure: a Persona layer that captures how they think and speak, and a Work Skill layer that captures their technical standards and workflows. The project runs across five hosts, including DeepSeek Harness through native filesystem skill discovery. The community gallery has 215 skills from 165 contributors.

OpenBiliClaw is a different kind of agent: a local-first, self-evolving content discovery system that builds a psychological profile from your cross-platform behavior and searches Bilibili, Xiaohongshu, Douyin, YouTube, X, Zhihu, Reddit, and more for content you'd like but haven't found. All data stays in local SQLite. It ships a DeepSeek Harness plugin that registers 22 Agent Bridge tools, so agents inside DSH can read recommendations and learn from feedback. The README even tells you to paste an install instruction into Claude Code or Codex and let the agent deploy itself.

Google's 5-Day AI Agents Intensive Vibe Coding course drew 353,000 registered participants, with 392,000 active on the Kaggle Discord and 6,000+ capstone projects. The scale tells you where developer attention is going: not just using agents, but building them.

What the community is saying: I've read a lot of takes on the Pi vs Claude Code split, and the pattern is consistent. People who build serious tooling on top of agents tend to praise Pi. OpenClaw was built on it, and prominent framework authors describe it as the agent they use almost exclusively. The minimalism argument lands hard, and the sharpest framing I've seen calls Claude Code the starter pack and Pi the endgame. But the daily-usage reality is the opposite. Most of the same people still run Claude Code for everyday work, because the guardrails are real and the subscription pricing is predictable. I found the OAuth block that cut third-party harnesses off from Claude subscriptions genuinely hostile when it landed, and it quietly pushed the economics back toward the subscription anyway. And the sharpest critique of Pi's ecosystem lands where you'd expect: a lot of community extensions are themselves vibecoded, and since Pi extensions run inside the agent process with full system access, that's a supply-chain concern. The docs tell you to review before installing. Most people won't.

Common Pitfalls ​

What trips people up with agent harnesses:

  1. Writing prompts like scripts instead of goals. The 36kr test is the cleanest example: 1,817 words plus five hard requirements produced a broken result, and a one-sentence goal worked on the first try. For frontier models, a detailed prompt can constrain the model into failure. State the goal. Let the model handle the how.

  2. Benchmarking harnesses with different models. The Pi vs Claude Code eval only means something because both ran DeepSeek V4 Flash. Compare harness A with model X against harness B with model Y and you're measuring the model, not the harness. Same model, or the comparison is noise.

  3. Ignoring token overhead. Claude Code burned 741,659 tokens per task versus Pi's 558,885. In long-context agentic work, token overhead is the hidden cost driver, and cache hit rate swings the bill by double digits. Measure tokens per task, not just pass rate.

  4. Treating RL datasets like SFT datasets. The Nemotron dataset is built for RLVR: each sample is a decision point, and the judge scores your policy's next action against the reference. If