Appearance
The agent stack is growing in two directions at once
Pick any agent framework this month and you'll see the same growth pattern. Multi-agent orchestrators spawn specialist subagents. Runtimes add permission systems, optional bundles, and async interaction modes. The tracking repos now list 1,000+ Claude skills, curated Codex plugin marketplaces with scanner gates, and MCP gateways counting integrations in the thousands.
A quieter set of results asks whether any of that machinery earns its keep. One paper finds that elaborate harnesses give no advantage over a single session of a minimal coding agent under the same backbone and time budget. Another finds that counting the rows a tool returns, a step every agent performs and almost nobody tests, silently fails in most frontier models. A third shows that a claim can be true somewhere in the evidence but attributed to the wrong source, and the verification layer that would catch it is only now being built.
Three threads tie this cluster together: how much scaffolding agents actually need, where tool use breaks in practice, and how we judge, verify, and audit the results.
The harness question
"How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?" runs a direct test. Same frontier backbone, equal time budget, MLE (machine learning engineering) benchmark tasks. The elaborate setups include multi-agent orchestrators and dedicated retrieval subagents. The baseline is a single session of a minimal coding agent with read, write, and bash primitives.
The result is blunt. The harnesses provide no advantage over the minimal baseline. On current MLE benchmarks, the backbone drives performance, and the machinery layers become redundant in the coding agent setting. Hand-crafted harnesses around strong models are a poor return on effort.
I've built my share of orchestrators, and this matches experience more than I'd like. The layers I added solved problems the underlying model had already solved, just slower. The caveat is scope. These are structured, long-horizon coding benchmarks, and open-ended research tasks may still reward scaffolding. But the burden of proof has shifted. If you add a retrieval subagent, you should be able to point at a run where the same backbone without it loses.
Free training grounds live in fictional worlds
The harness paper cuts one kind of complexity. PhantomEnvironments cuts another, the data bottleneck in RL for agents. RL training usually needs environments that give verifiable rewards, support long-horizon interaction, and scale cheaply. Human-curated data is expensive. LLM-generated environments risk hallucinations and benchmark contamination.
PhantomEnvironments builds multi-turn RL environments entirely from rules, in fictional worlds. Agents search a corpus of templated articles to answer multi-hop questions. Generation needs no LLM and costs nothing per episode.
The surprise is that it transfers. Agents trained on worlds that share no facts with reality do well on real-world multi-hop search benchmarks, often beating real-world training data on newer benchmarks. The ablation explains why: hop count drives transfer more than constraints or comparisons. The agent learns structural difficulty, not surface facts. Qwen models even learn to scale their search budget roughly linearly with question difficulty, an emergent search-scaling behavior that comes from environment interaction alone.
Teams building agents should read this as a free resource. The instinct to hand-curate realistic training data may be misplaced. A rule-generated world with the right hop structure teaches search better than a pile of real queries.
Quick Take: The consistent pattern across this week's results is that small, measurable loops beat elaborate machinery, and the verification stack is where the real gaps are.
When agents do use tools, retrieval has to co-evolve
EvoDuet starts from a specific failure: evolutionary search with LLMs stalls when progress needs external knowledge the model lacks. Adding a web search tool helps, then stops helping, because the same pages keep getting returned as the solutions change. The fix is to treat retrieval as part of the optimization loop, not a bolt-on.
At each iteration, a retrieval gate lets the LLM assess its own knowledge gap and choose one of three paths: fetch new documents, reuse stored ones, or proceed without retrieval. An inner loop refines the queries and ranks documents by the solution scores they're predicted to yield. An outer loop generates candidates in parallel from those documents and records evaluated outcomes for later searches.
The numbers, across 21 optimization tasks at one candidate per iteration: EvoDuet lifts OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash. The best runs beat previously reported scores on eight tasks and match on three more. The method also improves other scaffoldings, Top-K and EvoX, on Sums/Diffs and Denoising.
One result deserves attention: Qwen3.5-9B does not benefit. At 9B parameters, small enough to run on a single GPU, the model can't use the extra documents well enough to justify the noise they add. Retrieval co-evolution is a capability of the backbone, not a property of the loop. If your model is small, more retrieval context can hurt.
EviRover makes the same point in a different modality. Visual perception is normally a one-shot prediction from a single glance. EviRover reformulates perception as an agentic process that gathers evidence beyond the first look. It's a 4B model, small enough to run locally, trained on 5K supervised examples and 12K RL episodes, and evaluated on EviLens, a human-verified benchmark of 688 instances across five categories. It beats its own backbone by 30 points on average, which means the agentic loop, not scale, is doing the work. It reaches parity with advanced proprietary models, and the gains carry over to WebEyes and BrowseComp-VL, which improves by 15 points.
Both papers reframe the same operation. A one-shot step, a query or a glance, becomes an interactive loop with state. That's the direction of travel for tool use.
Key numbers from this cluster: EvoDuet lifts discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash across 21 tasks. PhantomEnvironments generates training worlds by rules, with zero marginal cost and no LLM in the loop. The counting benchmark separates models at 330 ids: over 2,700 output tokens per question means 15 to 21 of 21 lists correct, under 600 tokens means 0 to 10. ProvenanceGuard caught 138 of the 139 claims that human experts flagged, at roughly half a second per answer.
The counting problem: where tool use breaks
Now the failure mode that everyone hits and nobody tests. An agent's search, database, or API tool returns a list of records. The user asks how many. The model does the counting.
The Count It or Compute It benchmark asks ten models the same 68 counting questions with two versions of one tool. count_ids returns the exact count. list_ids returns the matching ids. Everything else is fixed, so every miss can be traced to either the query or the arithmetic.
With the count in the tool, every model but one answers correctly at a flat cost of 35 to 451 output tokens per question. With the ids, the models split in two. The ones that spend 2,700 to 8,200 tokens per question count 15 to 21 of 21 lists correctly. The ones that answer in under 600 tokens count 0 to 10. No model hedges. Each wrong answer is a single confident number.
| Model | Correct with rows tool (of 21) | Output tokens per question | Correct with count tool (of 21) |
|---|---|---|---|
| Gemini 3.8 Flash | 21 | 2,742 | 21 |
| Gemma 4 26B A4B | 20 | 8,153 | 21 |
| Gemini 3.7 Flash | 20 | 2,785 | 21 |
| gpt-oss-20b | 15 | 2,684 | 19 |
| Claude Sonnet 5 | 10 | 401 | 21 |
| Claude Opus 5 | 9 | 579 | 21 |
| Claude Haiku 4.5 | 5 | 143 | 21 |
| GPT-5.4 mini | 5 | 39 | 21 |
| Gemini 2.5 Flash | 0 | 580 | 21 |
| GPT-5.4 nano | 0 | 35 | 21 |
Two findings stand out. When I first looked at these numbers, I assumed model size or price would predict the split. Neither did. The query was right every time, so every miss on the rows tool was the model counting a correct list wrong. The failure sits in the arithmetic, and it's invisible unless you separate it from retrieval. Claude Opus 5, the most expensive model in the lineup, spent 579 tokens and counted 9 of 21. Gemma 4 26B, an open-weight model you can self-host, spent 8,153 tokens per question and counted 20. At 35 tokens for GPT-5.4 nano, there is no room to count anything; 35 tokens is less than a sentence.
Output caps are the hidden knob. In an earlier run with 1,100 ids in the prompt, capping output at 8,192 tokens dropped Gemini 3.7 Flash from 19 and 15 of 21 lists correct to 1 of 21. It still produced a confident number every time.
The fix is boring and cheap: return the count, the minimum, and the maximum in the tool. Every model in the lineup then answers correctly at a fraction of the tokens. Tool shape is an accuracy setting, and most teams never look at it.
Judges: treat disagreement as a signal
JuryFlow attacks the other end of the pipeline, evaluating agent output. A single LLM judge is unreliable, and a panel of judges leaves a residue, because when judges disagree, majority voting discards the conflict instead of resolving it.
The design is precise. JuryFlow decomposes each candidate response into atomic claims, has a panel of heterogeneous judges assign per-claim verdicts, and builds a disagreement graph. Nodes are scored by verdict entropy. Edges encode structural similarity between claims. A human acts as a structural guide, picking which disagreement to resolve with one minimal intervention. The correction then propagates along graph edges and to historically similar cases, and resolved cases crystallize into rubric entries that all judges inherit, so the evaluator improves over time.
In the automatic configuration, where focal selection is done by entropy ranking, JuryFlow improves agreement with gold labels over single-judge and majority-vote baselines on MT-Bench and LLMBar. The resource shift is the point. Instead of relabeling responses, you resolve a handful of disagreements and let the graph do the rest. The human role changes from labeler to structural guide.
Verification with a source attached
ProvenanceGuard targets a failure mode that source-blind verifiers miss: cross-source conflation. A claim that is true somewhere in the evidence, but attributed to the wrong source. A customer support agent says "according to the account record, this plan includes a 30-day refund window." The refund is in the policy document. Pool the sources and the claim looks supported. Keep them separate and the attribution is wrong, which in a clinical or financial setting can be as damaging as a wrong fact.
ProvenanceGuard runs as a post-generation gate on top of a black-box MCP agent. It preserves source identity through the whole pipeline: break the answer into claims, route each to the most relevant source, check support, compare the source against the one the answer names or implies, then emit per-claim verdicts and a global allow or block decision. Literal values get strict treatment. A number, date, or identifier that isn't in the source can't pass because the sentence sounds plausible.
On 361 claims from 40 medical answers, human experts flagged 139 as should-not-pass. ProvenanceGuard caught 138 and let one through. In a controlled test where the named source was swapped in 50 cases, it caught all 50. Where similar sources compete, it blocks well but picks the exact source correctly about half the time, which the authors flag as the remaining weakness.
| Verifier | Reject/block F1 | Emits claim-to-source ID |
|---|---|---|
| ProvenanceGuard | 0.802 | Yes |
| MiniCheck | 0.783 | No |
| RAGAS Faithfulness | 0.758 | No |
| AlignScore | 0.662 | No |
| SummaC-ZS | 0.436 | No |
The other checkers score competitively on blocking but never tell you which tool output supported the claim. The claim-to-source connection is the product. Blocked answers go through a RARR-style repair loop, which resolved all 173 blocked cases in the main test, though 144 ended in fallback text rather than a rewrite. The system prefers an unverifiable refusal over a manufactured answer. Overhead is roughly half a second per answer on the local configuration.
This is already moving into production stacks. NVIDIA NVFlow merged an optional grounding-verification stage for its finance agent that checks answers against the SEC excerpts the agent retrieved. That's the adoption signal for source-aware verification.
Skills, plugins, and the audit layer
The tooling layer is standardizing fast. Claude Skills, the format Anthropic released as an open standard in December 2025, is now supported across Claude Code, Claude.ai, Codex, Cursor, Gemini CLI, Antigravity, and Windsurf. The key design choice is progressive loading: at session start the agent sees only each skill's name and description, roughly 100 tokens per skill, and the full body loads only when relevant. That's how a single agent hosts hundreds of skills without bloating its context window.
The layering matters. MCP defines how an agent connects to external systems: auth, transport, tool discovery. Tools are the functions an agent invokes. Skills define the workflow, the guardrails, the order of operations. In production all three run together. MCP for access, tools for actions, skills for behavior.
The curated lists show how quickly this is consolidating. awesome-claude-skills tracks 1,000+ production skills. awesome-codex-plugins runs a marketplace with a scanner gate: a centralized scan must report a numeric score of at least 80 out of 130 before a plugin is merged. awesome-llm-apps ships 100+ open-source agents that install with a single npx command. When I pointed Codex at the marketplace repo, the one gotcha was that the marketplace command clones a Git repository, so a raw marketplace.json URL fails with a remote 404. You add the repo URL, not the file. The gate is the more interesting part: plugin curation is becoming a scored, machine-checked process.
Runtimes are maturing along the same lines. DeepSeek harness 0.2 moved session-local reminders into an optional bundle, added a permission-diagnosis skill for Windows Sandbox that detects Access Denied causes and applies recoverable fixes after authorization, and made "ask the user" asynchronous so the agent keeps working until the answer arrives. The permission fix is the kind of change that only ships