Appearance
The Long-Horizon Wall: Inside the Agent Stack's Next Phase
The long-horizon wall
Give an LLM agent a realistic multi-step task. Split a bill. Find a song. Reconcile an order across nine simulated apps. When it fails, it usually isn't for lack of knowledge. It mis-paginates an API. It resolves the wrong person. It returns a value when none was asked for. The model knows the APIs. What it hasn't internalized is how to use them reliably over a long horizon.
That's the shift the industry is stuck on. Chatbots answer. Agents execute. Execution means holding a goal across dozens of steps, recovering from failures, and doing it cheaply enough to be worth running at scale. Every source in this cluster attacks a different layer of that problem, and taken together they tell a clear story: the bottleneck has moved from the model to the scaffolding around it.
| Layer | Problem it solves | Recent work |
|---|---|---|
| Harness design | The agent loop itself is static and hand-written | AutoDesign |
| Communication | Text tokens lose information between agents | StateBridge |
| Credit assignment | One reward per trajectory can't train multi-turn behavior | CrEST |
| Memory | Agents repeat the same mistakes across tasks | ALTK-Evolve, ACE |
| Tool interface | Software isn't agent-readable | CLI-Anything, MCP |
| Orchestration | Cross-cloud identity and coordination | A2A federation |
The through-line: nobody is claiming a new model fixes this. The wins are coming from the loop, the memory, the transport, and the boundaries.
The harness is the model
The strongest evidence for that claim comes from AutoDesign, a framework that treats the harness as the thing to optimize. The test bed is paper-to-poster generation: take an academic paper, produce a conference-quality poster. That's a genuinely long-horizon agentic task. It's multimodal, it involves dozens of tool calls, and it has a clear quality signal in human preference. It's also a task people actually pay for, which keeps the benchmark honest.
AutoDesign's loop is simple in retrospect. A meta-harness optimizer watches rollouts, reads the feedback, and writes improved harness code. A code agent implements the changes. The new harness runs again. The loop repeats.
The results are hard to wave away. On PosterBench, a new benchmark with a 100-paper main track across five disciplines, AutoDesign scores 78.32, beating Claude Design, a closed commercial system, by 7.45 points. On a 100-point scale measuring whether posters reach conference quality, that's the difference between work that reads as clearly AI-generated and work reviewers can't distinguish from human output.
The transfer results matter more than the headline score. Across seven controlled code-agent-model configurations, adding the learned DesignHarness improved the average PosterBench score from 54.99 to 67.39, a +12.4% gain. That's the part that makes this more than a demo: the harness isn't overfit to one model. It generalizes across configurations.
78.32 on PosterBench: AutoDesign beats Claude Design by 7.45 points.
+12.4%: average gain from the learned harness across seven model configurations (54.99 → 67.39).
<$3: cost of a fully autonomous run, 253 tool calls and 11 editing turns in 40 minutes.
22 / 26: model-task pairs where StateBridge wins or ties the strongest baseline, with zero training.
6.7x: token reduction per task from calibrated memory retrieval on gpt-oss-120b (777K → 116K).
The cost number is the one I'd underline. A fully autonomous run executes 253 tool calls and 11 editing turns in 40 minutes, for under $3. That's what makes recursive self-improvement practical. The enabler is a feedback loop you can afford to run hundreds of times. When a single iteration costs less than a sandwich, you can let the harness evolve all day.
A system-blind human study backs the scores up: AutoDesign gets the highest human preference among all evaluated systems. The implication is uncomfortable for anyone who assumed the model was the ceiling. If the harness is where the intelligence lives, then the biggest performance you can buy might be a better loop, not a better checkpoint.
Talking without tokens
Multi-agent systems have a communication problem that nobody talks about enough. They talk in text. Text is a discrete bottleneck. When a sender's continuous hidden states get converted to tokens, information that token identities can't capture is gone, permanently, before the next agent ever sees it.
StateBridge attacks this directly. Instead of text, agents transmit hidden representations. The clever part is that it's training-free. Previous latent communication methods either injected working memory layer by layer across the transformer, which is expensive, or required trained projectors, which kill portability. StateBridge aligns the sender's final-layer hidden states to the receiver's input space with a closed-form orthogonal transformation. That's a linear algebra operation, not a training run. The orthogonal part matters: it preserves the geometry of the hidden space, so the sender's representational structure survives the alignment. Lightweight norm calibration and vocabulary anchoring keep the aligned states compatible with the pretrained input distribution, and the result gets prepended to the receiver's input as a continuous prefix.
The evaluations cover math reasoning, code generation, and question answering across four models from two families. StateBridge achieves the best or tied-best score on 22 of 26 model-task pairs, consistently beating the strongest baseline.
22 of 26 means it wins or ties on 85% of configurations, with zero training cost. That's the part I find compelling. Latent communication has been a research curiosity because it required fine-tuning. A closed-form alignment that works across model families changes the economics. You can bolt it onto existing agents without GPU hours or data collection.
This is early, and I'd say that plainly. The evaluations are reasoning tasks, not long-horizon agentic loops. Nobody has shown latent communication holding up across a 40-minute, 253-tool-call trajectory yet. But the direction matters. If agents can hand off continuous representations instead of compressed text, the information loss at every hop drops, and multi-agent systems stop being a game of telephone played with tokens.
Who gets credit?
Training agents is a credit assignment problem, and the current tooling handles it badly. Reinforcement learning with verifiable rewards, RLVR, gives you a verifier-bounded ceiling: the reward signal is only as good as the verifier that checks the outcome. But trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward. You ran a 20-step trajectory, one turn was wrong, and the whole thing gets one number. The model can't tell which turn to fix.
Distillation gives you dense per-token supervision instead, but it's teacher-bounded, and it's prone to gradient concentration collapse. That's what happens when the distillation signal piles onto a few tokens and the rest of the trajectory gets no learning signal at all. You trade one ceiling for another.
CrEST, a hierarchical credit assignment framework, tries to split the difference. It keeps RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher. The resolution happens at two levels: turn-segmented verified advantages handle inter-turn dilution, and entropy-gated self-teacher modulation refines intra-turn token contributions.
The paper's framing is the part that matters most: the teacher's role in policy optimization can be reduced from determining update directions to modulating update magnitudes. That's a real division of labor. The verifier decides whether the outcome was right. The teacher decides how much each token contributed. Direction comes from the verifier. Magnitude comes from the teacher.
On BFCL V3 and WildToolBench, CrEST consistently outperforms both RL and distillation baselines across two model scales, with the largest gains on long-trajectory and strict session-level metrics. No single headline number, but the direction is uniform. And that last part is the tell: the harder the credit assignment problem, the more the hierarchical approach pays off. If you're training a two-step agent, this doesn't matter. If you're training a 20-step tool user, it's the difference between a model that learns "the whole trajectory was wrong" and one that learns "turn 4 was wrong, and within turn 4, the tool call was the problem."
Quick Take: The pattern across this wave of work is consistent. The bottleneck has moved from the model to the scaffolding. The biggest measured wins, a 12.4% harness gain, a 6.7x token reduction, come from fixing the loop around the model, not swapping the model.
Memory: count it, don't collapse it
Two recent systems, ACE (Agentic Context Engineering) and ALTK-Evolve, solve the same problem from opposite directions: turning an agent's past trajectories into reusable lessons, fed back at inference time, with no weight updates and no human labels. They agree on the hard part, and they disagree on delivery. The disagreement shows up in the token bill.
Both refuse to compress. ACE names the failure modes precisely: brevity bias, where optimization collapses toward short generic instructions, and context collapse, where a model asked to rewrite its whole context each step summarizes the detail away. Its answer is a rich itemized playbook with a helpful/harmful counter on every bullet. ALTK-Evolve reaches the same conclusion from the other direction, keeping a support count on every guideline, the number of independent episodes that produced it. A lesson five different tasks discovered is a different object from one that appeared once. Both systems land on the same principle: count the lessons, don't collapse them.
The difference is delivery. ACE injects the comprehensive playbook on every step, the same way regardless of model or task. ALTK-Evolve treats delivery as a dial: a small fixed core of high-support guidelines, extended per task with a handful selected for the task at hand, or the full consolidated set when the model has the headroom to use it.
On AppWorld, with the same base ReAct agent, the numbers are stark:
| Model | System | TGC | SGC | Tokens/task |
|---|---|---|---|---|
| DeepSeek-V3.2 | ACE | 80.4 | 73.2 | 634K |
| DeepSeek-V3.2 | ALTK-Evolve | 89.3 | 80.4 | 263K |
| gpt-oss-120b | ACE | 54.8 | 35.7 | 777K |
| gpt-oss-120b | ALTK-Evolve | 56.0 | 37.5 | 116K |
On the strong model, ALTK-Evolve wins both metrics at roughly 40% of ACE's inference cost. On the weak model, the accuracy is a statistical tie, 56.0 to 54.8, at about one-seventh the cost. The same lessons, delivered differently.
The by-difficulty breakdown explains where the accuracy comes from, and it's not where you'd guess. On gpt-oss-120b, ACE's full playbook wins Easy and Medium tasks. There's enough of the task solved by generic instruction-following that a comprehensive prompt helps more than it distracts. But on Hard tasks, where the model has to pick the right lesson rather than wade through all of them, curated retrieval pulls ahead, 31.8 vs 23.8 TGC. Hard tasks decide the aggregate.
On DeepSeek-V3.2 the story flips: the stronger model absorbs the full playbook well enough to edge out ALTK-Evolve on Medium, but ALTK-Evolve wins Easy, Hard, and Overall. More capacity to spare means more lessons, delivered selectively, keep helping instead of crowding each other out.
The practical lesson is blunt: a large context overwhelms a weaker model rather than helping it. If you're building agent memory into a system, the content of the memory matters, but the delivery mechanism matters just as much. Injecting everything, every step, is the expensive default. It's not the right default.
The tool boundary
Agents need to touch software, and most software wasn't built for them. Two projects are attacking this from complementary angles.
CLI-Anything's bet is that the command line is the universal interface for both humans and agents. It's structured, composable, self-describing through --help flags, and deterministic when you ask for JSON output. The project generates agent-native CLIs for arbitrary software through a seven-phase pipeline, and it maintains CLI-Hub, a registry where agents can discover and install harnesses autonomously. The catalog is already absurd in the best way: FreeCAD with 258 commands across 17 groups, GIMP, Blender, Obsidian, QGIS, Godot, MuseScore, even Slay the Spire II. Each generated CLI ships with a SKILL.md, an agent-discoverable skill definition.
The security hardening in the changelog is the part I'd point at. Path traversal fixes, symlink escape guards, untrusted XML routed through defusedxml. This is the unglamorous work that determines whether agent tooling is safe to run at all. An agent will happily call a compromised tool. The boundary has to be the defense.
The complementary approach is MCP, and a recent Dart walkthrough shows the pattern in miniature. A Dart agent built on the unofficial adk_dart library exposes a greeting tool over Server-Sent Events, with a Shelf server implementing the handful of MCP methods a demo needs: initialize, tools/list, tools/call. The author is explicit that it's a narrow demo, one agent, one deterministic tool, enough MCP to show the request flow.
What the community is saying: the comments on that post are where the real lessons surfaced. I found myself agreeing with the security critique immediately. When I tested this pattern with a local MCP server, the first thing that bit me wasn't the agent logic, it was the transport. The sessionId in the URL is effectively a bearer capability, and URLs get copied into access logs and traces. Randomness alone shouldn't authorize a POST to an existing SSE channel. The fix is to bind sessions to authenticated principals, apply short idle and absolute TTLs, delete on disconnect, and keep the query string out of logs.
The framing that stuck with me was the three-layer view of agent failures. Mechanical: did the tool execute? Evidence: did the receipt match what was asked? Semantic: did the model draw the right conclusion? The dangerous failures happen between the second and third layers. The tool can succeed, the receipt can be valid, and the agent still draws the wrong conclusion. MCP is valuable because it gives you a clean place to keep the first two layers deterministic, leaving the model to handle only the part that actually requires reasoning.
One commenter extended that in a way I wish more tooling did: treat the tool schema as a first-class contract, not