Skip to content

The Agent Stack Is Growing Up: Cost, Memory, and Tooling

#llm-agents #agent-memory #inference-cost #coding-agents #agent-efficiency #agent-tooling

The Agent Stack Is Growing Up: Cost, Memory, and Tooling ​

The problem: the demo works, the bill doesn't ​

Say you need to classify a million product listings. A frontier LLM can do it. It'll get most of them right. And it'll cost you a fortune, because you're paying the same per-token price for query number one and query number one million.

That's the gap the agent community is attacking right now. A cluster of October 2026 research and product work shows the same conviction: agents need to be cheap at scale, honest about memory, and fit into real workflows. The demo question is settled. The production questions are not.

The BOTTLED benchmark paper names the first problem directly. LLMs solve narrow tasks well, but querying them separately for millions of related instances is prohibitively expensive. Can an agent turn its general capability into something cheap and reusable? Can it bottle the capability?

Bottling: capabilities into artifacts ​

BOTTLED gives agents an entire unlabelled workload and asks them to finish it under fixed time, compute, and LLM API budgets. No instructions about how. The agent can train a small model. Write a reusable program. Whatever works. Ten models, three tasks, sixty runs.

The results should worry anyone shipping agent demos straight to production. Strong zero-shot performance does not reliably transfer to bottling. Models with nearly identical zero-shot scores diverged substantially once they had to build something reusable. 48 of 60 bottling runs scored below the lower bound of their own model's zero-shot 95% confidence interval. 31 of 60 underperformed the stronger of two small-model distillation baselines given the same token budget.

But when bottling works, it works absurdly well. On query-product relevance classification, Opus 5 kept about 82% of its zero-shot macro-F1 at roughly 657 times lower reported cost. And against Jev, a model built specifically for cheap repetitive inference, bottled Opus 5 recovered about 94% of Jev's macro-F1 at a quarter of Jev's projected full-workload cost.

657x cheaper. Bottled Opus 5 keeps 82% of its zero-shot macro-F1 on query-product relevance. 48 of 60 bottling runs scored below their model's zero-shot baseline. 31 of 60 runs underperformed the stronger small-model distillation baseline on the same token budget. 94% of Jev's macro-F1 recovered by bottled Opus 5, at a quarter of Jev's projected cost.

The practical read: for repetitive workloads, the agent's job is to build the thing that answers. That's a different skill, and the benchmark suggests it's underdeveloped in most frontier models. If you're running a classification pipeline, the cost curve is the product.

Memory: forget summaries, manage state ​

Cost is one wall. Memory is the other. Long conversational agents must remember what was said months ago, but instructions and context change. A good memory system should preserve both current and historical states, tell stale information from active knowledge, and retrieve evidence appropriate to the query. And it shouldn't keep invoking an LLM to rewrite prior interactions.

MINDSET, a memory controller out this month, treats this as a state-management problem. Conversations arrive as immutable episodes. The controller organizes them into versioned schemas by making minimum-energy state transitions. Each new episode can reinforce a schema, supersede it, split it, or create a new one. The transition decision weighs representation distortion, contradiction, historical damage, fragmentation, and internal inconsistency. Hysteresis stops one isolated contradiction from rewriting stable memory prematurely.

The results are strong. Tested against five memory systems on 850 questions (700 LoCoMo, 150 MemoryAgentBench), MINDSET posted the highest LoCoMo answer F1 and significantly better retrieval ranking: Recall@8, MRR, and nDCG@8 all beat the second-best system, LightMem, at p<0.01 after Holm correction. Ablations point to controlled fragmentation and schema-aware assignment as the largest contributors. A 700-question cross-model run with GLM-4.7 and Gemma-4-31B showed it isn't tied to one model family.

The finding I keep coming back to: memory as constrained state management beats continual summarization. Summaries lose the past. Versioned schemas keep it, and let retrieval walk back to evidence instead of a paraphrase.

Quick Take: Long-term agent memory behaves more like a versioned database than a summary log, and the benchmark numbers agree.

Memory's second front: parameters ​

There's a complementary line of work on where memory should live at all. A survey from the same week maps the in-parameter memory space: instead of stuffing everything into the context window, reusable memory gets represented in model parameters, adapters, or parameter-like objects composed into the forward pass at inference time.

The survey organizes the field with two axes: Parameter Placement (embedding, attention, FFN layers, or hybrid) and Parameter Acquisition Time (online during deployment, offline before it). The motivation is the cost problem again. In-context learning is flexible, but it eats context capacity and pays repeated encoding costs that grow with context length. Parametric memory pays once.

How the main memory approaches stack up:

ApproachWhere memory livesUpdate mechanismPractical catch
In-context (ICL)Context windowRe-inject each turnToken cost grows with context; repeated encoding on every call
In-parameter (adapters, LoRA)Model parametersGradient-based updates, online or offlineInterference between memories; needs training compute at deployment
External vector storeDatabase of embeddingsEmbed and indexRetrieval quality depends on chunking and embedding choice
Versioned schema (MINDSET)Episode store + schemasMinimum-energy transitionsController complexity; hysteresis tuning

Open problems remain: interference between old and new parametric memories, safety of memory-bearing weights, co-design with ICL, and recursive self-improvement where the agent updates its own memory objects. All of these are research problems now. All of them are production problems soon.

Agents enter real workflows ​

The research is chasing efficiency. Products are chasing the same thing from the other side. Rabbit's OS3, opened up last month, is a personal agent operating system with a pointed design stance: people shouldn't reorganize their thinking for the agent. One continuous conversation window, no project folders, no work/life tabs. The founder calls it "brain dump": you say whatever's in your head, and the system figures out which tool and which computer handle it.

OS3 connects up to five of your existing computers and executes locally on them, using their files and logged-in sessions. That's a real engineering answer to a real problem: a cloud agent that logs into Amazon from a Minnesota datacenter looks like an intruder to the website. Local execution also becomes a privacy boundary. If the file stays on your machine, it doesn't get uploaded. Users bring their own model API keys, which is both a cost-control bet and a hedge: no model vendor wins outright, so any model improvement becomes Rabbit's improvement.

The honest part of the interview is that demand isn't validated yet. Trying Meta's Muse, the founder asked it to buy a case of Coke, then watched it wrestle with a forgotten password and a password-manager detour. The gap between "agent can operate software" and "agent is more convenient than two clicks" is still wide. The brain-dump design only works if memory is good enough to hold a page of half-formed thoughts and retrieve the right context months later.

What the community is saying: When I compared agent shells back to back, the token meter was the first thing I checked. One user reported OS3 consuming noticeably more tokens than OpenClaw on the same model, and the Rabbit team acknowledged it and promised an optimization pass. The buying-a-coke test keeps coming up as the minimum bar: when an agent needs three steps and a detour to do what a human does in two clicks, the demo stops feeling magical. The sentiment I keep hearing from people building on these tools: capability is no longer the complaint. Cost and convenience are.

OpenAI's product moves this month push the same direction from enterprise and education angles: computer use with Ironclad for contracting workflows, the Atlassian partnership connecting frontier models to enterprise knowledge, and College Planner for ChatGPT for Teens. Some of these are evaluations, some are pilots. All of them are agents doing work with consequences. Contracting workflows don't forgive a hallucinated clause.

Coding-agent tooling is exploding ​

The most concrete evidence of agent maturity is the tooling ecosystem. This week's GitHub trending lists show two very different examples of the same pattern: agents are getting hands.

text-to-cad gives agents local workflows for generating 3D models as STEP, GLB, STL, or 3MF files, with manufacturing checks, engineering drawings, and hooks into fabrication services. It installs as a plugin or skills set across Claude Code, Codex, Cursor, Gemini, and Grok. The skills list is a small factory: CAD modeling, off-the-shelf part lookup, DXF drawing, URDF/SRDF/SDF robot descriptions, DFM analysis with cited rules, G-code slicing, and direct-to-printer jobs via Bambu Labs. The MCP server runs locally through uv, and models render as interactive viewer cards in the conversation.

The install handling is unusually careful. Pinned versions, update buttons, rollback commands, anonymous usage telemetry that's off until you opt in. Someone who has watched non-technical users fight an agent toolchain wrote this README.

I ran into the Windows case myself. The underlying CAD kernel ships an unsigned native module, and Windows 11's Smart App Control blocks it by default. Every cadgen command fails with an OCP import error, logged as Event ID 3077. There's no per-app exception; you either disable Smart App Control (one-way, requires reinstalling Windows to re-enable) or run under WSL. That's the kind of edge case that never shows up in a demo and always shows up in a ticket.

gpt-instruct shows the verification side of coding agents. It's a prompt and evaluation toolchain for Codex, focused on first-round execution, process continuity, artifact verification, and runnable rollback. The project runs a three-layer gate system before any prompt ships: an A gate on original samples with fresh-run probes, a 66-case issue regression suite, and a 120-case medium-difficulty test set that only runs after A and B both pass.

The current release candidates tell the story honestly:

VersionStatusA gate (fresh runs)B gate, non-cloudB gate, cloud repeatsArtifact gates
gpt-5.6-sol-v45Stable productionn/an/an/an/a
gpt-6-astra-v2-rc1Prerelease3/4, 3/442/50 cases, 48/56 turns23/48 attempts16/16
gpt-6.1-sol-v1-rc2Prerelease3/4, 3/434/50 cases, 40/56 turns22/48 attempts15/16

Both candidates pass the A gate twice at 3/4. Both miss the B gate's 66/66 hard bar. Cloud runs are reported separately from non-cloud runs, because reruns drift. Results stay scoped to identical model, reasoning level, and method identity. The author also runs JailbreakBench suites and refuses to let a fiction benchmark substitute for the hard safety gates. That's the rigor agents need: the agent produced the artifact, but someone has to verify the artifact, and today that's bespoke tooling.

Common pitfalls ​

A few mistakes keep showing up across these projects.

Assuming zero-shot scores predict production performance. The BOTTLED result is unambiguous: models with similar zero-shot scores diverged wildly on bottling. If you evaluate an agent only on single-call accuracy, you'll be surprised by how it behaves when it has to build something reusable.

Treating memory as summarization. Roll the conversation into a summary every N turns and you lose the evidence. MINDSET's wins come from keeping immutable episodes and versioned schemas, with retrieval that walks back to actual interactions. Stale summaries silently degrade retrieval, and you can't roll back a summary.

Ignoring amortized cost for repetitive workloads. If you're doing millions of calls on the same task shape, paying frontier prices per call is the expensive option by roughly three orders of magnitude. Bottling, distillation, or a small fine-tune is the difference between a service that survives pricing pressure and one that doesn't.

Shipping plugins without testing the install path. Unsigned native modules break on Windows 11's Smart App Control with confusing import errors. Agent tooling that expects terminal familiarity fails for the exact users consumer agents are supposed to reach. That's the same lesson Rabbit learned when its most common support request was "how do I install OpenClaw."

Merging cloud and non-cloud eval results. gpt-instruct's cloud reruns consistently scored below non-cloud runs on the same test set. If your eval harness averages the two together, you're folding in a source of variance you don't control.

One thing to remember ​

The agent stack in late 2026 is no longer bottlenecked on raw capability. The bottlenecks are amortized cost, memory that stays honest, and tooling that survives real machines. Every source in this cluster, from the BOTTLED benchmark to a CAD plugin README, is an attempt to clear one of those three hurdles.

The bottom line ​

If you're building agents for large repetitive workloads, adopt bottling. Give the agent a budget, let it produce a reusable artifact, and evaluate the artifact's quality. A 657x cost reduction at 82% quality retention beats any prompt optimization you'll find.

If you're building long-running conversational agents, structure memory as versioned state, not summaries. Immutable episodes plus schema transitions give you recall that summaries can't, and they give you rollback when a memory turns out wrong.

If you're shipping agents to regular users, watch the BYOK plus local-execution pattern. Token-cost transparency and cross-device execution are becoming default expectations, and products that hide API costs behind a subscription will get squeezed when users see the meter.