Appearance
A $6,531 mistake
An AI agent tried to join DN42, a hobbyist network where people practice BGP, DNS, and other internet backbone tech. Its stated goal: "create an index of the network." Its method: five AWS m8g.12xlarge instances, each with 48 vCPUs, 192 GiB of RAM, and 22.5 Gbps of network bandwidth, running full-port scans every hour.
The bill was $6,531.30. The operator ended up begging for donations from the community the agent wanted to scan.
Key Numbers:
- $6,531.30 — the AWS bill that bankrupted the operator
- 5 × m8g.12xlarge — instances, each with 48 vCPUs and 192 GiB of RAM
- 22.5 Gbps — network bandwidth per instance, aimed at a network where most participants run 100 Mbps VPSes
- 100 Gbps — the aggregate scanning bandwidth the agent claimed it needed
The agent didn't just spin up one box. It wrote a justification for the infrastructure: "Scanning the entire DN42 prefix space at 20 Gbps requires multiple high-bandwidth interfaces and CPU cores to handle packet capture, filtering, and state tracking without dropping packets." The reasoning is coherent. It's also absurd in context. DN42 participants mostly peer over cheap VPSes with traffic measured in hundreds of gigabytes. Five 20 Gbps instances would have DoS'd whoever peered with them and burned through traffic quotas in minutes.
The IRC channel noticed immediately: "what's this dn42 they know about where everyone has enough bandwidth to easily spare 100G."
Two things went wrong. Nobody with authority reviewed the plan before money started burning. And the agent had tools and a goal, but no permission system between them. That pattern repeats across every incident in this article.
The permission gap
The DN42 agent had a binary choice: deploy instances or don't. Binary permissions are reachability, not authorization. The same tool is harmless in staging and dangerous in production. The same read is fine on public docs and risky on customer data.
The numbers back this up. About 18% of MCP server deployments implement any access scoping. 80% of orgs admit agents have taken actions beyond intended scope. OWASP classifies agent tool misuse as a first-class risk.
The fix emerging in the community is a four-state decision model. Agent ToolTrust, a permission engine I've been testing, returns one of four decisions instead of two:
| Decision | Meaning | When it fires |
|---|---|---|
| allow | Let it through | Low-risk, in-scope calls |
| audit | Allow but log everything | Reads on sensitive data |
| escalate | Stop and get a human | Destructive actions in production |
| deny | Block outright | Unknown tools, malformed input, policy violations |
The engine sits outside the model. The LLM proposes, policy disposes. No amount of prompt engineering can override a deny, because the deny happens in code, not in the prompt.
Fail-closed everywhere. Unknown tool → deny. Malformed input → deny. Engine crash → deny. The alternative is fail-open, which means an attacker who can crash the engine gets unrestricted tool access.
Testing against real agents, not mocks
I've been through the mock-agent trap myself. Unit tests pass. Demos look clean. Then real agents run and everything breaks. On one eval harness I built, the pass rate was 9% once real agents were involved. The integration, not the judge, was the problem.
The ToolTrust field test is the counterexample. 83 real agents across 10 frameworks. Zero mocks. 30 scenarios, including 10 adversarial ones: prompt injection, Unicode obfuscation, replay attempts, blank tool names, grant-bypass attempts.
The math: 83 agents × 30 scenarios = 2,490 runs. Each run calls a local 4B model and takes 30-80 seconds. That's about 12 days of compute. The covering design cut it to one afternoon. Plan A runs one scenario per agent, 83 runs, covering all 30 scenarios, all 10 frameworks, all 5 agent classes. Plan B proves each framework handles all four decision types, 123 runs. Together: 206 runs instead of 2,490. Same coverage. 12x reduction.
The 7 failures were the most valuable part. All 7 shared one pattern: not-available. The guard never fired because the LLM didn't call the tool. A mock agent always calls the tool. A real 4B model, given 5 tools at once, sometimes answers textually instead.
Never interpret not-available as a policy failure. It means the LLM didn't call the tool. That's different from unexpected-decision, where the guard ran and made the wrong call. Only the latter is a real regression. If your CI fails on not-available, the gate is flaky because of model nondeterminism. Fail only on unexpected-decision and the gate is strict but stable.
The framework quirks only surface with real agents. LangGraph's ToolNode isn't callable in v1.x, so you wrap the tool before it enters the graph. Google ADK's InMemorySessionService.create_session() is a coroutine you have to await. LlamaIndex's legacy ReActAgent has no .query() or .chat(), and execution is driven by async for over handler.stream_events(). A separate await handler yields nothing. I spent an hour on that one.
Quick Take: The expensive resource in agent testing is real-agent execution, not code coverage. Spend it on what only real agents can prove, and use covering designs to cut the matrix.
Memory is a write problem
The DN42 agent had no memory of consequences. The ToolTrust engine has audit logs but no long-term memory. The next layer of the stack is durable memory, and the community is converging on a hard lesson: vector databases aren't enough.
A vector database is one implementation of durable memory. It is not the architectural definition. Durable memory is a policy, not a place.
The failure state has a name: the Digital Attic. Everything is kept, nothing can be found. When an attic gets queried, it hands your agent a poisoned working set: current requirements, obsolete notes, conflicting observations. That un-sieved context hits the context window and the agent starts thrashing, spending inference cycles reconciling contradictory history instead of making progress.
Memory is a write problem. Teams debate embedding strategies, chunk sizes, hybrid search, and re-ranking. Almost nobody asks the opposite question: should this be remembered at all? Every write is a promise to your future retrieval system. A system that remembers everything eventually remembers nothing particularly well.
What the Community Is Saying: The write-boundary framing got the sharpest pushback in the comments. One engineer I respect put it this way: curation fixes volume, but it doesn't fix jurisdiction. A clean durable memory with no attic in it can still hand back two entries that both passed the write boundary honestly and still disagree. Nothing in the ranking knows which one governs. Another thread pushed on write-time authority versus read-time authority. I hit this testing against a live certificate authority. A stored fact stays valid by its own TTL while the source underneath revokes the grant. Nothing about the entry is stale. The authority it inherited is. You can cache a prediction. You can't cache news. Expiry is locally computable. Revocation only exists at the source.
Context as infrastructure
OpenViking is one attempt to build this properly. It treats agent context as a virtual filesystem under viking://. Memories, resources, and skills are files. Agents browse with ls, tree, and find instead of querying a black-box vector store.
Every entry gets processed into three tiers on write:
| Tier | Content | Token cost | Used for |
|---|---|---|---|
| L0 abstract | One-sentence summary | ~100 tokens | Quick relevance check |
| L1 overview | Structure and key points | ~2k tokens | Planning |
| L2 details | Full original data | Full | Only when needed |
Retrieval drills down: vector search locates the highest-scoring directory, then loads layers only as deep as the task requires. Each query preserves its browsing trajectory, so when a result looks wrong you can see which path produced it.
The benchmark numbers are the interesting part. On LoCoMo, a long-conversation user memory benchmark, three agent integrations with OpenViking land at 80-83% accuracy, up from 24-57% on native memory. That's more than double the accuracy. Input tokens drop by 34.3-91.0%, which for a long-running agent is the difference between fitting in a context window and paying for a bigger one. Query latency drops by 58-66%.
The session handover paper formalizes the same problem from the theory side. When a task continues in a new session, context limit hit, app restart, another agent takes over, what should you pass on? The paper distinguishes exact recovery of earlier material from preservation of the target distribution. Under an exogeneity condition, predictive equivalence characterizes the coarsest deterministic sufficient handover. In plain terms: you don't need to pass everything. You need to pass enough to keep the downstream task's distribution intact.
The practical output is a three-part record. Store decisions and constraints exactly. Use task-justified statistics for repeated evidence. Retain original observations whose effect isn't preserved by those statistics. That maps almost one-to-one onto the write-side custody questions from the memory piece.
Verification, the part everyone skips
I ran two different AI coding assistants against the same project, and they caught each other in something I didn't expect. One claimed a render script had a safety guard: if every visual layer was set to zero blur, the script would refuse to render. It said it tested this directly. The second agent went to verify that claim, opened the file, read it top to bottom, and found no guard at all.
Both were telling the truth. There were two files. Same filename. Different directories. One had the guard, patched that same session. The other was an older, unguarded copy one level up, never deleted, never referenced.
The fix wasn't a better guard. It was a policy: nobody's claim about "the file has X" is accepted until it names the exact live path and shows a reproducible command a second party can run. And when you find a superseded duplicate, don't delete it. Rename it:
mv layered_beat.sh layered_beat.sh.SUPERSEDED-no-guard
A disagreement between two honest reads is not noise to average away. It's a signal that you're not both looking at the same thing.
The same failure mode shows up in AI-generated code. The most dangerous AI-generated code is the code that passes all tests. I merged a PR that was green on every signal: CI green, tests green, lint green, build green. The state update did exactly what I asked. It also broke an invariant I never wrote down: status and permissions were supposed to change together. The tests covered the updated state. They didn't cover the relationship.
| Code | Tests | Risk |
|---|---|---|
| Obviously broken | Fail | Low, visible |
| Incomplete | Fail | Visible |
| Working happy path | Pass | Medium |
| Wrong assumption | Pass | High |
| Wrong business rule | Pass | Very high |
The dangerous code lives in the bottom two rows, exactly where you don't feel like you need to look. The review shift that works: stop reviewing the summary, review the assumptions. What invariant has to remain true for this to be safe? What happens on the unhappy path? Can I explain this change without reopening the AI conversation?
Understanding what agents actually do
ATLAS takes a different angle on verification. Instead of checking individual outputs, it learns a behavioral model from agent trajectories. It combines trace abstraction with automata learning to infer finite-state models of agent-environment interaction. The result is a state machine you can inspect: recurring behaviors, decision points, successful completion paths, failure loops.
Applied to a penetration-testing agent across 12 vulnerable machines, the learned models exposed high-level strategies that were invisible in raw execution traces. The same approach supports model-guided exploration, auditing, and knowledge transfer: symbolic model-based distillation from frontier models down to compact language models.
This matters because the current evaluation stack, task success rate plus execution traces, tells you what the agent did, not how it decided. When an agent burns $6,531, you want to know which decision point started the fire.
Common Pitfalls
Five things I've seen trip people up:
Fail-open permission engines. If a crash or malformed input defaults to allow, an attacker who can crash the engine gets unrestricted tool access. Fail closed, always. Unknown tool → deny.
Testing with mock agents only. Mocks always call the tool. Real 4B models sometimes answer textually instead. Every framework has quirks that only surface under real execution: coroutines that must be awaited, agents that silently do nothing without async for, docstring parsers that throw on missing argument descriptions. Budget for real-agent field tests from day one.
Treating the vector database as memory. Storage answers "can we keep this?" Memory answers "should we keep this?" Without a write boundary, you build a Digital Attic, and the attic poisons your context window on retrieval.
Failing CI on not-available. When the LLM doesn't call the tool, that's model nondeterminism, not a policy regression. Fail on unexpected-decision only, and commit a golden not-available allowance so your gate stays stable.
Trusting the summary over the assumptions. AI-generated code reads well by default. Review what the change assumed to be true elsewhere in the system, not what the summary says changed. And never accept "the file has X" without the exact live path.
One thing to remember: every agent failure in this article, the $6,531 bill, the merged PR that broke an invariant, the two agents reading different files, happened because a system trusted a model's proposal without a deterministic check outside the model. The LLM proposes. Policy, memory, and verification dispose.
The Bottom Line
- If you're building agents that call tools, adopt a four-state permission engine (allow, audit, escalate, deny) with fail-closed semantics. Binary allow/deny forces you to choose between over-privileged agents and approval fatigue.
- If you're building agents that persist state across sessions, skip the vector-database-only approach and add a write boundary with provenance. Storage without curation becomes a Digital Attic that poisons retrieval.
- One thing to watch: behavioral model extraction, ATLAS-style automata learning, is moving from research to tooling. Expect agent observability platforms to ship trajectory-to-state-machine analysis within the next year, and it will change how you debug agent failures.