Skip to content

The Best Agent Tools Don't Trust the Agent

#ai-agents #mcp #agentic-tooling #llm-tools #agent-skills #open-source

Six projects, one question ​

Six agent projects landed in the same week. A markdown publishing kit that ships one article to four platforms. A Blender pipeline that made Spider-Man swing on a web without a single hand-keyed frame. An open-source video studio that produced a 60-second animated short for $1.33. An agent that settles disputes between AWS's own documentation. A self-hosted quant workstation with 12 MCP tools and a full-market screener that runs in milliseconds. And Meta open-sourcing the hardware layer of its consumer agent, Muse.

Different domains, same lesson. The tools that work don't trust the model. They wrap it in deterministic code the LLM is not allowed to override, connect it to the world through MCP, and make failures loud before anything ships.

The reason is consistent: the weakest part of any agent system is model confidence, and confidence scales faster than accuracy. Every project here found the same fix. Decide what the model is allowed to touch, then make the exact parts run as code.

Key numbers from the cluster: 4 publishing destinations from one markdown source · 3 conflicting AWS figures for a single quota: docs say 32, the record says 5, live AWS says 16 · $1.33 total cost for a 60-second animated short · 78 Mixamo animation clips, downloaded one click at a time · 5000 units in the first free batch of Muse Home Link.

MCP is becoming the universal connector ​

Start with blender-mcp, because it's the cleanest example of the architecture. A first-time Blender user used it to rig Spider-Man and chain motion-capture clips into a coherent 16-second performance. The setup has two halves: an add-on inside Blender opens a local socket server, and an MCP server outside speaks the Model Context Protocol to Claude Code on one side and forwards JSON commands down that socket on the other.

The most powerful tool it exposes is execute_blender_code, which runs arbitrary Python inside Blender through bpy, Blender's Python API. Every button in Blender is a thin wrapper over bpy, so an agent that can write bpy can press every button. The loop that makes it work is the screenshot tool coming back the other way: the agent changes something, grabs the viewport, looks at it, and decides what to fix next. Same edit-run-inspect cycle you use for code, pointed at a 3D scene.

One honest warning from the author: execute_blender_code is arbitrary code execution by design. Treat it like local shell automation, and don't point it at a machine full of secrets. The same warning applies to every browser-driving helper described here.

The quant workstation, tick-stock-panel, pushes the MCP pattern further. It exposes 12 MCP tools, but each one is filtered through a token gateway with six permission scopes, and the permission check reuses the same decision point as its 61 REST endpoints. Paper trading is deliberately not exposed to the AI. The MCP server is a thin bridge with zero business logic. Human users get the password session, external programs and AI clients get the token gateway, and neither crowds the other's rate limit. The full-market screener scans thousands of A-shares in milliseconds because metrics are precomputed in an enriched columnar layer instead of recalculated per query.

Write once, publish everywhere ​

publishing-kit is an agent skill for Claude Code, Codex, and Antigravity. One dev.to markdown file goes in, and a dev.to draft, a Medium story, an AWS Builder Center draft, and a LinkedIn post come out. The author's framing is the whole point: markdown is the easy part. Publishing is the part that fails, silently, four different ways per destination.

DestinationPublishing APITablesMulti-line codeCover
dev.toFull REST APINativeNativeURL, cropped to 2.381:1
MediumNone since 2023DroppedFlattened to one lineFirst image in body
AWS Builder CenterNoneNativeNative, line-numbered1200x675 upload
LinkedInPosts API only, no draftsNoneNoneLink card or upload

Every one of those failures is silent. You get a plausible-looking page with a missing table, a code block run together on one line, or a cover with the title cut off, and you see it after it's live.

Medium is the worst, because it has two entry points that disagree with each other. The importer drops tables, flattens code blocks, renders only two heading sizes, and caches by URL while ignoring the query string, so re-importing a fixed page with ?v=2 brings back the old one. Pasting into the editor keeps multi-line code but drops any data: URI image with no placeholder left behind. The REST API that would avoid all this is archived, and Medium's own docs open with "the Medium API is no longer supported." So the kit drives the browser editor, quirks known in advance.

The design worth copying is the split. make-medium.py, make-builder.py, and make-linkedin.py generate one artifact each. The prompt-level skill tells the agent what each destination needs. preflight.py checks the cover, front matter, prose, links, and numbers, and exits non-zero on any failure. The agent runs the scripts, it doesn't improvise the parts that should be exact. I ran the pre-flight mid-production and it failed on the LinkedIn links, because the Medium and Builder Center versions were still pending. That failure saved a post full of dead links.

Quick Take: An agent skill is just a prompt plus scripts that refuse to run when something looks wrong. The refusal is the feature.

When the model is not allowed to decide ​

The strongest example of the don't-trust-the-model pattern is the AWS source-of-truth agent. The problem it attacks is real: for one AWS fact, three official pages can give three different numbers. An old User Guide says one thing. The Service Quotas console says another. A pricing page says a third. All look official.

The demo case is brutal. For the EC2 On-Demand Standard vCPU quota, the docs say 32, the reconciled knowledge-base record says 5, and the live account says 16. For EBS gp3 max IOPS, the stale legacy guide contains the exact phrase you'd type into a search box: "maximum IOPS per volume for a general purpose SSD EBS volume." Plain keyword retrieval ranks the wrong 16,000 answer at 0.86, more than four times the score of the correct 80,000. Keyword search hands you the wrong number and sounds sure of it.

The architecture is where it gets good. Facts are typed awsFact documents in Sanity with fields like source.kind and effectiveDate, not prose. The agent pulls them over the hosted Context MCP with knowledge_base_search and knowledge_base_read. But the model writes the sentence, it does not get to decide the number. A plain function reconciles the value using source precedence (console and pricing pages beat changelogs, which beat official docs, which beat blogs), with the newest effectiveDate breaking ties. A guard then checks the model's answer against the deterministic result. On one run, Amazon Nova Pro muddled its own wording, the guard caught it, and swapped in the correct answer without human help.

Then the twist: even the reconciled record can be stale. So the agent calls the live AWS API, read-only, and asks whether the record is still true right now. That's where DRIFT comes from.

FactReconciledSupersededLive AWSResult
EC2 vCPU quota5 (console)32 (old guide)16DRIFT
EBS gp3 max IOPS80,000 (console)16,000 (old SSD guide)unavailabletrusted
S3 Standard $/GB-mo0.023 (pricing)0.021 (stale blog)0.023AGREE
RDS PostgreSQL oldest major13 (release notes)11 (old tutorial)11DRIFT
Lambda concurrent executions1000 (dev guide)n/aunavailabletrusted
Graviton4 (R8g) availabilityAvailablen/aAvailableAGREE

The DRIFT label means the record and reality disagree. That's the honest answer: this is the current record, and reality has already moved past it. A number without its predecessor is unfalsifiable, which is why the agent shows the replaced value next to the live one.

tick-stock-panel applies the same philosophy on the human side. Its AI assistant has 21 query tools and 4 action tools for things like generating a signal strategy or running a backtest. Every write action goes through a confirmation card with the full parameters visible, and a 120-second timeout auto-cancels if nobody confirms. The model is explicitly told not to retry. Propose, then let someone with authority dispose.

From prompt to final render ​

OpenMontage makes the boldest claim in the group: "the first open-source, agentic video production system." You describe a video in plain language, and the agent handles research, scripting, asset generation, editing, and composition. The distinction it insists on matters: it can make real videos from actual motion footage, not just animate a handful of stills. The documentary pipeline builds a CLIP-searchable corpus from Archive.org, NASA, Wikimedia Commons, and free stock sources, retrieves real clips, edits them into a timeline, and renders the finished piece.

The economics are the story. "The Last Banana," a 60-second Pixar-style short, used 6 Kling v3 motion clips via fal.ai, Google Chirp3-HD narration, royalty-free piano, TikTok-style captions, and Remotion composition. Total cost: $1.33. "Reimagine Your Universe" ran about $4. A seven-world showcase ran about $5. At that price, iteration stops being scary.

The pipeline discipline matters more than the models. Every production runs through pipeline_defs, stage director skills, and a tool registry. The storyboard is an approval gate: asset generation pauses on a scene-by-scene contact sheet with takes, prompts, per-asset cost, and quality scores. Before you see anything, a multi-point self-review runs ffprobe validation, frame sampling, audio level analysis, and subtitle checks. Provider choices are scored across 7 dimensions with an auditable decision log. The agent is the project manager, not the person improvising.

The Blender and Mixamo pipeline behind the Spider-Man piece shows what's under the hood. A 3D character is a mesh at rest. To move it you need an armature (a tree of bones, like a transform hierarchy), skin weights (every vertex stores how much each bone influences it, which is what lets an elbow bend instead of folding like a paper straw), and a rest pose, the reference every animation is measured against. Forward kinematics sets every joint angle and hopes the hand lands on the doorknob. Inverse kinematics says where the hand should be and lets a solver find the joint angles. Blender ships exactly that constraint solver.

Mixamo is rigging as a service: export a T-pose FBX, place five landmarks (chin, wrists, elbows, knees, groin, symmetry on, so really three clicks per side), and it infers the whole skeleton. Download one clip with skin, the rest without, because every clip targets the same 65-bone skeleton and the skin only needs to come along once. Chain clips on the Nonlinear Animation timeline with crossfades, stitch root motion by reading where the hips ended in one clip and offsetting the next to begin there, same idea as stitching GPS tracks. Iterate at 360p, because an 8-minute render is fine once and miserable forty times. Make the feedback loop fast first, then make the output good.

Even loading the model had lessons. The texture files pointed at the original author's Windows desktop folders, and missing normal and specular maps rendered the suit hot pink, the 3D equivalent of a 404. The agent relinked what it could and disconnected the rest. The face looked like a horror movie until someone realized the 143-bone face rig was set to draw in front, painting every bone on his skin. One checkbox fixed it.

Agents are leaving the screen ​

Muse Gadgets is the curveball. Muse is Meta's personal AI agent: it handles email, shopping, trip planning, and background tasks. Last week Meta open-sourced the hardware around it, so anyone can build Muse peripherals.

The practical details: developers can run Muse firmware on an ESP32 board, which puts agent interaction logic on a microcontroller that costs a few dollars, or use the Linux SDK on a Raspberry Pi 5-class device for more ambitious edge setups. Meta published reference forms: color E Ink displays as low-power desktop reminder boards, HDMI display sticks that project the agent's interface onto a TV, and a touch pendant that gives you a direct line to your agent without smart glasses. Alongside the open source, Muse Home Link is a USB-C powered hub that connects Muse to your home network so it can talk to TVs, speakers, and anything with an HTTPS endpoint. First batch: 5000 units, free to qualifying Muse subscribers.

The onboarding path is agent-first. Get a token from gadgets.muse.ai, import the repo into Claude Code, Cursor, or GitHub Copilot CLI, and let the coding agent generate the device driver code from the protocol. People are already shipping. One developer turned an Xteink e-reader into an always-on Muse display, mounted on the back of a phone with MagSafe, showing the agent's status updates. Another built a wine-cellar tracker from a ~$50 SenseCAP Watcher: photograph a bottle label, Muse identifies the winery and vintage, and a Tailscale connector syncs it to a personal cellar database. Under an hour from unboxing to working device.

The structural shift is worth sitting with. Hardware used to flow one way: company designs device, user buys it, developers write apps for it. Meta's bet is the reverse: the company provides the agent capability, and developers invent the device forms. When Meta announced the project, Alexandr Wang said the reasoning was simple: building gadgets is fun, and the company wants to see what developers make when the hardware layer is open. Muse is the brain. The community gets to build the bodies.

What the community keeps tripping on ​

The best commentary this week came from people running agents in production, and most of it reads like war stories. I run a model-routing layer across 32 models behind one API key, and the scaffolding is the only reason it holds together. When a provider throttles mid-task, the state machine decides retry versus failover versus error. The LLM can't make that call reliably. Debugging a flaky agent is usually debugging a missing state machine around a good model.

The AWS agent's own failure log tells the same story. A wrong factType guess used to lose a real fact, because the typed filter matched nothing and the fact vanished. The fix was mechanical: a zero-row fetch retries on service and region alone. When Nova dropped a tool argument and passed an empty string to the reconcile function, the guard fail-closed to "Not verified" instead of guessing. The fix was caching the last real fetch server-side, so facts always come from Sanity, never from whatever the model relayed. Every bug in that log was a scaffolding bug, not a model bug.

The sharpest edge case I hit was Claude Code's MCP tool output limit. The docs say 25,000 tokens, but on v2.1.273 the count only ran once a result was already long in characters, around 45,000 to 52,000, so 24,000 characters of CJK text went in uncounted and grew the next request by roughly 50,000 tokens. The page wasn't out of date. It just doesn't say the limit is only checked once a result passes a certain length.

I also found a problem with the DRIFT flag: it mixes two different questions. GetServiceQuota returns the applied quota for an account, including any increase that account requested. GetAWSDefaultServiceQuota returns the default. A live value of 16 versus a record of 5 can mean "this account asked for more