Appearance
The edge deployment math just changed
Three things landed this month that make edge deployment practical. A post-training token-merging method for recursive vision transformers that cuts peak activation memory by 37% with zero retraining. A 3.1B vision-language model that decodes 228 tokens per second on a laptop and fits in about 3GB. And a 45M-parameter tool-calling agent that runs a full session in 28MB of RAM, shipped as a single 14MB binary.
These aren't incremental tweaks. They attack three different bottlenecks in the edge stack: the compute and memory cost of vision backbones, the model size floor for useful VLMs, and the memory budget for agentic tool calling. Here's how they stack up.
| Approach | What it is | Footprint | Best for | Accuracy / capability cost |
|---|---|---|---|---|
| MergeOver (SReT) | Post-training token merging for recursive ViTs | 37.3% less peak activation memory at batch 1 | Vision backbones on edge CPUs and small GPUs | 1.47 pp top-1 drop on ImageNet-1K |
| LFM2.5-VL-3B | 3.1B vision-language model | ~3GB, runs on phones and laptops | Screens, documents, grounding, tool calling | Matches most 4B-class VLMs on average |
| Needle 2 | 45M tool-calling model | 14MB binary, ~28MB RAM total | Agents, structured extraction, embedded devices | Trades wins with comparable models 5-70x its size |
None of them required a new training approach. MergeOver is purely post-training. Needle is a dense small model compressed to 2-bit weights with a baked-in engine. LFM2.5-VL-3B is a 3B model with day-one support across llama.cpp, MLX, vLLM, SGLang, and ONNX. The efficiency comes from tooling and architecture choices, not bigger training runs.
Key Numbers
- 37.3%: peak activation memory reduction from MergeOver at batch size 1 (38.4% at batch 16)
- 1.47 pp: ImageNet-1K top-1 accuracy drop for the selected MergeOver config
- 228 tok/s: LFM2.5-VL-3B decode speed on an M5 Max; 116 on a Ryzen AI Max+ 395; 20 on a Galaxy S26 Ultra
- 28MB: total RAM for a full Needle 2 session, including the 256-token context window
- ~3GB: memory footprint for LFM2.5-VL-3B, small enough for a phone
MergeOver: token merging without retraining
Vision transformers have a scaling problem on edge hardware. The quadratic attention cost means token count is the real budget, not parameter count. Recursive weight sharing, as in the Sliced Recursive Transformer (SReT), cuts parameters by reusing weights across blocks, but it leaves the token count untouched. You've solved the memory problem for weights, not for activations.
Token merging (ToMe) attacks the activation side by progressively merging similar tokens. But ToMe was designed for standard ViTs. Recursive transformers shuffle tokens across spatial permutations between iterations, which breaks the assumption that a merged token stays merged. MergeOver (arXiv:2608.13141) fixes this with three mechanisms: an unmerge tracking stack that restores tokens before spatial operations, a constraint-safe merge-rate adjustment that respects the recursive structure, and synchronized token-mass tracking across permutations.
The schedule matters as much as the mechanism. MergeOver reduces tokens once at the first block of each stage, then holds the sequence length fixed through the recursive iterations. That single-shot schedule avoids the compounding drift you'd get from merging at every block.
The ImageNet-1K results have a sign flip you need to catch. On GPU, peak activation memory drops 37.3% at batch 1 and 38.4% at batch 16. Throughput drops 21.7% at batch 1 but rises 21.7% at batch 16. On a Raspberry Pi 5, latency drops 2.4% at batch 1 and 17.6% at batch 16.
The batch-1 throughput loss is the unmerge tracking overhead. At batch 16, the memory savings let the kernel pack more work, and the overhead amortizes away. If you're benchmarking this yourself, test at the batch size you'll actually serve, not the one that fits in a demo notebook.
Quick Take: Token merging on recursive ViTs is a free lunch only if you track the unmerge; the 1.47-point accuracy drop buys back a third of your activation memory.
LFM2.5-VL-3B: a phone-sized VLM that holds its own
Liquid AI's LFM2.5-VL-3B pairs a SigLIP2 400M vision encoder with the LFM2.5-2.6B text backbone. It was pre-trained on roughly 34T tokens with 4x more vision data than the previous release, and the tokenizer was extended to 128K vocabulary in place to handle non-Latin scripts. Post-training runs in two stages: SFT with distillation from a larger teacher, then multi-reward RL.
The benchmark story has two headline jumps. Grounding on RefCOCO goes from 57.1 to 87.9, which is the difference between a model that vaguely points at objects and one you can use for actual detection workflows. Screen understanding goes from 6.0 to 78.7 on ScreenSpot-v2 Desktop, a 72-point swing that turns the model from useless to credible for UI automation.
| Benchmark | LFM2.5-VL-3B (3.1B) | LFM2-VL-3B (3.1B) | InternVL 3.5 4B (4.7B) | Qwen3.5-4B (4.7B) |
|---|---|---|---|---|
| MMStar | 63.3 | 57.7 | 65.5 | 59.3 |
| RealWorldQA | 73.1 | 71.1 | 67.7 | 67.1 |
| MMBench (EN) | 81.0 | 80.0 | 81.1 | 78.4 |
| MathVista (mini) | 68.5 | 62.1 | 67.1 | 63.6 |
| ChartQA | 81.3 | 80.4 | 86.2 | 84.2 |
| DocVQA | 91.1 | 89.8 | 91.8 | 94.8 |
| TextVQA | 84.3 | 83.0 | 77.5 | 81.2 |
| RefCOCO-avg (grounding) | 87.9 | 57.1 | 88.8 | 86.6 |
| BLINK (multi-image) | 61.5 | 50.2 | 57.2 | 58.7 |
| ScreenSpot-v2 Mobile | 81.2 | 7.6 | 87.8 | 81.4 |
| IFEval | 82.3 | 72.9 | 35.4 | 86.2 |
| ToolSandbox | 59.5 | 26.4 | N/A | 65.0 |
| BFCL V4 | 32.5 | 20.5 | N/A | 53.6 |
| Average | 69.4 | 57.2 | 69.4 | 70.1 |
Scores normalized to 0-100, evaluated with vLLM 0.26.0 in non-reasoning mode. InternVL 3.5 models don't support function calling.
The 3.1B model matches InternVL 3.5 4B's 69.4 average and sits just below Qwen3.5-4B's 70.1, while being roughly a third smaller. On tool use it lands between the 2B and 4B Qwen models. For a model that fits in 3GB, that's a strong position: you're getting 4B-class capability at 3B-class cost.
The speed numbers translate directly to deployment choices. 228 tokens per second on an M5 Max means real-time interaction with no perceptible lag. 20 tokens per second on a Galaxy S26 Ultra is slow enough that you notice it, but it's fully on-device, which matters for privacy-sensitive workloads. On the server side, the model hits about 11K output tokens per second at high concurrency, roughly 2x the 4B-class models. That's nearly a billion output tokens per day on a single H100, which changes the cost math for high-volume vision pipelines.
Needle 2: a tool-calling agent in 28MB
Needle 2 from Cactus Compute is the most aggressive footprint in this cluster. 45M parameters, compressed to CQ2-bit, baked into a single 14MB binary that runs a full session in about 28MB of RAM. To put that in context: 28MB is less than a single browser tab. It runs on devices where a 1B model won't even load.
The architecture is a Simple Attention Network: a Hadamard MLP replaces the FFN, GQA attention with engram key-value memory, multi-lane hyper-connections. The Hadamard transform is a fixed matrix applied in n log n time with no weights to read, which gets the mixing benefit of a large MLP without the parameter cost.
The design decisions around the engine matter more than the architecture. Tool calls come back as structured JSON, with a byte-level grammar compiled from your schemas constraining every token. The model can't produce malformed output; the grammar won't let it. Every response carries a calibrated confidence score from a learned head, so you can set a threshold, act above it, and escalate below it. And if you declare a large tool catalogue, a retrieval head renders only the top five tools per turn, with the grammar constrained to that subset.
The memory bound comes from a 256-token sliding window with tools pinned as KV sinks. Total memory stays near 28MB no matter how long the conversation runs. The tradeoff is obvious: this is not a model for long multi-turn reasoning. It's for single-shot tool calls, structured extraction, and device control, where the 256-token window is plenty.
The fine-tuning story is what makes it practical. LoRA on the frozen base, merge the adapter at export, and you get a tuned .cact file that runs on the same engine with no recompilation. The whole loop, from synthesized data to tuned binary, is a few CLI commands. For a model this small, that's the difference between a demo and a deployable product.
What the community is saying
The benchmark tables tell one story. The community is telling another. I ran the Q8 GGUF of Qwen3.8-27B on my Framework Desktop and one-shot a Super Mario clone with it. It's not fast, but it's extremely smart for overnight batches and background jobs. I'm still looking for ways to speed it up without losing accuracy, and I'm curious about MTP and the other quantization variants.
That post captures the mood around local models right now. People aren't waiting for vendor benchmarks. They're running 27B models on desktop hardware, spinning up browser demos of small models like minimax-h3 (the two HF Spaces versions have 254 and 104 likes respectively), and sharing results that read more like field reports than evaluations. The Mario clone is a better testimonial than any benchmark table, because it's a complete, working artifact produced by a model running in an office.
The through-line across these community reports: the constraint isn't accuracy anymore, it's speed. The Q8 quant of a 27B model is smart enough to ship, but slow enough that you schedule it like a batch job. That's exactly the gap that smaller models with better engines, like Needle 2 and LFM2.5-VL-3B, are trying to close.
Common Pitfalls
Merging tokens without tracking the unmerge. If you apply ToMe-style merging to a recursive transformer and skip the unmerge tracking, spatial alignment silently breaks for detection and segmentation heads. MergeOver's unmerge stack exists because this fails in ways that don't show up in classification accuracy but destroy localization tasks.
Benchmarking at the wrong batch size. MergeOver's GPU throughput is -21.7% at batch 1 and +21.7% at batch 16. If you benchmark at batch 1 and deploy at batch 16, you'll make the wrong call. Test at the batch size your service actually runs.
Writing vague tool descriptions. Needle's docs put it bluntly: describing your tools well is the whole game. The model reads your docstrings to decide what to call and how to fill arguments. A tool described as "gets weather" will get called with wrong arguments. A tool described with concrete parameter semantics will not.
Assuming the 256-token window keeps history. Needle pins tools as KV sinks, but the rest of the context is bounded. Long agent loops will lose earlier turns. Design for single-shot calls, or re-inject state explicitly each turn.
Trusting normalized benchmark scores across eval setups. The LFM2.5-VL-3B numbers were produced with vLLM 0.26.0 and each model's recommended generation parameters, in non-reasoning mode. Normalized scores hide harness differences. If a number looks too good, check whether the competing model was run under the same conditions.
One thing to remember
The pattern across all three of these is that efficiency is being added after the fact, not designed in from scratch. MergeOver is post-training. Needle is a dense model compressed to 2 bits with a purpose-built engine. LFM2.5-VL-3B is a 3B model with day-one support across every major inference runtime. The models aren't getting dramatically smaller through new training runs. The tooling around them is getting smarter, and that's where the deployment wins are coming from.
The Bottom Line
- If you're deploying a vision backbone on edge hardware, try MergeOver-style token merging on a recursive ViT before you retrain anything. You recover 37% of peak activation memory and 17.6% latency on ARM for 1.47 accuracy points, at zero training cost.
- If you need agentic tool calling in a truly constrained footprint, embedded, air-gapped, or battery-powered, Needle 2's 28MB session is the only option in this cluster that fits. Design around the 256-token window and you get a deployable agent on hardware that can't run anything else.
- If you need general VLM capability on-device, LFM2.5-VL-3B at ~3GB is the sweet spot for phones and laptops, matching 4B-class benchmarks at two-thirds the size. Watch the 2-bit quantization trend: Needle's CQ2 approach is likely coming to larger VLMs within a year, and that will shift the footprint floor again.
Sources
- MergeOver: Post-Training Token Merging for Recursive Vision Transformers, arXiv
- LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge, Liquid AI
- [multimodalart/minimax-h3](https://