Appearance
The Training-Free Inference Stack: KV Virtualization, Input-Adaptive Compute, and Speculative Decoding
The new bottleneck is the second token
The cheapest token is the one you don't generate. The second cheapest is the one you generate without a full matrix multiplication. The third is the one a small model drafts and a big model verifies in parallel, instead of decoding one step at a time.
That's the throughline of the last month in LLM inference. Five papers landed within days of each other, and every one of them is training-free. No fine-tuning, no weight changes, no distillation. They reclaim dead KV cache memory, skip matrix multiply slices, rotate quantization outliers, and draft tokens ahead of the main model. The serving layer is moving just as fast: OpenAI previewed a Cerebras-powered API tier that hits 750 output tokens per second, Baseten joined Hugging Face's Inference Providers, and SGLang released a multi-stage runtime for speech and omni models.
The shift makes sense if you look at where the money goes. Training a frontier model is a nine-figure event. Serving it is a nine-figure recurring cost. When you can't afford to retrain, you optimize the serving path. And the serving path has four attack surfaces: memory, compute, precision, and decoding strategy.
Memory is the first wall. The KV cache grows with sequence length and batch size, and it's the reason your GPU fills up before your compute does. Compute is next: every generated token is a sequence of high-dimensional matrix multiplications, and a large fraction of that work is redundant for any given input. Precision is the classic lever: drop from fp16 to 4-bit and you cut memory bandwidth, but you pay in outliers and perplexity. Decoding strategy is the newest lever: generate draft tokens cheaply, verify them in parallel, and only pay for the tokens the big model actually accepts.
There's a pattern across all five papers that I want to name up front. They don't just report speedups. They report where the speedup comes from, which component of the model tolerates the intervention, and where the technique breaks. That's the difference between a benchmark and a deployment tool. The vToken paper quantifies stranded memory. The RMM paper identifies attention as the reducible half of the Transformer. The quantization paper explains exactly why its theoretically clean rotation fails. Those details are worth more than the headline numbers.
This article walks through each layer with the latest results, then looks at the serving infrastructure that packages it all. The short version: the optimization stack is deep enough that the interesting question isn't which single technique to adopt. It's which combination, and in what order.
Four layers, one goal
Five papers, four layers, one goal: more tokens per dollar without touching the weights.
| Technique | Layer | Method | Headline result |
|---|---|---|---|
| vToken | Memory | Token-level virtualization with async repacking | 27-72% fewer retained KV blocks, up to 1.37x throughput |
| RMM | Compute | Input-adaptive slice selection on contraction dims | Wall-clock gains at long context, works from 1B to 70B |
| RoPE-aligned rotations | Precision | Per-pair Q/K rotation for 4-bit PTQ | Negative result: full-head mixing still wins |
| SPADE | Decoding | Edge draft model, cloud verifier | 76% fewer cloud calls, zero accuracy loss |
| mlx-dspark | Decoding | DeepSeek DSpark drafters on MLX | 2.45x mean speedup on Apple Silicon |
| SGLang-Omni | Serving | Multi-stage runtime for speech and omni | Day-0 support for new TTS and music models |
| GPT-5.6 Sol Ultrafast | Serving | Cerebras hardware tier | Up to 14x faster, 750 tok/s |
The layers feed into each other. Memory, compute, precision, and decoding optimizations all land in the serving infrastructure, which is where they become something a developer can actually call.
Key numbers27-72% fewer retained KV blocks per request with vToken's repacking. 76% fewer cloud model calls with SPADE's edge draft, cloud verify split. 2.45x mean speedup on an M4 Pro with mlx-dspark's 8-bit Qwen3.8-27B. 750 output tokens per second on the Cerebras-powered Ultrafast tier. 50 lines to integrate vToken into vLLM, down from 500+.
Before going layer by layer, one framing point. These techniques target different bottlenecks, which is why they compose. vToken frees memory, RMM skips compute, the rotation work tries to reduce quantization error, and speculative decoding trades a little extra compute for a lot of latency and cost reduction. The results are also measured in different settings: vToken on vLLM with eviction policies, RMM on A100s with custom kernels, SPADE across edge and cloud, mlx-dspark on a laptop. Direct comparison is impossible, and that's fine. The point is the shape of each result, not a leaderboard.
Layer 1: reclaiming the KV cache
The KV cache is the first wall you hit when you try to serve at scale. Every token in the context contributes a key and value vector per layer per head, and the cache grows linearly with sequence length and batch size. At long context with a large model, the cache alone can dwarf the weights. This is why vLLM's PagedAttention became the default answer: fixed-size memory blocks, virtual memory-style mapping, minimal allocator fragmentation.
But there's a mismatch hiding underneath. The recent generation of KV eviction policies, H2O, Scissorhands, Random, operates at token granularity. They score individual tokens, decide which ones matter, and evict the rest. PagedAttention manages memory at block granularity. When a policy evicts one token from a block, the block stays allocated as long as any token in it is live. The dead tokens' memory is stranded. The vToken paper calls this intra-block fragmentation, and it's the reason your KV memory stays pinned even after eviction has supposedly freed it.
vToken is a lightweight token-level virtualization layer that decouples logical token liveness from physical block placement. It maintains a token table, an indirection layer that gives the model a stable logical view of the tokens, and it realizes physical reclamation by repacking live tokens asynchronously. The model's attention computation never sees the movement. The token table handles the indirection, and the repacker compacts live tokens into denser blocks in the background.
The design constraint that makes this production-relevant: it preserves PagedAttention kernels and CUDA Graph compatibility. That matters more than it sounds. vLLM's performance is baked into those kernels, and any optimization that forces you off them is a non-starter. vToken works inside the existing execution model, which is why the integration footprint is small enough to survive contact with a production fork.
The results, measured against a paired Naive-Evict baseline across H2O, Random, and Scissorhands policies and multiple models: retained KV blocks per request drop by 27.2% to 72.3%. SLA-constrained throughput goes up by up to 1.37x. Under a constrained active-KV budget, the maximum feasible concurrency doubles. And the integration footprint in vLLM drops from 500+ lines to under 50.
Let me translate those numbers into deployment terms. 27-72% fewer retained blocks means 27-72% more requests fit in the same GPU memory, depending on the eviction policy and how dead the evicted tokens actually are. 1.37x SLA-constrained throughput means 37% more requests served within their latency targets on the same hardware. 2x concurrency means the same memory budget serves twice the batch. The 50-line footprint is the detail that makes it credible. A virtualization layer that needs 500 lines of surgery won't survive contact with a production fork.
The broader point: eviction policies were already squeezing value out of the KV cache, but the memory they freed was partly fictional. vToken is the layer that makes the eviction real. If you're running H2O or Scissorhands-style eviction on vLLM today, you're leaving a measurable fraction of your KV memory stranded. This is the cheapest memory you'll ever recover: no retraining, no new hardware, no accuracy loss.
Quick Take: The biggest near-term wins in LLM serving are training-free, and they stack, so the sensible deployment question is which combination to run, not which one to pick.
Layer 2: skipping matrix multiplications you don't need
RMM attacks the compute layer, and its premise is almost embarrassingly simple. Transformer inference is repeated high-dimensional matrix multiplications. Most of those multiplications are doing work the current input doesn't need. RMM selects informative slices along the contraction dimensions of the matrix products, per input, and skips the rest. No weight modification. No training.
The contraction dimension is the shared inner dimension of a matrix product. When a weight matrix multiplies an activation, the inner dimension is where the per-input information lives. RMM's insight is that you can keep a subset of those slices, chosen adaptively for each input, and control the trade-off with a single retention-ratio knob. That gives you a smooth, predictable accuracy-efficiency trade-off, which is exactly what you want when you're tuning a production deployment. You set the knob, you measure the quality impact, you adjust. No cliffs, no sudden quality collapse.
The results span models from 1B to 70B parameters, and the shape is consistent. Reduction tolerance depends on the model family, the task, the component, and the retention ratio, but it often improves with model scale. Bigger models tolerate more reduction. If you're serving a 70B model, you can be more aggressive than a 1B model, which is convenient because the 70B model is also the one where the compute savings matter most. A 1B model has less redundancy to exploit; a 70B model has plenty.
Under moderate reduction, RMM stays robust across discriminative tasks, autoregressive generation, and long-context settings. It also extends to multimodal vision-language inference, which suggests the principle isn't specific to text Transformers. That's a wider applicability claim than most inference optimizations make, and it held up in the paper's evaluations.
The mechanistic finding is the one that matters for deployment: attention-side computations are substantially more reducible than MLP components. That's a structural asymmetry inside the Transformer, and it tells you exactly where to apply the reduction. Be aggressive on the attention projections. Be conservative on the MLP blocks. A uniform retention ratio across both leaves performance on the table, or breaks quality, depending on which direction you tune it. The per-component tuning is the difference between a technique that works in a paper and a technique that works in production.
The wall-clock benchmarks on an NVIDIA A100 with custom kernels show the savings translate into runtime gains, especially at longer sequence lengths. That last clause deserves emphasis: the gains show up when the workload is compute-bound, which is exactly what long-context generation is. At short sequence lengths, the overhead of slice selection can eat the savings. If your deployment is mostly short prompts, benchmark carefully before adopting this. If you're serving 8K or 32K contexts, this is where the compute budget actually goes.
RMM's position in the stack is complementary to everything else here. It reduces the compute per token, which is orthogonal to the memory savings from vToken and the latency savings from speculative decoding. The retention-ratio control also makes it predictable to operate: you get a dial, not a cliff.
Layer 3: the quantization result nobody wanted
This is the section where I recommend a paper that mostly reports a negative result. That's rare, and it's valuable.
Rotation-based post-training quantization, the QuaRot and SpinQuant family, applies an orthogonal transform across an entire attention head to spread outlier mass and reduce quantization error. RoPE complicates the picture. It partitions each head into 2D frequency pairs, and a rotation that respects that decomposition has to commute with RoPE. Prior work established which per-pair rotations commute. This paper proves the converse: for distinct frequencies, no other single-head orthogonal map commutes with RoPE. The design space is fully mapped, which is a useful theoretical result on its own.
The empirical question was whether a pairwise rotation that respects the RoPE decomposition could beat full-head mixing. The answer, tested in a dynamic W4A4KV4 setting across four checkpoints, is no. Replacing the full-head Hadamard with the head-shared pairwise configuration increases perplexity at both short and long context lengths. Composing the pairwise rotation with the Hadamard satisfies the selected ±0.05 PPL interval criterion under the default estimator, which is a pass but not a win. Estimating the shared angle from K alone improves the pairwise-only configuration on every checkpoint, but doesn't close the gap to full-head mixing.
The explanation is the part to internalize. The analytic objective controls a position-averaged second moment of a pooled calibration covariance. The dynamic quantizer sets its step from a tokenwise group range. Those two things don't align. The pairwise transform also has only two-channel mixing support. Along a controlled interpolation from two-channel to full-head mixing, K range, relative quantization error, and perplexity degradation all decrease as support increases.
In plain English: the rotation that's optimal for the smooth surrogate objective isn't optimal for the quantizer that actually sets the scale. Full-head mixing wins because more channels participate in absorbing the outlier mass. A pairwise rotation can only spread an outlier across its two channels, and that's not enough room. The math was clean. The quantizer didn't care.
What should you do with this? If you're choosing a quantization scheme, check that the optimization objective aligns with the quantizer's scale-setting statistic. If it doesn't, the theoretically optimal rotation can make things worse than the plain baseline. The paper's title says it: when local variance optimality is not enough. The surrogate was optimal. The quantizer was measuring something else.
This is also a reminder that the quantization layer is the most fragile of the four. Memory, compute, and decoding optimizations preserve output distributions by construction. Quantization changes the numerics, and clever rotations can make the numerics worse if the objective is misaligned. Measure perplexity on your own workloads, at your own context lengths, before trusting a quantization paper's headline number.
Layer 4: speculative decoding goes everywhere
Speculative decoding is the rare optimization that improves latency and cost at the same time. The mechanism: a small draft model generates candidate tokens quickly, the target model verifies them in parallel, accepted tokens are kept, and only rejections trigger correction. The draft tokens are cheap, the verification is parallel, and the output distribution is unchanged because the target model checks every token. That last property is why it's called lossless, and it's the reason speculative decoding has become the default answer to the question "how do I make my LLM faster without changing the model?"
Two projects this month show the same principle at very different scales.
SPADE splits the draft and verify steps across devices. A compact draft model runs on the edge and generates candidate tokens rapidly. A large verifier model on the cloud validates them in parallel. Accepted tokens are retained, and only rejections trigger verifier correction. The result: 76% fewer cloud model calls with zero loss in accuracy, no retraining. That's roughly a 4x reduction in API cost for the same output quality. The latency profile improves too, because drafting is local and verification is parallel, so the perceived speed is the draft speed, not the round-trip speed.
The deployment pattern is worth spelling out. You put a small model on the edge device, where it costs nothing per token. It drafts. The cloud model only gets involved to verify and correct. Since acceptance rates are typically 60-80% on in-distribution text, the cloud does a fraction of the work it would otherwise do. The paper evaluates across SpecBench and CNN/Dailymail, and the 76% reduction holds with zero accuracy loss. The plug-and-play design matters: no retraining, no distillation, no shared weights between draft and target. You can pair any small drafter with any large verifier.
mlx-dspark co-locates everything on one machine. It's an MLX port of DeepSeek's DSpark speculative-decoding drafters, plus z-lab's DFlash, with a single lossless verify loop. The latest release adds Qwen3.8-27B support via RadixArk's drafter, and the numbers on an M4 Pro with 48GB are the best local-inference results I've seen reported this year. The 27B model