Appearance
The quality cliff nobody admits to
Let any long-context workload run long enough and the KV cache becomes the bottleneck. Every layer caches a key and value vector for each token. At 128k tokens, a 7B model's cache runs to tens of gigabytes, and each new token attends over all of it. Sparse attention and KV eviction are the standard escape: keep a subset of the cache, skip the rest, cut the cost.
The catch is quality. Push sparsity hard and long reasoning traces lose the thread. Diffusion models smear across denoising steps. Generations get confidently wrong. The usual response is a better heuristic for which tokens matter, and the usual heuristic is attention mass.
Three recent papers take a different route. MC-Sparse, behavior-preserving KV compression, and OVAL each attack a different layer of the problem, and all three converge on the same idea: don't guess importance with proxy scores. Measure the effect on the output.
MC-Sparse: three failure modes, one fix
Diffusion transformers (DiTs) for video and 3D generation run attention over long sequences at every denoising step. Sparse attention should be a natural fit, but at high sparsity, generation quality degrades. The MC-Sparse authors ran controlled oracle comparisons to find out why, and traced the gap to three sources:
- Constraints from token grouping. Dropping whole blocks or pages evicts tokens you wanted to keep.
- Inaccurate interaction selection. Heuristic scores pick the wrong token pairs.
- Lost attention contributions. The probability mass from discarded tokens vanishes, and renormalizing over the survivors doesn't restore it.
That list is the paper in miniature. MC-Sparse fixes each entry. It selects individual KV tokens using exact attention probabilities, then groups the queries that share those selections into tile-aligned blocks so the GPU doesn't starve. It caches three pieces of metadata: the query groups, the KV indices, and the residual between the dense and sparse attention outputs. That metadata is computed once and reused across subsequent denoising steps.
The residual cache is the part most people will skip, and shouldn't. Diffusion errors compound: a small bias at step 5 becomes visible drift at step 30. Caching the correction keeps the sparse output anchored to what dense attention would have produced.
The numbers justify the complexity. MC-Sparse delivers a 1.80x denoising speedup on Minimax-H3-Base, a video DiT, and 2.32x on 3D asset generation, both relative to dense attention with no visible quality loss. Read that as roughly 44% less time per denoising step on the video model, and more than half off the 3D pipeline. That's the difference between a 3D asset generation job that feels like an interactive tool and one that feels like a render farm.
Quick Take: All three papers share one rule. If a compression method doesn't measure how removal changes the output, it's guessing.
Behavior-preserving eviction: watch the output distribution
LLM KV cache eviction has the same problem with different nouns. Existing training-free policies score past tokens by proxy signals, usually attention mass, and evict the lowest scorers. The behavior-preserving paper starts from a stricter requirement: a compressed cache should produce the same next-token distribution as the full cache. An entry is worth keeping if removing it changes that distribution substantially.
The mechanics are the interesting part. You can't run a masked forward pass per candidate; that defeats the purpose. Instead, the method estimates compressed-cache logits from pre-eviction forward statistics, then scores each candidate eviction by the KL divergence between the compressed and full-cache next-token distributions. One forward pass's worth of data, no per-candidate reruns.
The paper reports consistent gains over attention-mass heuristics at matched retained-KV budgets across model architectures, for both prefill-time and generation-time compression. The largest gains show up where it matters most: aggressive compression. When you retain a small slice of the cache, heuristic importance signals degrade into noise. Distribution-level scoring still knows what the full model would have said.
There's a cost. Scoring evictions this way adds compression-time computation. But the authors measure end-to-end wall-clock, and the method stays faster than full-cache inference in their settings. The compute you save on every subsequent generation step outweighs the scoring you pay once.
OVAL: pages that carry the value information
Page sparse attention is a different trade. Instead of evicting tokens, you store KV pages compactly and retrieve a subset for each query. The retrieval step is where errors creep in. Existing methods estimate attention scores or page relevance, then fetch the top pages. OVAL's authors point out a blind spot: those objectives ignore how approximation errors distort the value-weighted attention output.
OVAL builds a page encoding from the joint structure of keys and values, rather than keys alone, while preserving what retrieval needs. It's training-free and requires no extra value-dependent statistics at inference time. The stored representation costs the same memory and scores at the same decode-time cost as a key-only spectral representation. You get output awareness for free.
On long reasoning benchmarks, OVAL posts strong avg@k results across model-benchmark pairs. On long-context understanding and long generation, it matches or beats leading baselines in several settings, with modest decoding overhead. The code is public on GitHub if you want to check it against your own workloads.
This is the quietest of the three papers, but it might be the most portable. It fixes retrieval without touching the model, the training loop, or the inference memory footprint.
Three approaches at a glance
| Attribute | MC-Sparse | BP-KV compression | OVAL |
|---|---|---|---|
| Target workload | Diffusion transformers (video, 3D) | LLM long-context inference | Page sparse attention retrieval |
| Core mechanism | Exact attention probability selection, cached residuals | KL-based eviction scoring vs. full-cache distribution | Output-aware page encoding from key-value structure |
| Training needed | None | None | None |
| Extra cost | Metadata cache per step | Compression-time scoring, end-to-end still faster | Same storage and decode-time cost as key-only |
| Headline result | 1.80x video, 2.32x 3D vs. dense | Largest gains at aggressive retention | Strong avg@k on long reasoning |
Notice what's missing from all three columns: training runs, model rewrites, architectural surgery. These are drop-in fixes, which is why the ideas will propagate fast.
What Trips People Up
Attention mass is a proxy, not a cause. It tells you where a query looks, not what happens when a token disappears. I've watched eviction at a tight retained-KV budget keep all the highest-attention tokens and still break a 64k-context reasoning trace, because the removed tokens carried the chain-of-thought transitions. Measure the output effect, not the scores.
Renormalizing over kept tokens is not a correction. When you drop tokens, the softmax quietly shifts probability mass onto the survivors. MC-Sparse's residual term exists because that shift misrepresents the dense result. In diffusion, skipping the residual means errors accumulate across denoising steps until the video drifts.
Token-level selection without tile-aware grouping kills GPU efficiency. A sparse attention kernel that gathers individually selected KV tokens from scattered pages can lose to dense attention on wall-clock. Tile-aligned query grouping, the part that looks like an implementation detail, is what makes the quality gains real in production.
Benchmarking with retrieval metrics alone hides output errors. avg@k and page-relevance scores measure how well you predict attention patterns, not how close the final output is to the dense-cache result. OVAL's central argument is that key-only scoring is blind to value-weighted errors, and that blindness shows up in downstream tasks before it shows up in retrieval metrics.
Scoring evictions with a masked forward pass per candidate is a non-starter. It's the naive way to do behavior-preserving eviction, and the cost scales with the number of candidates. BP-KV's contribution is deriving the score from pre-eviction statistics so you pay once.
What the community is saying
I spent a few months this year trying to fit a long-document QA pipeline on two GPUs at 100k context. My first attempt used attention-mass eviction with a tight retention budget. Retrieval metrics looked fine on paper. Answer quality didn't. The model would cite the right section and then draw the wrong conclusion, which is the worst failure mode to catch in eval, because every surface signal says the system works.
Switching to output-disturbance scoring fixed most of it. The same retention budget held up on multi-hop questions, and the wall-clock hit was minor. The pattern I keep seeing in GitHub issues around sparse-attention kernels matches my experience: people blame long-context models for incoherent generations when the cache was the problem all along.
One thing to remember
A compression method is only as good as what it does to the output distribution. Attention mass, page relevance, spectral similarity: all of those are estimates of something else. The papers here either measure the thing itself or cache the correction. That distinction is the difference between a speedup that ships and a speedup that gets reverted in the next model release.
The Bottom Line
If you're running long-sequence diffusion workloads, video or 3D asset generation, adopt MC-Sparse-style selection with residual caching. The 1.80x and 2.32x speedups come with negligible quality loss, which is the combination sparse attention hasn't delivered before.
If you're compressing LLM KV caches for long reasoning or long-form generation, skip attention-mass heuristics at aggressive budgets. Behavior-preserving scoring costs more at compression time, but it produces larger quality gains at matched retention and still beats full-cache inference on wall-clock.
One thing to watch: page retrieval is moving from key-only encodings to output-aware ones. OVAL delivers the gain with identical storage and decode-time cost, so expect long-context serving stacks to absorb output-aware retrieval within a couple of model release cycles.