Appearance
Long-context inference has a dirty secret. Compute scales with the tokens you process, but the KV cache scales with the tokens in the window, and the cache is what breaks your GPU first. For a 32-layer transformer with 32 KV heads and 128-dimensional head states, every token you keep in context costs 512 KiB at fp16. A single 128K-token sequence carries a 64 GiB cache. That's more HBM than the weights of an 8B model.
Two 2026 papers attack this wall from opposite ends. WUSH-KV (arXiv 2609.38121) makes the KV cache survive at 2 bits by quantizing in a coordinate system built from the actual statistics of K and V. Delta-Matching (arXiv 2609.37852) removes the last precision barrier in native FP8 training, the softmax backward pass. One shrinks what inference has to hold; the other shrinks what training has to spend. Same bottleneck, two sides.
The KV cache is the wall
The math is simple and brutal. Per token, the cache stores K and V for every layer:
2 (K and V) × 32 layers × 32 KV heads × 128 head_dim = 262,144 values per token.
At fp16, that's 512 KiB per token. Multiply by 131,072 tokens and you get 64 GiB per sequence. One sequence. An 8B model at fp16 is about 15 GiB of weights, so the cache alone runs more than four times heavier than the model producing it. When I ran this arithmetic for my own serving setup, the conclusion was uncomfortable: past roughly 32K context, the cache decides how many sequences fit on a GPU, long before the weights do.
Per-token cost: 262,144 values for K+V in a 32-layer transformer, 512 KiB per token at fp16. At 128K context that's ~64 GiB per sequence, more than four times the weights of an 8B model. At 2 bits it's ~8 GiB, small enough that one GPU can hold the cache and the model together.
The bandwidth side is just as bad. Decode is memory-bound: every generated token reads the entire cache. At fp16, a 128K-cache generation step shuffles 64 GiB through HBM; at 2 bits, 8 GiB. Quantization is the only lever that attacks capacity and bandwidth at the same time.
This is why KV quantization isn't a nice-to-have. It's the difference between serving one 128K sequence per GPU and serving eight.
WUSH-KV: quantize in the right coordinate system
Naive 2-bit quantization destroys long-context quality because it rounds values in the raw activation space, where outliers and correlated structure wreck the error profile. WUSH-KV's premise: quantization error depends on the basis you're rounding in. WUSH constructs a data-aware transform from the second-order statistics of both factors in a matrix product, so the quantization happens in a coordinate system where the values are better behaved for clipping and rounding.
The KV-specific design choices matter more than the transform itself. WUSH-KV builds separate transforms for K and V from calibration data, because the two tensors play different roles. V gets contracted against the attention probabilities, so its error interacts with how confident the model is about each position. K gets rotated by RoPE inside the QK dot product, so its coordinate system is tied to the rotation.
Two implementation details make this practical. The value transform folds into the output projection weights, which means the V cache is quantized in the transformed basis at zero inference cost. The key transform is applied after RoPE, not before, because RoPE is a rotation and any transform applied ahead of it gets tangled with the frequencies you're trying to preserve. The key transform does run at inference time, but that compute is noise compared to the bandwidth you're saving.
WUSH-KV pairs the transform with clipped quantizers. For QuEST INT, the paper shows the WUSH transform is near-optimal under mild assumptions: with that quantizer family, you're not leaving much on the table. In layerwise reconstruction error and end-to-end perplexity, WUSH-KV beats the other tested transforms. The evaluation that matters for serving stacks is the SGLang integration, which uses OSCAR-style percentile-clipped affine quantization. At 2 bits, WUSH-KV is comparable to or better than the OSCAR transform across every tested model and downstream task.
| Variant | What changes | Footprint at 128K | Reported quality |
|---|---|---|---|
| fp16 baseline | no quantization | ~64 GiB per sequence | reference |
| 2-bit affine | percentile clipping only | ~8 GiB per sequence | quality degrades on long context |
| 2-bit + OSCAR transform | fixed transform from calibration stats | ~8 GiB per sequence | strong baseline |
| 2-bit + WUSH-KV | data-adaptive K/V transforms, V folded into weights, K after RoPE | ~8 GiB per sequence | lowest end-to-end perplexity; matches or beats OSCAR on tasks |
The result is that the 8x memory reduction doesn't cost you the perplexity you'd expect from dropping to 2 bits. The transform redirects the error away from the structure the softmax depends on.
Quick Take: Both papers fix error structure rather than raw precision: WUSH-KV picks a coordinate system that keeps 2-bit rounding from destroying softmax inputs, and Delta-Matching preserves a gradient invariant that keeps FP8 error from accumulating.
Delta-Matching: the stale-delta problem
Training has the mirror-image problem. FP8 linear GEMMs in forward and backward became routine after DeepSeek-V3 showed the pattern at production scale. The stubborn holdout is attention: the backward pass through softmax needs the probabilities saved from the forward pass, and when forward and backward disagree on precision, the delta term gets computed against stale values. What should be a corrective refinement becomes a slowly compounding bias.
Delta-Matching's first contribution is diagnosis. The paper derives how forward-backward inconsistencies produce the stale delta and shows what it does to training dynamics. Then comes the empirical part, which is where the warning lives. A stale-delta hybrid shows a modest loss gap at 569M parameters. At 1.67B and 5.29B, the same setup produces substantial loss increases and downstream degradation.
The fix targets the invariant that keeps attention gradients honest. Each row of the softmax probability matrix sums to 1, so the gradient with respect to the scores must sum to zero across the row. FP8 rounding breaks that property, letting a net bias creep into every attention gradient. Delta-Matching restores the zero-row-sum invariant under the paper's stated numerical assumptions. That's what allows native block-scaled FP8 in every forward and backward attention matmul, with no architectural changes, no smaller global batches, and no auxiliary forward outputs.
The 569M trap
The scale dependence is the result to internalize. Small models and short runs conceal accumulated optimization error, and the paper's mitigation attempts make that explicit. QK normalization bounds the magnitudes feeding the QK product, NoPE removes the positional rotation from the equation, and lower learning rates during context extension shrink the steps the optimizer takes. All three mitigate or delay the degradation. None eliminates it.
| Mitigation | What it does | Outcome |
|---|---|---|
| QK normalization | bounds QK^T magnitudes before softmax | delays degradation, doesn't eliminate it |
| NoPE (no positional encoding) | removes rotation sensitivity in QK | delays degradation |
| Lower-LR context extension | smaller optimizer steps in long-context phase | delays degradation |
| Delta-Matching | restores the softmax gradient's zero-row-sum invariant | matches BF16 loss and downstream performance |
Any pipeline that validates FP8 attention on a sub-billion-parameter run and declares victory is only validating the concealment. The fix shows itself at production scale. If you're training at 1B and above, that's the result that should make you check what precision your attention backward actually runs in.
Common pitfalls
Treating K and V as the same kind of tensor. V gets contracted against attention probabilities; K gets rotated by RoPE in the QK product. The error structures are different, which is why WUSH-KV builds separate transforms. A shared transform at 2 bits leaves quality on the table.
Applying the key transform before RoPE. A data-adaptive transform placed ahead of the rotation scrambles the coordinate basis your calibration measured. The transform has to live after RoPE, where the rotation is already applied.
Validating FP8 training on a small model and calling it done. The stale-delta hybrid looks fine at 569M and falls apart at 1.67B and 5.29B. Even QK-norm, NoPE, and lower learning rates only postpone the problem. Validate at production scale before you sign off on a training pipeline.
Reaching for smaller global batches to stabilize FP8 attention. That's a workaround with real throughput cost. Delta-Matching's point is that you don't need smaller batches or auxiliary forward outputs; if you're tempted by either, fix the delta term instead.
Calibrating the KV transform once and forgetting it. WUSH-KV's transforms come from calibration data. Serve different traffic, longer contexts, or different languages, and the covariance structure drifts. Recalibrate on representative production traffic, the same way you would for any activation quantization scheme.
Error structure beats bit width
Low-bit quantization errors are structural, not numerical. Per-element rounding looks the same in any basis, but what changes between transforms is how the error correlates across a matrix product and across training steps, and that correlation is what determines downstream quality. KV quantization is a choice of coordinate system so the error doesn't destroy the softmax structure. FP8 training is a choice of numerical invariant so the error doesn't accumulate through the backward pass. Both papers converge on the same shape: fix the structure of the error, and extreme bit widths become affordable.
That's also why the two results land in production paths rather than staying on the research bench. WUSH-KV ships in SGLang, so serving stacks can adopt it without writing custom kernels. Delta-Matching promises released implementation, trained models, and data recipes, which turns a numerical proof into a reproducibility story. The interesting question for the next few months is whether the zero-row-sum fix holds up in production training pipelines at 10B+ scale, where the paper's own results suggest the accumulated error is worst.
Recommendations for practitioners
- If you're serving long-context traffic on GPU clusters, adopt 2-bit KV quantization with a data-adaptive transform and percentile clipping, because the per-sequence footprint drops from ~64 GiB to ~8 GiB at 128K context and decode bandwidth shrinks 8x. Recalibrate the transforms on your own production traffic before trusting the published perplexity numbers.
- If you're pretraining or continually training models above 1B parameters and you've kept attention in BF16 out of caution, adopt a Delta-Matching-style correction for the softmax backward pass, because the stale-delta alternative degrades loss and downstream tasks at 1.67B and 5.29B while the invariant-preserving fix matches BF16 without smaller batches or architecture changes.
- If you're memory-constrained on a single GPU and holding off on quantization until it's safe, start with 4-bit transform-based caching rather than raw 2-bit, because that's already a 4x footprint cut with much smaller quality risk, and move to 2-bit after you've validated the transform on your own workloads.