Skip to content

Cut Tokens, Not Just Parameters: The Local Inference Stack in 2026

#quantization #model-compression #local-inference #kernel-compilers #llm-optimization

Someone on r/LocalLLaMA did the electricity math on their home rig: a 7900 XTX and a 9800X3D running Qwen Flash Next at 0.25€/kWh. Total came to 0.12€ per hour of active inference. Their question was direct: if frontier APIs are faster and cheaper per token, why run local at all?

The honest answer is that the gap between "model exists" and "model is cheap to run" is closing from several directions that rarely get discussed together. Compression that learns where to cut, quantization that accounts for how reasoning models think, small routers that keep big models quiet, and kernel compilers that squeeze the hardware. All four moved in the last few weeks, and each changes the local-versus-API calculation in a different way.

The cheapest parameter is the one you never load ​

Low-rank weight factorization is attractive because it keeps matrices dense. Dense means standard kernels, happy hardware, and you still cut memory and compute. The hard part has always been choosing which subspace to remove from each weight matrix. Existing methods use local closed-form criteria: activation energy, layer-wise reconstruction error, or a quadratic approximation of the loss. Those criteria look at one matrix in isolation and ignore how errors propagate through the network. At high compression, errors compound with depth and performance collapses.

Learnable Subspace Projections (LSP) learns the discarded subspaces end-to-end instead. Each linear layer, or tied group of layers that read the same activations, gets an orthogonal projector. All projectors train jointly against a global objective, either KL divergence to the dense model's output distribution or the original training loss, while the pretrained weights stay frozen. Projectors start from a whitened SVD truncation, and ranks are allocated by the output KL each projector induces per parameter saved.

The high-compression numbers are the proof. On Llama-2-7B at -70% compression, LSP lands at 10.9 WikiText-2 perplexity and 42.2% mean zero-shot accuracy, against 13.3 and 36.0% for the strongest baseline. A gap from 13.3 down to 10.9 is the difference between a model that rambles and one that stays coherent. A 6.2-point accuracy gap is the difference between a toy and a tool.

LSP also gives attention a trick untied factorizations can't use. Because tied groups share one factor, the model caches one narrow latent in place of full keys and values. At a 128k-token context, combined weights plus KV cache shrink 13.5x, versus at most 6.5x for untied baselines. That's the difference between a 128k context that fits on a laptop and one that needs a server. The factorized model also decodes up to 1.6x faster than dense at small batch sizes, which is exactly the regime single-user local inference lives in.

13.5x reduction in combined weights plus KV cache at 128k context, thanks to LSP's shared latent 1.6x decode speedup over the dense model at small batch sizes -70% compression, where LSP keeps Llama-2-7B at 10.9 perplexity

Quantized reasoning models think too much ​

Post-training quantization reads like a free lunch: drop to 4-bit, keep accuracy, cut memory. On reasoning models, the lunch has a hidden cost. PTQ doesn't just degrade reasoning performance, it amplifies overthinking. The model generates longer chain-of-thought trajectories, and those extra tokens eat the savings from lower precision. A quantized model that reasons 50% longer for the same answer is not cheaper.

RATIO attacks this in two steps. QRBA (Quantization-aware Reasoning Behavior Analysis) identifies overthinking tokens by comparing what the full-precision and quantized models do at each position. TSPD (Token-Specific Penalty Determination) then assigns each identified token a tailored penalty, using full-precision guidance and no additional training. The results: up to 9.8 points accuracy improvement and up to 51.3% shorter chain-of-thought compared to quantized baselines.

Quick Take: On quantized reasoning models, token count is an efficiency metric, and RATIO's 51.3% chain-of-thought cut at +9.8 accuracy points shows the quantization was never the whole problem.

The same blind spot shows up in context compression. CAFD, a ground-truth-free self-distillation framework for OmniLLMs, adapts a compressed-context student using a full-context teacher's soft targets, with no reference answers or rewards. On Qwen2.5-Omni-7B across five audio-video benchmarks, five compression pipelines, and five deployment budgets, it improved 120 of 125 conditions with an average accuracy gain of 1.44 points. The pattern repeats: compress the representation, then adapt the model to live inside the compressed representation. Compression without adaptation leaves accuracy on the table.

The backward pass reads roundings too ​

Low-precision training rounds tensors, and the backward pass reads those tensors again, often more than once for different gradients. Each read can take the forward's rounded value, the original full-precision value, or a fresh random rounding. This backward-state policy looks like a memory detail you settle once and forget. The paper "Backward-State Policy Is Part of the Learning Algorithm" argues it's part of the learning algorithm itself, and neither of the usual checks catches it.

In three pairs of 390M-parameter runs with an emulated FP8 backward, training fails when attention's backward reuses the forward's rounded output and succeeds with a fresh rounding from the same distribution. Copy accuracy doesn't predict that outcome. Final loss doesn't catch the other failure mode: reading the original for every use persists in trained models without surfacing in planned loss comparisons.

The concrete trap: a normalization output stored in low precision feeds two gradients. The gain's gradient needs the original value. The next layer's weight gradient needs the rounded value that the layer actually multiplied. No single stored value serves both. If you write a low-precision training loop, derive which value each use must read, operator by operator, before launching a full run. In tests with PyTorch and Transformer Engine, the per-use requirements predicted exactly when reuse changed what the backward computes on average.

Small models in front of big ones ​

Most decisions an agent makes in front of a local model aren't open-ended. They fall into a handful of recurring kinds: is this a prompt injection, which tool do I call, how urgent is this ticket, is this answer grounded in the sources. Jeff-Qwen3.5-0.8B is a small "System 1" model that returns calibrated probabilities over options you define in a single forward pass. The author then trained nine LoRA adapters, one per job, about 40 MB each, loaded alongside the untouched base model.

The headline result, measured on the same test rows:

MeasureQwen 27B aloneJeff + adapters, 27B only when unsure
Mean accuracy (8 adapters)86.6%95.3%
Time per decision8.1 s0.25 s
Wrong answers13.4%4.7%
Memory28.6 GBUnder 2 GB

Roughly 38x faster decisions at higher accuracy, with the big model consulted only on the rows the small one is unsure about. The caveats are up front: the 27B ran 8-bit with step-by-step reasoning off, and with reasoning on, the speedup would be larger. Jeff wins outright on 8 of the 9 adapters and ties on grounding (96.3% vs 96.7%) at 20x the speed.

The defer threshold is the entire game. It must be calibrated on real traffic, or the router either over-escalates to the big model and you've bought nothing, or it silently answers questions it should have routed onward and your accuracy quietly erodes.

The complement to routing is cutting the model itself. Victoria prunes Qwen3.8-Flash-Next down by 44% of its experts (512 to 288 per layer) using the REAP technique, then retrains at NVFP4. Training in the format it ships in beat the previous train-dense-quantize-after build by 7.5 points on Terminal-Bench 2.1 (70.0% vs 62.5%). The draft head pushes single-stream decode from 135 to 280 tok/s on a B300. If your deployment format is fixed, train in it.

The kernel layer decides what the hardware actually does ​

A perfectly compressed model still underperforms when the kernels are generic. TileLang, a Pythonic DSL on top of TVM, targets CUDA, ROCm, Metal, Ascend, and LLVM CPU from one language. Its recent changelog reads like a checklist of efficiency work: MXFP4 block-scaled GEMM on AMD gfx950, NVF4 block-scaled MMA on Blackwell SM120, a ~1.9x win on DeepSeek V3.2's sparse-attention top-k selector, FlashAttention on SM100, cooperative-tensor GEMM on Apple M5, and an Ascend 950 backend. The formats new hardware wants, NVF4, MXFP4, block-scaled everything, don't have cuBLAS entries. You write the kernel or you leave 1.5-2x on the table.

The llama.cpp side has its own efficiency push. The Qwen4Exp multi-token prediction (MTP) PR merged, bringing speculative draft heads to the local stack. When I tested it, the q8 MTP head OOM'd at 256k context. Loading the Q4 MTP head and dropping ubatch to 512 got it in, and short-context generation went from about 40 to 50 tokens per second. The draft head doesn't share the token embedding when loaded as a separate GGUF, so requantization is often needed. And a custom GGUF with a draft head that mainline llama.cpp doesn't know about yet throws "expected 1256, got 1224". The download is fine, the binary is wrong. Use the fork.

Hardware reality check: from 192GB desktops to a Tandy 286 ​

Local stacks run on an absurd range of hardware. Framework opened preorders on a desktop with the AMD Ryzen AI Max 400 series and 192GB of unified memory. One user finished a four-card Tesla T4 build, 64GB total, in a Fractal Torrent and immediately wanted more cards. Another fed eBay listings to Claude and charted the cheapest 32GB VRAM GPUs under $1600.

The 0.12€/hour question resolves against this backdrop. Electricity is the wrong axis. A 27B model on a 7900 XTX is cheap per hour, but hardware cost only amortizes if you actually use the machine. The API wins when your usage is low and sporadic; local wins when data can't leave the machine, when latency has to hit a hard ceiling, or when per-token pricing compounds. Across these threads, the community verdict lines up: people run local because it's private, because it's fast, or because it's fun. The electricity math rarely justifies it alone.

The fun end is extreme. Someone runs full chat and image generation on a 286 Tandy: DeskMind, a native DOS program, talks over WiFi (PicoMEM 2 with mTCP) to a Python server driving Qwen3.8-27B and Krea 2. The 286 never sees JSON, base64, or a PNG, just plain text lines and pictures ready for video memory. "Draw me a 286 AI logo" takes about 9 seconds. The B.L.O.O.M. project sits at the opposite end: an Arduino UNO Q cyberdeck in a $10 vintage train case running a local Gemma, classifying birdsong with Edge Impulse, and writing Grinnell-style field notes to SQLite. Nothing leaves the device until you press sync, which posts to a Netlify function that opens a GitHub pull request. I hit the same trap the B.L.O.O.M. author did, Chromium popups opening behind windows on Debian, and the fix is the same: launch Chromium in kiosk mode. Local deployment finds new failure modes at every hardware tier.

Common pitfalls ​

Five things trip people up in this cluster, and I've hit versions of all of them.

  1. Measuring accuracy but not thinking. PTQ on a reasoning model can hold accuracy while overthinking pushes CoT 50% longer. The quantized model is then more expensive per correct answer than the full-precision one. Measure token count per solved problem, not just pass rate.

  2. Assuming local means private. The B.L.O.O.M. author made "nothing leaves the device unless you press Sync to Cloud" a hard requirement because it doesn't happen by default. App Lab bricks, Edge Impulse models, and so-called offline runtimes can still call out. Audit network calls before you trust the word offline.

  3. Treating backward-state policy as a memory detail. The backward pass reads rounded tensors multiple times, and which value each read gets changes whether training converges. The normal checks, copy accuracy and final loss, don't catch the wrong policy. Verify per-use requirements on individual operators.

  4. Shipping an untuned defer threshold. The router is only as good as its calibration. Set the threshold on held-out rows and it will over-escalate or under-defer on real traffic. Calibrate on real traffic, and re-check after any input distribution shift.

  5. Quantizing after training instead of training in the reduced format. Victoria's NVFP4 retrain beat the quantize-after build by 7.5 Terminal-Bench points. If you know the deployment format, bake it into training. Use post-hoc quantization only as a fallback.

One thing to remember ​

Every win in this cluster is a version of the same move: stop generating tokens that don't need to exist. LSP removes parameters that don't matter. RATIO stops a quantized model from thinking in circles. Jeff keeps the 27B from answering at all when a 0.8B is calibrated enough. The draft head predicts several tokens at once. Pick the layer that wastes the most tokens in your workload and fix that one first.

The bottom line ​

  • If you're serving a quantized reasoning model and your effective cost per solved problem is up, apply a token-level overthinking intervention like RATIO before buying more GPUs. Cutting CoT by half at equal accuracy is the cheapest capacity you'll ever add.
  • If you're building an agent that makes a handful of recurring decisions, put a calibrated small router with task adapters in front of your big model. The evidence here: 95%+ accuracy, 38x latency, under 2GB of memory. Just calibrate the defer threshold on real traffic.
  • If you're shipping a fixed kernel shape, dequant GEMM or FlashAttention on Blackwell, RDNA4, or Ascend, learn TileLang. Generic kernels are leaving 1.5-2x on the table, and speculative decoding via MTP heads is landing in llama.cpp now, so expect draft heads on every local stack within months.