Appearance
The local inference stack moved on every layer at once
A 12GB laptop GPU runs a frontier-class model from six months ago at 51 tokens per second of generation and 1500 tokens per second of prefill. The machine is a 5070 Ti laptop with 64GB of DDR5. The model is a heavily quantized Qwen flash variant. The engine is called Strata, and most people haven't heard of it.
That's the news this week in local LLMs. The stack got faster on every layer at once: quantization quality, inference engines, browser runtimes, memory overclocking, server silicon. No single one of those is a dramatic jump on its own. Together they multiply. The r/LocalLLaMA threads reporting these numbers ended on a giddy note about the era we're in, and honestly, it was earned.
Here's where each piece sits:
Strata vs. llama.cpp: what engine-level tuning buys you
When I tested Strata against every llama.cpp fork I could find on my 5070 Ti laptop, nothing came close. Same quant, same context, roughly double the generation speed. Stock llama.cpp pushed 23 t/s at 43k context depth; Strata held 51 t/s. Generation at that rate is a sentence per second, smooth enough for real conversation. 23 t/s is usable but laggy in comparison.
Prefill was the bigger shock. Reading a 32k-token context, the same quant went from about 100 t/s on stock llama.cpp to 1500 t/s on Strata. In practical terms, the stock engine needed over five minutes to ingest that history; Strata does it in 21 seconds. A full-context read stops being a stall and becomes a normal step.
Key numbers: this week's local inference reports, straight from the r/LocalLLaMA threads: 51 t/s generation and 1500 t/s prefill on a 12GB laptop GPU (Strata + ISTA-DASLab IQ3_XXS) 23 t/s generation and ~100 t/s prefill with stock llama.cpp on the same quant 75 t/s generation for a 26B-class model running across two 8GB 4060s 200+ WebGPU kernels open-sourced by Hugging Face 1248 GB/s on a 5070 Ti after uncapping GDDR7 clocks (+40%)
Strata is a tuned instrument for exactly one model family, Qwen3.8 Flash Next, and it only accepts the ISTA-DASLab GGUFs built for it. It's also NVIDIA-only for now, with AMD support still experimental. That specialization is why it wins. llama.cpp has to serve hundreds of architectures; Strata serves one. The developer shipped a rough first build with KV cache bugs and CPU throttling, and the community noted the fixes landed fast. Living software, moving quickly.
Because llama.cpp isn't slow. The thank-you thread for Mradermacher's IQ quants had someone running a 26B-class Gemma at 75 t/s generation and 1500 t/s prefill across two 8GB 4060s, served through LM Studio. Faster than most people read, on GPUs that were budget cards two years ago. The gap between "stock but good quants" and "specialized engine plus good quants" is the real story.
Quant quality is the silent multiplier
The engine gets the headlines, but the quant does more of the heavy lifting. Mradermacher's IQ quants have a reputation in r/LocalLLaMA as the best quality-per-bit you can download, and the thank-you thread this week is the norm. ISTA-DASLab's GSQ-RCO line does the same for the Qwen flash models: IQ3_XXS, the 3-bit tier, is good enough for daily chat, and their IQ3_S claims to recover the full model's coding benchmark performance. That's a wild claim for a quant that small. Test it on your own prompts and judge for yourself.
The consequence: you stop asking whether a model fits and start asking whether the compression shows. With the better IQ quants, usually it doesn't.
Quick Take: quants, engines, kernels, and drivers improved simultaneously, and their compounding effect is why a 12GB laptop now runs what needed a data center six months ago.
The browser became an inference runtime
Hugging Face open-sourced its WebGPU kernels collection this week: 200+ kernels for common ML operations that run entirely in the browser. No server, no CUDA, no install. The team is upstreaming the work into Transformers.js, ONNX Runtime Web, and LiteRT.js. For a model builder, that means the inference runtime you ship can be a URL. Users get the model weights and a WebGPU kernel set on page load, and their own GPU does the work. Privacy wins, deployment cost dies.
The ceiling is real: this is the 1-8B tier, not 320B hybrids. But the 1-8B tier is the one most consumer products actually need, and it now runs on integrated GPUs in a browser tab. The kernels page filters by platform for a reason, so check kernel support for the browsers you ship before you commit.
Memory bandwidth is the actual bottleneck
Token generation is a memory problem. Every generated token streams the model's entire weight set out of memory. Bytes per second decide your decode rate. FLOPS barely enter it. That's why the GDDR7 overclocking news matters as much as the new engines. A tool called mlock, which removes the GDDR7 clock ceiling, showed the hardware was leaving a lot on the table: a 5070 Ti measured 1248 GB/s, up from roughly 896 GB/s stock. A 40% bandwidth gain, and decode speed follows it almost linearly.
Same story on the integrated side: AMD's Linux 7.4 work boosts Radeon iGPU LLM performance by 18-23%, per Phoronix benchmarks. Driver-level kernel efficiency is another engine, just compiled into the graphics stack. When you're bandwidth-bound, any improvement to how the GPU moves weights shows up directly in tokens.
Server silicon went absurd in parallel
AMD's EPYC 9006 "Venice" price sheet landed with the expected absurdity: a 256-core flagship and a pricing structure where the biggest chip is the best value per core. The 16-channel DDR5-12800 memory subsystem is the part that matters for inference, roughly 1.6 TB/s of bandwidth, about 91% of what an RTX 5090 moves at 1.79 TB/s. In practice that's enough to stream a 70B model in 8-bit through memory about twenty times per second. Big-DDR5 servers stay the budget route to long-context local inference.
Here's a sample of the lineup:
| Model | Cores / Threads | 1KU price | Price per core | TDP | L3 cache |
|---|---|---|---|---|---|
| EPYC 9996 | 256 / 512 | $14,904 | $58.22 | 600W | 1024 MB |
| EPYC 9966 | 192 / 384 | $14,079 | $73.33 | 600W | 768 MB |
| EPYC 9756 | 128 / 256 | $12,498 | $97.64 | 500W | 512 MB |
| EPYC 9556 | 64 / 128 | $8,008 | $125.12 | 300W | 384 MB |
| EPYC 9016 | 8 / 16 | $700 | $87.50 | 130W | 48 MB |
The 64-core chip costs more per core than the 256-core flagship, an inversion that rewards buying up. Dual-socket boards take it to 512 cores, though one commenter pointed out Windows 11 still caps at 128 cores by default, so that machine is Linux territory. And the 16-core 9176F packs 192MB of L3 for per-core licensing savings, a reminder that AMD now prices silicon for lawyers as much as engineers.
One training-side data point: DeepSeek announced it's training on Ascend 950, two years after Liang Wenfeng argued that someone had to push onto the frontier. When leading labs run on domestic silicon, the consumer GPU supply picture calms down, and that trickles down to what you can buy for local inference.
llama.cpp absorbed a 320B hybrid this week
The llama.cpp PR that added GLM-5.3-Flash support is a good window into how the project absorbs an architecture it hasn't seen before. The model is a 320B hybrid: 34 KDA linear layers mixed with 11 DSA layers, MLA attention, DeepSeek-style MoE, plus vision. The hard parts were the hybrid recurrent state and DSA cache, plus a new vision encoder that shares a family resemblance with glm4v but needs its own preprocessing.
What the community is saying: testing the pre-release on a 4090 with 128GB of DDR5, an early adopter saw roughly 300 t/s prefill at 256K context, with generation starting near 9 t/s and sagging to about 6 t/s by mid-window. The model stays coherent the whole time, but that drop is the "pooled indexer keys" behavior, and it only surfaces when someone runs real long-context workloads. The same tester flagged a memory quirk: checkpoint files grow alarmingly as context fills, nearly 1.6GB by 90K tokens. And the review thread showed the integration friction: tensor renames orphaning GGUF files uploaded days earlier, vision projector naming conflicts, and a Metal fusion baseline that someone had to regenerate on an M5 Max. One maintainer joked about still teaching their AI pair-programmer canonical llama.cpp architecture.
The takeaway: open-source engines absorb frontier architectures within weeks. The long tail is the integration work, quant compatibility, vision towers, Metal baselines, and the context-length behaviors that only show up in the field.
Common Pitfalls
- Don't feed Strata a random GGUF. It only accepts select ISTA-DASLab builds for Qwen3.8 Flash Next, and it's NVIDIA-only for now. Pull the engine's supported quant list before you clone the repo.
- Don't trust short-context benchmarks for long-context work. GLM-5.3-Flash decode starts near 9 t/s and sags to about 6 t/s by 128K context. Measure at the context depth you'll actually run, not at 1K where everything looks fast.
- Don't overclock the core clock expecting faster tokens. Decode is memory-bound. The mlock gains came from GDDR7 memory clocks, a 40% bandwidth jump and a near-linear token rate increase. Core clock moves a few percent.
- Don't assume WebGPU kernels are portable. A kernel tuned for Chrome on NVIDIA can be slow or absent in Safari's Metal path. The Hugging Face kernels page has per-platform filters; use them, and test every browser you ship.
- Pin your engine version after downloading quants. The GLM-5.3 merge renamed tensors and orphaned GGUFs from days earlier. "My quant stopped loading" is usually a version mismatch, not a corrupted download.
One thing to remember
Token generation is a bandwidth problem. Every generated token streams the model's weights from memory, so decode speed tracks how many bytes per second your hardware can deliver. Judge each new benchmark by the bandwidth it implies, then ask whether your machine actually has it.
The Bottom Line
- If you're a llama.cpp user running a flash-class Qwen, try the matching ISTA-DASLab GGUF with Strata before you buy new hardware: the same quant delivered roughly 2x generation and up to 15x prefill, but confirm your GPU vendor and the engine's supported quant list first.
- If you're capped at 8-12GB of VRAM, your daily-driver speed is decided by the quant, not the engine: the Mradermacher and ISTA-DASLab IQ3-line is the current quality-per-bit standard, and it's the difference between a demo and a daily driver.
- If you're shipping an on-device product, target the 1-8B tier with WebGPU kernels and Transformers.js rather than bundling a runtime; the browser path ships today, and the 8-30B tier should become browser-runnable within the next year as engine specialization and driver work keep stacking.