Appearance
The benchmark that split the subreddit
I have an R9700 GPU, an RTX 5060 Ti, and 64GB of DDR5. Nothing special by 2026 standards. My daily driver is Qwen3.8-27B at Q6 running on the AMD card, around 35 t/s. Comfortable for chat, fine for small coding tasks.
Then a new runtime showed up and half the local-AI subreddit wouldn't shut up about it. The claim: big MoE models that crawl on general runtimes run several times faster on an engine built specifically for them. I tested it. Same machine, same Qwen3.8-Flash-Next weights at IQ3_XXS: llama.cpp gives 21 t/s. The new engine gives 60 t/s.
That's the difference between reading output token by token and holding an actual conversation.
The engine is Strata, the loudest example of a growing category: inference runtimes that deliberately support a tiny slice of models and squeeze them hard. Community reaction is a mix of genuine enthusiasm and visible fatigue. One of the top posts this week is a Konosuba meme captioned "yes bots we get it, Strata is good now please stop." Both reactions are fair. The speedup is real. The hype is real too.
What an overfit engine gives up
A blog post making the rounds this week named the category: Strata, ninfer, DwarfStar, Splash, llamAmpere, gufo. Different projects, same bet. They skip what llama.cpp and vLLM do best, which is running almost anything with a safetensors file, and instead hard-code around a small set of architectures, sometimes down to one hardware family. Strix Halo APUs get their own engines because 128GB of unified memory changes which offloading tricks make sense.
What you get in exchange: the engine knows the model's expert count, routing patterns, KV cache shape, and the hardware's memory topology. It can pin the KV cache to VRAM, keep hot experts resident, and page the rest from NVMe with a prefetch policy tuned to that model's traffic. Strata's user-facing flags tell the story: --mmap-experts, --resident-cpu-experts, --expert-cache auto. A general runtime can't assume any of that. Its job is to work everywhere.
Every step in that chain is tunable when you only support one model. General runtimes implement the average case. Overfit engines implement the exact case.
The numbers on the board
Here's what the current field looks like, pulled from this week's threads:
| Hardware | Model and quant | Engine | Generation speed |
|---|---|---|---|
| R9700 + 5060 Ti, 64GB DDR5 | Qwen3.8-27B Q6 | llama.cpp | 35 t/s |
| Same box | Qwen3.8-Flash-Next IQ3_XXS | llama.cpp | 21 t/s |
| Same box | Qwen3.8-Flash-Next IQ3_XXS | Strata | 60 t/s |
| 16x 3090, 100GbE, 6kW | Qwen 397B INT4 | vLLM/SGLang | 50-60 t/s |
| 4x DGX Spark, 400W | Qwen 397B INT4 | vLLM/SGLang | 30 t/s |
| 16x DGX Spark | Kimi K3 (2.8T), 300k context | vLLM/SGLang | 20 t/s |
| 2x SQRL FK33, 75MHz | Qwen3.5-9B INT4 | custom VHDL | ~3 t/s |
60 t/s is conversational. You don't perceive a typing delay. 21 t/s is readable, but you can feel the model working. 8 t/s, where the Threadripper-with-512GB-RAM crowd lives on 671B models, is roughly reading-aloud pace: fine for real tasks if you learn to batch your questions. 3 t/s is barely interactive. It's also running on $280 of ex-mining FPGA hardware, which puts it in a different category entirely.
The same 397B model that needed a 6kW rack to hit 50-60 t/s runs at 30 t/s on four DGX Sparks drawing 400W. That one comparison is the story of local inference in 2026. At 20 t/s with a 300k window, you can drop a multi-repo codebase into context and still keep a conversation going.
Key numbers: nearly 3x the tokens per second from a runtime swap with identical weights. 400W versus 6kW to run the same 397B model. A quant shift from roughly 93GB to 70GB, plus an engine swap, took Flash-Next from 15 to 60 t/s on a 64GB box. 20 t/s at 300k context is the home-cluster ceiling for the 2.8-trillion-parameter Kimi K3.
Quick Take: for the models that get dedicated engines, your choice of runtime now matters more than your choice of GPU.
What breaks when you squeeze
My 64GB machine hit 60 t/s on Flash-Next, and it also hit 96% system RAM usage with the model loaded. ComfyUI on the 5060 Ti needs system memory too, and video models eat a lot of it. For a while it was pick one workload.
Workarounds exist, and they're the kind of flags you only find deep in forum threads. Strata side: --mmap-experts, --resident-cpu-experts, --expert-cache auto. ComfyUI side: --fast-disk. Together they got both workloads running at about 90% RAM. The catch shows up in storage I/O: when the video model loads its weights, NVMe reads spike to 1GB/s and Flash-Next inference drops to the 25-40 t/s range. Once loading settles, reads sit around 30MB/s and the LLM climbs back to 50-60 t/s.
A different kind of failure came from the other side of the hype. I asked the miracle engine's IQ3_S run to give an opinion on a subreddit thread and measure throughput on a long generation. The thinking trace started fine. Ten thousand tokens later it was looping on "Need maybe maybe include 'Use for code, list constraints'." The same weights in llama.cpp produced a coherent answer. No doom loop.
That one stuck with me. The pitch for these engines is that they make big models usable on modest hardware, and this engine spiraled on a trivial task. The cause isn't visible from the outside: a sampler difference, a template bug, a quantization interaction, or something in how the engine handles thinking traces. Whatever it is, tokens per second tells you throughput. It doesn't tell you whether the output is sane.
The subreddit's fatigue is understandable. The Konosuba meme post nails the pattern: every thread eventually lands on the same recommendation, and the counter-reaction has become its own subgenre. Nobody disputes the benchmark numbers. What people want is proof that the engine holds up on real workloads, not just long generations where nobody reads the output.
The hardware frontier: fuses, Sparks, and FPGAs
While most people fight over quant sizes, a smaller group keeps pushing what home means. Someone posted their full journey this week, and it reads like a compressed history of local inference: a 3090 for LLaMA 33B, a second 3090 for 65B, a Threadripper with 512GB of DDR4 running DeepSeek 671B at 8 t/s with experts in RAM, then 16x 3090s across P620 nodes on 100GbE. That cluster ran Qwen 397B beautifully, right up until the house fuse blew because the electric oven was on at the same time.
The pivot came with the GB10. Four linked Sparks ran the same 397B model at 30 t/s while drawing 400W instead of 6kW. He bought four more, talked his brother into an 8x setup five minutes away, and joined the clusters. Kimi K3, the 2.8-trillion-parameter open-weight release, runs on the combined 16x at a usable 20 t/s with a 300k context. That took iteration: the first attempt was an unusable 7 t/s at 100k context, and he rebuilt vLLM and SGLang images until it worked. Four more Sparks are on order now, a dedicated always-on node running a smaller model (GLM 5.3 Flash) 24/7 so the big cluster isn't wasted on daily chores.
Notice the pattern: he keeps the smaller cluster running all the time while the big one switches between GLM 5.3 and Kimi K3 depending on the task. That's the same split my box fell into with Flash-Next and ComfyUI. Sizing the stack to the workload is becoming the standard home setup.
The other extreme is the FPGA build. One developer posted a complete Qwen3.5 inference implementation in VHDL, running a 9B INT4 model on two SQRL FK33 boards bought for $280 each. Those are ex-mining FPGAs with 8GB of HBM2 apiece. At 75MHz it generates about 3 t/s, dropping to 2.4 t/s past the 2-3k context mark. He verified outputs layer by layer against llama.cpp, which is the right way to ship a custom inference stack. His projections for the 27B on a two-die board at 200MHz look solid: around 16 t/s prefill, 8 t/s generation, or 15 t/s with tensor parallelism across the dies. Two dies run out of room around 45k context because the 27B's KV cache no longer fits beside 14.5GB of weights. Four dies are needed for the full 262k window. His ASIC estimate for the same design on N3 at 2GHz: 294 t/s and 125-340W, which would beat most hosted APIs.
FPGAs won't become the mainstream path. But that project makes the same point as Strata, at a different price point: a stack designed for exactly one model beats a general one for that model.
The model side is overfitting too
The same specialization is happening one layer up. PewDiePie's Ajax launched this week as part of Odysseus, his self-hosted AI workspace. Ajax is a fine-tuned Qwen3.5-9B meant to run as an always-on agent on a home PC: search, browsing, email, calendar, all processed locally. The model page is upfront that refusal behavior was removed for a freer experience on your own hardware. The setup sits in the same space as OpenClaw and Nous Research's Hermes Agent.
The method is abliteration, automated with an open-source tool called Heretic. It finds the directions in weight space that produce refusals and removes them, ideally without collateral damage. The project's docs warn against removing too much, and the warning is earned: over-abliterated models lose general judgment, not just politeness. The line was drawn at content that harms people, and a legal review kept the model away from actionable dangerous instructions.
Then there's the OpenAI ban drama. His account was deactivated twice, one ban explicitly citing distillation: using another model's outputs to train or improve your own. He wanted to distill "just a little bit" to improve the fine-tune, and later ran the model again to create seed data. Community reaction splits along familiar lines. Some see distillation as unwanted customer behavior, no different from scraping against an API's terms. Others expect a court to eventually call a distilled model a derivative work, which would land hard on a lot of open-weight development. The open-model camp points out that frontier labs trained on the open web themselves, and that "they worked hard stealing all of our data to make their AI machines."
Ajax is the model-side version of Strata: small, shaped for one harness, aimed at hardware you own. The creator's argument is blunt, and it's the same one the runtime crowd makes: why poke a trillion-parameter beast for basic tasks when a 9B tuned for the job is cheaper, faster, and private.
I set up a local voice agent this week and the experience lines up with the same trend. Breeze for speech, Opus 5.5 doing the thinking, a Live2D avatar for company. The first sound lands about 500ms after I stop talking, roughly the delay of a phone call; about 1-1.5s with thinking mode on. It watches for background agent sessions to finish, summarizes them out loud, and asks for decisions. Past 500k tokens of context it still remembers we're in a live session and keeps messages short. That used to be the hard part. None of this needed a datacenter.
Common Pitfalls
The first trap is quant choice. I almost dismissed Strata entirely because an earlier attempt ran at 15 t/s. That attempt used IQ4_XS at roughly 93GB. The quant everyone benchmarks with, IQ3_XXS, sits around 70GB and changes the whole memory picture. Hold the quant fixed before you assign credit to any engine, because the weights determine the offload strategy.
The second trap is trusting t/s. The doom loop produced plenty of tokens, and all of them were useless. Before switching your daily driver, run your actual tasks on the new engine and compare outputs against a known-good llama.cpp run. A ten-minute A/B test beats a week of regret.
Then there's shared RAM. At 96% usage, your second workload fails before it starts. Plan for the whole machine, not just the model. The flags that got both workloads running concurrently were --mmap-experts, --resident-cpu-experts, and --expert-cache auto on the inference side, plus --fast-disk on the ComfyUI side. And NVMe is the hidden variable: another process hammering the disk can halve your tokens.
If you go the abliteration route, expect to re-benchmark after. Over-abliterated models lose subtle reasoning before you notice it, and you only spot the damage on tasks that require careful judgment.
Clock speed doesn't scale for free either. On the FPGA side, pushing past 75MHz needs RTL optimization and higher core voltage. The design ceiling hits before the clock ceiling. And some ex-mining boards have no fast path for loading weights without soldering, which turns every iteration into a hardware project.
One Thing to Remember
The local inference stack is fragmenting on purpose. Runtimes overfit to specific models, models get fine-tuned for specific harnesses, and hardware clusters get built around specific weight sizes. The compatibility cost is real