Appearance
The 16GB ceiling is moving
I spent last weekend pushing over a million tokens through Qwen 3.8 27B on a 16GB RTX 5060 Ti. The setup ran a 73,728-token context window, used native MTP speculative decoding, and built a complete REST API and MCP server for a legacy vBulletin forum with only three prompts. The whole machine cost less than a mid-range laptop.
Artificial Analysis scores the model at 52 on its Intelligence Index, against a median of 9 for comparable open-weight models. The community framing is blunter: GPT-5.6 Luna compressed into 27B. Either way, this is frontier-adjacent quality running on consumer silicon.
No single trick made that setup work. It worked because four separate layers of optimization finally lined up: a compact architecture, a sub-quadratic attention mechanism, a compressed KV cache, and a runtime tuned to squeeze every megabyte of VRAM. This is the efficiency stack, and it's the reason "frontier-class on consumer hardware" stopped being a meme.
The pieces are arriving at the same time. Neural Architecture Search papers are producing compact models. Attention is going sub-quadratic without the quality cliff. KV cache compression is getting attention-aware. And the local inference toolchain just hit its first semantic release. Each layer buys 2-5x. Together they multiply into the 50x that separates "can't run it" from "runs great."
Layer 1: architecture
NAS has always been the right idea with the wrong price tag. Discrete search over neuron and activation choices explodes combinatorially, and training candidate architectures is brutal. New work on neuron gating and mixed activation sidesteps this by relaxing those discrete decisions into continuous ones, then solving a bilevel optimization problem: the upper level optimizes architecture against validation performance, the lower level trains the weights. The three formulations, NAS-NG, NAS-MA, and NAS-NGMA, produce compact models that hold their own against much larger ones. NAS-NGMA hits 98.68% on MNIST with a 7.69M-parameter MLP, NAS-NG reaches 99.63% with a 0.26M-parameter CNN, and the family consistently outperforms vanilla DARTS on CIFAR-10.
The parameter counts matter. 0.26M parameters is about a megabyte of weights. That's not edge-device territory, that's microcontroller territory. The more practical use case is slimming down models you already have: NAS-NG can take an over-parameterized, literature-optimal architecture and improve accuracy while cutting parameters. You don't search from scratch, you compress what's already trained.
Mobius takes a different route. Instead of finding smaller architectures, it separates the two jobs a transformer block does. A globally shared FFN stores knowledge vectors, and multiple self-attention reasoners query it iteratively, using hidden states as both cache and carrier. The 7B version trained from scratch matches a standard 7B transformer while using 62.6% of the training data. The continually-pretrained Intern-S2-Mobius delivers nearly 4x end-to-end inference speedup over its dense baseline. 62.6% of the training data means the pretraining bill drops by more than a third. 4x inference speedup means the same quality at a fraction of the serving cost.
The community is already feeling this at the small end. Ling 3.0 Tiny packs 8B total parameters but only 1.3B active, and on a 4GB VRAM card it runs at 36 tokens per second. The same person reported Qwen 3.5 9B crawling at 5 tok/s on that hardware. Same memory budget, 7x the speed, comparable quality. Active parameters are becoming the spec that matters.
| Approach | What it optimizes | Result | What it means in practice |
|---|---|---|---|
| NAS-NG (neuron gating) | Neuron-level architecture | 99.63% MNIST with 0.26M CNN params | ~1MB of weights, microcontroller class |
| NAS-NGMA (gating + mixed activation) | Neuron and activation decisions | 98.68% MNIST with 7.69M MLP params | ~30MB, runs on a Raspberry Pi |
| Mobius (decoupled FFN/attention) | Knowledge vs reasoning separation | 4x inference speedup, 62.6% of training data | Pretraining cost cut by a third |
| Ling 3.0 Tiny (MoE) | Active vs total parameters | 36 tok/s on 4GB VRAM | Interactive speed on old hardware |
NAS is also expanding beyond accuracy. The SFLaaS framework folds carbon constraints and consumer participation into the search objective, transforming sustainability profiles into a feasible architecture region before federated execution. That matters if federated training is going to survive hard carbon limits rather than quietly dropping participants.
Layer 2: attention
Scaled dot-product attention is O(N²·d) in sequence length. Fine at 4K context, a wall at 128K. SSOG-Attention replaces the full similarity computation with a few learned Gaussian atoms per head, steered geometrically by the query token. Because the atoms factorize into a separable sum, complexity drops to O(N·√N·d).
The results are better than we're used to from sub-quadratic attention. SSOG beats SDPA outright on CIFAR-100, matches it on ImageNet, and converges much faster while using less memory as scale grows. The practical read: you no longer have to sacrifice quality to escape the quadratic wall. That's been the standard excuse for ignoring linear attention variants, and it's wearing thin.
Quick Take: No single technique gets you to 73k context on a 16GB card. The setups that work today stack 2-5x gains from architecture sparsity, sub-quadratic attention, compressed KV caches, and speculative decoding. Multiplied together, that's the 50x that separates "can't run it" from "runs great."
Layer 3: the KV cache
The KV cache is the memory bottleneck nobody sees. At 128K context, it can dwarf the weights. Standard quantization treats it uniformly: same bit width for every tensor, tuned to minimize reconstruction error in the cache itself. The AATC paper points out the flaw. Reconstruction error in the cache isn't what you care about. You care about how that error propagates through the attention softmax.
The paper proves that under a white-noise quantization model, the attention-aware distortion decomposes into additive key and value contributions that factor across tokens and channels. That decomposition makes bit allocation tractable: spend bits where attention actually depends on them, using transform coding and reverse water-filling, classical tools from rate-distortion theory. On Llama-3.1-8B-Instruct and Qwen-2.5-7B-Instruct, AATC stays near-lossless at roughly 5.8x compression across LongBench, RULER, GSM8K, MMLU-Pro, and MATH-500. Every baseline degrades in at least one setting.
5.8x compression is the difference between a 128K context needing 40GB of cache and needing under 7GB. That's the difference between cloud-only and local.
Key Numbers
- 5.8x KV cache compression with near-lossless accuracy across five benchmarks
- 73,728 tokens of context in 16GB VRAM using q4_1 KV cache quantization
- 4x end-to-end inference speedup from Intern-S2-Mobius
- 36 tok/s on 4GB VRAM with 1.3B active parameters
Layer 4: runtime
llama.cpp hit v0.1.0, its first semantic version after years of build numbers. It's a milestone that matters mostly as a signal: the local inference stack is stabilizing. The configs people run are getting surgical.
The 16GB Qwen 3.8 27B setup is a good example of how much the runtime layer absorbs. Q3_K_XL quant for the weights, q4_1 for the main KV cache, q5_1 for the MTP draft context, native speculative decoding with two draft tokens, and fit-target set to 128 MiB because the machine is headless. The result: 73k context in 16GB VRAM, with MTP pushing decode from 62 to 91 tok/s once warmed up.
Speculative decoding is the sleeper hit of this stack. A small draft model generates tokens, the big model verifies them in parallel, and you get 1.5x decode speed without quality loss. The Qwen 3.8 family builds this in natively with MTP, which is why community configs treat it as a default rather than an experiment.
The compiler layer matters too. PyTorch 2's TorchDynamo and TorchInductor showed a 2.27x inference speedup across 180+ models on an A100, which is what makes these architectures practical to deploy in the first place. And the TinyML survey makes the case that on-device inference is no longer a compromise for bandwidth-constrained use cases; it's often the only option that meets latency requirements.
What the community is actually seeing
I ran the Galaga test on Qwen 3.8 27B, and the difference from 3.6 was stark. The older model took 8 seconds of thinking and delivered a space invaders clone with none of the details. The new model at xHigh effort took 15 minutes of thinking and produced animated sprites, sound effects, an idle screen, and the ship capture mechanic from the original arcade game. It remembered details I never prompted it to include.
The catch is the time. 15 minutes of thinking for one code generation task is a lot. Medium effort cut that to 3 minutes and delivered about 90% of the result. Low effort matched the old model in 3 seconds. The reasoning budget is a dial, and most people are leaving it on the wrong setting.
The overthinking debate is real but misdirected. Yes, Qwen 3.8 27B burns far more reasoning tokens than 3.6 did. Artificial Analysis logged 160M output tokens during its evaluation, against a median of 43M for comparable models. But the other Chinese models in its class, GLM 5.3 and DeepSeek V4, run similar reasoning budgets. The tokens are doing work. If you're on slow hardware, hard-cap the reasoning budget at 8192 tokens and move on. That's a config line. Fix it and move on.
What the Community Is Saying: the loudest complaint in local LLM circles right now is about quant levels. Comparison posts routinely pit a 9B model against a 27B model and declare the small one the winner, then it turns out the 27B was running a 2-bit quant from an unknown converter. The community is petitioning for quant levels to be mandatory in posts. I found the same problem when testing: the gap between a good Q8 quant and a bad low-bit quant is larger than the gap between models.
Common pitfalls
- Comparing models at different quant levels. A Q3_K_XL 27B can lose to a Q8 9B on quality, and that says nothing about either model. Check the quant and the converter before trusting any comparison.
- Ignoring active parameters. Total parameter count tells you about quality. Active parameters tell you about speed. On memory-constrained hardware, a 1.3B-active MoE will beat a dense 9B at the same VRAM budget by 5-7x in throughput.
- Forgetting MTP draft KV caches double VRAM allocation. The draft model has its own KV cache. On a 16GB card, that's the difference between fitting 73k context and OOMing at 40k. The configs that work set q5_1 for the draft cache, q4_1 for the main cache, and keep fit-target at 128-256 MiB.
- Treating reasoning tokens as pure overhead. They're not. The Galaga test showed the difference between 8 seconds and 15 minutes of thinking is the difference between a generic clone and a faithful recreation. If you need speed, lower the reasoning budget instead of disabling it.
- Using uniform KV cache quantization when attention-aware allocation exists. Uniform quantization minimizes the wrong objective. Attention-aware allocation spends bits where error actually propagates through the softmax, and gets near-lossless results at 5.8x compression where uniform methods degrade.
One thing to remember
The efficiency stack compounds. Architecture search, sub-quadratic attention, KV cache compression, and runtime tuning each buy a few times in isolation, but they multiply together. The 16GB card running 73k context today would have needed a server rack two years ago.
What to run today
If you're on a 16GB card doing agentic coding, run Qwen 3.8 27B with a Q3_K_XL or Q4 quant, q4_1 KV cache, MTP speculative decoding enabled, and a reasoning budget capped at 5000-8192 tokens. That config gives you near-frontier code quality at interactive speeds.
If you're on 4GB VRAM or less, skip dense models entirely. Use an MoE with 1-2B active parameters like Ling 3.0 Tiny. You'll get 36 tok/s instead of 5, and the quality gap to a dense 9B is smaller than the speed gap is large.
If you're building long-context services, watch attention-aware KV cache compression and sub-quadratic attention. Both are leaving the paper stage. Expect 5-10x memory reduction and O(N·√N) attention to be standard deployment options within a year.