Appearance
The problem: inference is the new bottleneck
Inference used to be the cheap part of the ML lifecycle. You trained once, then paid a small per-query toll. That stopped being true around the time context windows hit 100K tokens and VLMs started ingesting video.
The cost shows up differently depending on the workload. Long-context models re-read entire documents on every generation step. Embodied VLMs search thousands of video frames for each question they're asked. Edge devices can't fit the weights, let alone the activations. And retraining a smaller model is often not an option: the weights are frozen, the training run was too expensive, or the data is gone.
The papers in this cluster all start from that constraint. They attack efficiency from six levers: inference-time recurrence, modular memory, key-frame selection, better pruning math, wavelet bases, and activation function choices. The surprising part is how many of them work on frozen models with zero retraining.
Recurrence without retraining: recirculation
Feedforward transformers have a hard limit: state updates are bounded by model depth. A 30-layer model gets 30 chances to update its belief state, and that's it. Recirculation treats that as an architectural bug and adds a specific form of recurrence at inference time.
The mechanism is simple in concept. The model's output feeds back as input for additional passes, letting it act as a dynamical system that tracks belief states. The clever part is where the cost lands: the recurrence happens in the prefill phase, so generation latency is essentially unchanged. The paper distinguishes this from chain-of-thought, which the authors argue is better reserved for complex multi-step inference, and from looping and recurrent transformer training, which require expensive training runs. This is a training-free architectural change applied to off-the-shelf models.
The results on the Gemma3 family are hard to dismiss. Adaptive recirculation freezes the original weights and tunes only a few hyperparameters. It cuts perplexity by 23% across a suite of datasets and lifts GSM8k accuracy by 21%. A 23% perplexity drop from a training-free tweak is the kind of gain you'd normally expect from fine-tuning.
Key numbers23%: perplexity reduction on Gemma3 from adaptive recirculation, training-free. 21%: GSM8k accuracy gain from the same method. 0: additional generation latency, since the recurrence runs in prefill. Frozen weights: no retraining, no fine-tuning of the backbone.
Modular memory for long context: MoNe
Long-context inference has a scaling problem. In-context learning reads every token on every query, so compute and peak GPU memory both grow with context length. MoNe (Modular Neural Memory) attaches a lightweight neural memory to any frozen transformer instead.
The design is two-phase. In preprocessing, the model reads the context in fixed-size segments and writes fast-weight memory networks using test-time learning with layer-localized gradient updates. At query time, the memory generates keys and values from the query tokens alone. No context tokens get re-read. That's what decouples inference cost from context length: O(N) preprocessing, O(1) query cost, and peak GPU memory that stays flat as N grows.
At 128K tokens, MoNe cuts both compute and peak GPU memory by roughly 80% compared to in-context learning, with 6.4% parameter overhead. 128K tokens is a 300-page book or a mid-size codebase, and MoNe reads it without letting GPU memory grow with the input. The authors also show it generalizes beyond the backbone's native context window, where in-context learning degrades sharply, on RULER's needle-in-a-haystack and word extraction tasks.
Quick Take: every method in this cluster moves work from query time to build time, and that single shift is what makes the efficiency gains possible.
Key-frame selection: MemTree3D
Embodied 3D question answering has a specific inefficiency: for every user query, the system visually searches thousands of video frames to find the ones relevant to the question. That's linear in video length, per query, and it burns compute on every single question.
MemTree3D flips the order. Instead of searching frames per query, it builds a compact, reusable 3D scene representation once, in real time, using camera 6-DoF poses. The memory tree captures multi-level 3D scene information. When a question arrives, an LLM queries the tree and a scoring-based frame selection pulls out the question-relevant key frames. The video stream is never reprocessed.
On OpenEQA, MemTree3D improves LLM-Match by 17.4% for GPT-4o and 5.8% for LLaVA-OneVision-7B over existing visual search methods. The 7B model is the one that fits on a single consumer GPU, so the efficiency gain matters for real deployment, not just benchmark runs.
Pruning with better math: DVBP + OB²C
Pruning is the oldest trick in the efficiency book, but the math has been sloppy. Variance-Based Pruning (VBP) selects neurons by activation variance, which sounds reasonable until you look at the covariance spectrum from a finite calibration set. It's full of statistical noise. And VBP's bias-only updates can't fully compensate for the structural error of removing neurons.
DVBP + OB²C fixes both problems. The denoising side uses random matrix theory to filter noise from the activation covariance spectrum, so neuron selection is based on signal, not sampling artifacts. The compensation side integrates mean-shift into the Optimal Brain Compression objective, and the authors prove this reduces the layer-wise Hessian exactly to the activation covariance matrix. That yields a closed-form update for the remaining weights using the same statistics gathered for selection. No retraining.
The practical numbers: at 50% MLP pruning, DeiT, Swin, and ConvNeXt variants retain over 90% of original Top-1 accuracy. Cutting half the MLP parameters is the difference between fitting on an edge device or not. Compared to VBP, the gain is up to 29.46% on ConvNeXt-T and 7.33% on Swin-S.
Architecture-level wins: wavelet bases, SFMformer, and SineKAN
Some efficiency wins come from changing the model itself, not the inference procedure. Three papers in this cluster take that route.
Wavelet convolution layers enlarge receptive fields through multiresolution analysis, but prior work fixed the basis to Haar or Daubechies at a chosen decomposition depth. The new paper characterizes a tradeoff nobody had mapped: filter length vs. decomposition levels (F-vs-L). Longer filters cost more per level but reduce the depth required for competitive accuracy. Coiflet-based wavelet convolutions match Haar at deeper levels with roughly 32% fewer additional parameters and 33% fewer additional FLOPs. If you're building wavelet layers, Haar is not the default answer.
SFMformer tackles lightweight image super-resolution from an interesting angle. Sparse attention has two places where quality matters: the top-k selection of which tokens survive, and the aggregation of the survivors. A token dropped at selection can't be recovered downstream. SFMformer pairs a dual-branch spatial enhancement before the attention with a wavelet-domain modulation after it. The gains compound rather than add: the joint gain exceeds the sum of the individual gains on nine of fifteen benchmark-scale pairs, and the sign of that discrepancy correlates with how much the weaker module contributes on its own (r = -0.72). Running spectral modulation once per block instead of once per layer keeps the effect at roughly one-sixth of the cost. The model stays under a million parameters at every scale and ranks first on 28 of 30 PSNR/SSIM entries across five benchmarks and three upscaling factors. Under a million parameters means it runs on a Raspberry Pi 5, which is exactly where the authors deployed it.
SineKAN is the outlier: a KAN variant using sinusoidal activations instead of B-splines. It's smaller in scope, but it's a useful reminder that activation function choice in KANs is still wide open.
What the community is saying: I found SineKAN the way I find most research, lying awake at 2 a.m. wondering whether anyone had tried sinusoids in a KAN. Fortunately, or unfortunately depending on your sleep schedule, someone had. The paper had been sitting on arXiv since mid-2024, with a peer-reviewed version in Mathematics and a GitHub repo, and I'd missed it entirely. The Reddit thread where I found it had the usual mix of people who had tried something similar and people asking why this wasn't in the KAN survey. The takeaway: sinusoidal activations bring periodicity that helps on some function classes and hurts on others, and nobody has mapped that boundary yet.
| Approach | Core idea | Efficiency lever | Headline result | Where it fits |
|---|---|---|---|---|
| Wavelet conv layers | Multiresolution analysis with basis choice | Longer filters (Coiflets) reduce required depth | 32% fewer params, 33% fewer FLOPs vs. Haar | Vision backbones needing large receptive fields |
| SFMformer | Separate selection quality from aggregation quality in sparse attention | Spectral modulation once per block, not per layer | First on 28/30 PSNR/SSIM entries, under 1M params | Lightweight super-resolution on edge devices |
| SineKAN | Sinusoidal activations replace B-splines in KANs | Simpler basis, no spline grid | Community-explored; peer-reviewed version published | Small MLPs and function approximation |
What trips people up
Five mistakes I keep seeing in this area, all grounded in these papers:
Pruning with raw activation statistics from a small calibration set. VBP's weakness is finite-sample noise in the covariance spectrum. DVBP exists because of this. If you're selecting neurons from a few hundred samples, denoise the spectrum first, or you're fitting noise.
Treating Haar as the default wavelet basis. It's cheap per level, but you may pay for it in decomposition depth. Coiflets match Haar at deeper levels with about a third fewer parameters and FLOPs. Basis choice is part of the efficiency budget, not a fixed constant.
Re-running visual search over the full video for every query. MemTree3D's core move is building the 3D representation once from camera poses, then scoring frames per question. If your system searches thousands of frames per query, you're paying the video-processing cost N times.
Confusing recirculation with chain-of-thought. CoT is for complex multi-step inference. Recirculation is for belief-state tracking and adds no generation latency. Use the wrong one and you either burn tokens or leave accuracy on the table.
Assuming long context means retraining or a bigger window. MoNe attaches to a frozen transformer and gets O(1) query cost with 6.4% parameter overhead. Before you plan a full retrain for longer inputs, check whether a test-time memory layer gets you there.
One thing to remember: the most reliable efficiency wins in this cluster come from changing when work happens, not from shrinking models. Build the memory, the scene tree, or the pruning mask once, then keep per-query cost flat. That pattern shows up in every method here, and it's the one worth stealing for your own systems.
The Bottom Line
If you're running a frozen backbone and need long-context inference, adopt MoNe-style modular memory instead of planning a retrain. You get O(1) query cost and roughly 80% less compute at 128K tokens for a 6.4% parameter overhead.
If you're deploying a ViT on edge hardware, use denoised variance-based pruning with bias compensation rather than raw VBP. At 50% MLP pruning you keep over 90% Top-1 accuracy, training-free, and you beat VBP by up to 29.46% on ConvNeXt-T.
One thing to watch: wavelet basis selection and inference-time recurrence are both moving fast. Expect learned basis selection in wavelet layers and recirculation-style recurrence in production inference stacks within a year. Both are training-free, which makes them cheap to adopt.