Appearance
Four Papers That Make Model Training and Inference Cheaper
The efficiency wall has more than one crack
Roughly once a week, a paper lands that finds compute you were wasting. LoopCD finds it in the discarded intermediate states of looped transformers. SoftServe finds it in curvature information that first-order optimizers throw away. A new universality result finds it in the ordering of frozen attention blocks. And a NeurIPS 2026 spotlight finds it in the time axis of RNN training, parallelizing a dimension everyone treats as strictly sequential.
All four share an assumption you should notice: the standard recipe of more parameters and more tokens is not the only path to better results. Sometimes the gains are already sitting in the pipeline you run.
Key numbers
- 11.45 points: the AIME 2024 pass@1 gain on Ouro-2.6B-Thinking from LoopCD (61.88% to 73.33%). That is the difference between a model that guesses at contest problems and one that reasons through a solid share of them.
100x: reported training speedup on chaotic time series from parallel-in-time DEER training with generalized teacher forcing. A run that took a month fits in an afternoon.
- 22.5% to 48.2%: forward FLOPs saved when LoopCD lets you halve the recurrent loops and still match the unguided baseline.
- 2: the number of frozen single-head attention blocks the universality paper needs to interpolate arbitrary token collections.
LoopCD decodes what standard decoding discards
Looped transformers run a shared block over and over. Each pass refines the representation, and every intermediate state can already decode the next token. Standard decoding only looks at the final loop. Everything before it is computed and thrown away.
LoopCD turns that waste into signal. Earlier loops embody less computation, so they naturally produce weaker predictions. Weak and strong predictions for the same token form exactly the aligned pairs contrastive decoding wants. Until now, getting those pairs meant training a small auxiliary model or running a second pass with a different configuration. LoopCD gets them from the model's own recurrence, at zero training cost.
Two variants. LoopCD-Logits contrasts in logit space and needs one extra output pass. LoopCD-Hidden contrasts in hidden-state space and runs with zero output overhead.
Both deliver at full recurrent depth:
The contrast signal is where the gains come from, and the gains convert directly into compute. Because the guided model matches or beats the full-depth baseline, you can halve the number of recurrent loops and still come out ahead. That is the 22.5% to 48.2% FLOP reduction. For a model that loops four times per token at serving time, skipping two loops is a real latency win, not a rounding error.
SoftServe makes quasi-Newton practical for deep learning
Quasi-Newton methods were designed for convex optimization, where they shine by estimating curvature and taking steps that plain gradient descent can't. Two things kept them out of deep learning. The loss surface is non-convex, so curvature estimates go sideways. And models have a billion parameters, so storing a Hessian approximation is laughable at scale.
SoftServe clears both hurdles. Curvature comes from a variational objective that produces positive-definite estimates even where the local surface bends the wrong way. Diagonal and Kronecker-factored variants keep the memory footprint reasonable. The coupled Newton-Schulz iteration replaces matrix decompositions with matrix multiplications, which is exactly what GPUs are good at.
The paper shows wins on problems with brutal conditioning: recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model. 136M parameters is modest by today's standards, but physics-informed objectives are stiff in a way that punishes first-order methods. Adam zigzags down a canyon. Muon and SOAP do better but still leave curvature on the table. SoftServe reaches lower loss than all three on these tasks.
| Optimizer | What it tracks | Where it tends to win | Where it struggles |
|---|---|---|---|
| Adam | Diagonal gradient moments | General-purpose training | Ill-conditioned stiff objectives |
| Muon | Orthogonalized momentum | Wide networks | Same stiff regimes |
| SOAP | Kronecker-factored statistics | Transformer-scale models | Approximate conditioning |
| SoftServe | Positive-definite quasi-Newton curvature | RNNs, autoencoders, PINNs | Not yet shown at LLM scale |
That last row is the honest caveat. The largest SoftServe demo is 136M parameters. The method is built to scale: Kronecker-factored, no line searches, GPU-friendly primitives. But the field will want to see it on a real pretraining run before declaring the optimizer wars over.
Quick Take: LoopCD turns discarded intermediate states into decoding signal. SoftServe turns discarded curvature into training signal. Both are free lunches hiding inside pipelines you already run.
Frozen attention blocks can be universal
Universal approximation is usually an infinite-width story. Take enough hidden units and a network can approximate any function. That claim is true and, for modern shared-parameter architectures, almost beside the point. This paper works the other side of the duality: approximation power from depth alone, with strong parameter sharing across layers.
The setup is a question. Is there a finite set of attention blocks, fixed in advance, such that applying them in the right order, with the right signs and durations, maps any collection of N sequences of n tokens to any other collection? The answer is yes. The set is only two frozen single-head blocks with Gaussian-initialized projection matrices. The weights never move. Everything learnable lives in the schedule: which block, for how long, with what sign.
That is a strong statement, and it holds at continuous and finite depth. The paper also characterizes what causal masking breaks. Masking restricts which token positions can influence which outputs, and the authors pin down exactly which interpolations become impossible under those constraints.
This matters for anyone training models. Parameter sharing is not an expressivity tax. The theory says a family of deep residual self-attention networks with reused blocks can express transformations that would normally demand fresh weights per layer. Looped transformers get their validity at the level where scaling laws can actually apply. It is the theoretical counterpart to LoopCD: the decoding trick works because shared-block schedules carry real expressive weight.
Training RNNs in parallel time
The Reddit thread on the NeurIPS 2026 spotlight makes the setup personal. Train a nonlinear RNN on anything with a positive Lyapunov exponent and backprop-through-time becomes a war of attrition. One step per timestep. Gradients that vanish or explode. Exposure bias creeping in as teacher forcing fades. I have burned weeks of cluster time on exactly this pattern, and watching the loss flatline while the clock runs is its own kind of slow motion.
DEER changed the mechanics. Instead of marching through T timesteps, it treats the forward pass as a fixed-point problem and solves it with Newton-type iterations across the whole sequence. That collapses sequential cost from O(T) to O[(log T)²]. The catch: chaotic dynamics break the iteration, and the paper's own references show runtime degrading back to O[T log T] with divergence.
Generalized teacher forcing (GTF) is the stabilizer. It prevents the chaos-driven divergence, and it reduces exposure bias relative to plain teacher forcing for state space models. The combination makes the speedup real. The authors report stable training on sequences past a million timesteps, over 100x faster, with results that beat Mamba and other state space models on dynamical systems reconstruction.
A million timesteps per sequence is past what most BPTT pipelines even attempt. The 100x claim means a training experiment that used to occupy a weekend of four GPUs now finishes in a couple of hours on one. That changes which research questions are even on the table.
What the community is saying on the Reddit thread echoes the failure mode I ran into. Plain DEER looks great on clean synthetic systems, then falls apart the moment the dynamics turn chaotic. GTF does more than suppress that divergence: it changes which fixed point the iteration converges to. Teams that had written off DEER after hitting the chaotic wall are revisiting it.
The common thread: stop wasting what you already paid for
Put side by side, the four results look different, but the mechanism is the same. Each one identifies a resource already embedded in the pipeline that standard practice discards.
| Method | Stage | What it exploits | Reported gain |
|---|---|---|---|
| LoopCD | Inference, decoding | Intermediate recurrent states already computed | +11.45 AIME points, 22.5-48.2% fewer FLOPs |
| SoftServe | Training, optimizer | Curvature information in the loss surface | Lower loss than Adam, Muon, SOAP on ill-conditioned tasks |
| Universal interpolation | Theory | Expressive power of block ordering and duration | Universality with two frozen attention blocks |
| DEER + GTF | Training, parallelism | The time axis, via fixed-point parallelization | >100x training speedup on sequences past T = 10^6 |
For foundation model teams the reading is direct: the field is shifting from buying more compute to extracting more from the compute you already run. That is a good shift for anyone who cannot order another cluster.
Common Pitfalls
Four things trip people up with these methods.
Picking the wrong contrast loop in LoopCD. Contrastive decoding needs the weak prediction to be aligned with the strong one. If you reach back too many loops, the early state disagrees with the current context, and the contrast signal becomes noise. Tune which intermediate pass you contrast against per model; the paper's gains assume the right alignment.
Diagnosing conditioning before switching optimizers. SoftServe shines where the problem is severely ill-conditioned. On a well-behaved transformer run, the per-step cost may not pay for itself. Check whether your loss surface looks like a canyon before you swap away from Adam. Gradient norm histograms and curvature proxies tell you more than benchmark gossip.
Running DEER on chaotic dynamics without stabilization. The raw Newton iteration diverges and silently degrades to sequential runtime. If your time series has positive Lyapunov exponents, GTF is required. Teams that skip it report exactly the O[T log T] degradation the authors describe.
Misreading the universality result as a training recipe. The interpolation schedules, block order, signs, and durations, are existential. The proof says the capacity is there. It does not say gradient descent will find those schedules. Treat it as a green light for shared-parameter architectures, not as a method that trains itself.
Treating FLOP reductions as wall-clock reductions. LoopCD's 22.5% to 48.2% FLOP cut assumes you actually run fewer loops end to end. The logit-space variant adds an output pass, and hidden-state contrast has memory bandwidth costs. Profile on your serving hardware before promising latency wins.
What to adopt, and where each method pays off
One Thing to Remember: every one of these papers converts something you already compute into something you can use. Intermediate states become contrast signals. Curvature becomes step direction. Block order becomes expressive capacity. The time axis becomes a parallel dimension. No new hardware, no bigger models.
If you serve or plan to serve a looped transformer, adopt LoopCD-Hidden first. It costs zero extra output passes and bought Huginn nearly 10 HumanEval points, and the logit variant is there when the latency budget allows one more pass.
If your training curves on recurrent nets, autoencoders, or physics-informed objectives look like a sawtooth, swap Adam for SoftServe. Ill-conditioned problems are exactly where curvature pays, and the Kronecker-factored variant is built to scale past 136M parameters, but start there and measure step-to-loss efficiency against Muon.
If long chaotic time series are blocking your state space model work, move to DEER with generalized teacher forcing. The >100x speedup on T > 10^6 sequences is the gap between an overnight run and a week-long cluster reservation, and the Mamba comparison on reconstruction tasks is favorable enough to test directly.
One thing to watch: three of these results appeared within weeks of each other, and all four target the same wall. The next six months will show which generalize past their reported setups. The LoopCD experiment on non-looped transformer families is the one I would bet on. If intermediate states in a shared-weight decoder carry that much signal, the same trick may work across ordinary deep transformers.