Skip to content

Same Weights, More Compute: The Looped Transformer Frontier

#looped-transformers #mixture-of-experts #recurrent-models #test-time-scaling #diffusion-models #scaling-laws

Same Weights, More Compute: The Looped Transformer Frontier ​

The core trade: depth without parameters ​

Transformers spend most of their cost in execution, not storage. Once a checkpoint is trained, every forward pass through those layers costs the same, whether you ask for a one-token answer or a long chain of thought. Looped transformers target exactly that point: apply the same block several times, and you spend more compute per token without growing the parameter count.

That idea has made looping one of the more active corners of efficient scaling research. Recurrence gives you depth at fixed parameters. Mixture-of-experts gives you capacity at fixed active compute. The two axes are independent in principle, and four papers released in the same window argue they combine cleanly in practice.

WorkQuestion it attacksCore mechanismHeadline result
Loop Scaling Laws (arXiv 2609.40316)How do recurrence and sparsity combine?Bounded, sparsity-conditional recurrence mapping in a joint scaling lawLooped MoE matches a ~2x larger non-looped MoE at matched training compute
What Limits Recursive Reasoning (arXiv 2609.39967)Why are recursive models unstable to train?Intermediate gradient horizon, large physical batches, controlled recurrent state updates13.6M-param model reaches 71.2% Arithmetic OOD, up from 36.2%
Shared Weights, Selected Computations (arXiv 2609.39892)Does every loop repeat the same operation?Hidden-state steering through a learned linear layer JAttention routing is the causal pathway for selecting transitions
Looped Diffusion Transformer (arXiv 2609.40305)Does looping help text-to-image generation?Deep supervision plus self-modulating attention260M model beats a 6.5x larger baseline at 4.9x lower inference compute

The connecting thread: weight sharing is not behavior sharing. All four works end up supervising or routing the intermediate loops, because naive weight tying alone has a ceiling. If you want more depth, you have to decide what each pass over the block is for.

Scaling laws that count the loops ​

Most scaling laws model recurrence or sparsity in isolation. That's a gap, because the two interact in practice. The Loop Scaling Laws paper introduces a bounded, sparsity-conditional recurrence mapping that captures how many effective parameters looping adds, and how sparsity raises that gain. The fitted laws recover standard dense and MoE scaling laws as special cases, and they predict held-out loss more accurately than earlier alternatives.

Concrete numbers: sparsity delivers about 3x active-parameter efficiency, meaning you need only a third as many active parameters to reach the same loss. Recurrence delivers roughly 2x total-parameter efficiency on reasoning tasks, which reads as a looped model with P total parameters behaving like one with 2P. The gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matched a ~2x larger non-looped MoE on reasoning benchmarks. And recurrence stays available at test time, so you can keep trading latency for accuracy after deployment.

The practical reading is that the law turns looping from a research trick into a design parameter. You can compute how many loops to run for a given parameter and compute budget instead of guessing. That matters because the obvious way to use loops, adding them everywhere, is wrong half the time.

Quick Take: recurrence and sparsity are independent efficiency axes, and a looped MoE behaves like a model roughly twice its size.

The winning recipe is plain ​

Prior recursive reasoning models, HRM, TRM, and URM, differ in architecture, gradient propagation, and training procedure at the same time. That makes it impossible to say which design choice drives their results. The study running them all through one pipeline across six algorithmic domains, then ablating each factor on representative domains, changes the question entirely: how much of recursive reasoning is architecture, and how much is just optimization?

The answer turns out to be mostly optimization. The winning recipe has three parts: an intermediate gradient horizon, not a full unroll and not single-step training; large physical batches; and controlled updates of the recurrent state. No explicit hierarchical architecture needed. That last point matters because the field has spent real effort on hierarchical designs, and this suggests the hierarchy was compensating for optimization problems rather than adding representational power.

Combined into one 13.6M-parameter model, the recipe delivers the strongest overall performance among the recursive baselines, with the biggest gains out-of-distribution. Arithmetic OOD accuracy jumps to 71.2% from 36.2% for the strongest baseline. Sudoku hits 98.41%. ARC-AGI-1 reaches 59.5% pass@2. To put the size in perspective: 13.6M parameters fits on a single GPU with room to spare. That's the whole pitch, these compact solvers are meant to be tools that a larger LLM calls on narrow algorithmic subproblems without burning its own context.

Key Numbers~3x active-parameter efficiency from MoE sparsity in looped models ~2x total-parameter efficiency from recurrence on reasoning tasks 71.2% Arithmetic OOD accuracy for the 13.6M recursive model, up from 36.2% 4.9x lower inference compute for the looped 260M diffusion model versus a 6.5x larger baseline

Every loop is doing something different ​

Weight sharing alone doesn't force identical computation. That sounds like a technicality, but it's the central question for looped architectures: if the same weights run every loop, what makes the model do different things at different iterations?

The graph walks study (arXiv 2609.39892) answers this directly. In the model's native trajectories, decoded predictions can advance by several graph steps in a single loop, or stop at a reached target. There's no fixed one-loop-one-step pattern. Then they show control is exercised through the hidden state: modifying the entering hidden state of a frozen loop steers it toward different transitions, and a learned linear layer J selects the desired transition without touching the shared transformer layers.

Activation patching closes the causal chain. The experiments find that attention patterns can recover the effects of J, and can switch which transition gets selected. Across five matched pairs of graph models, changing the intermediate supervision during backbone training changes which transitions J can induce. So J selects computations the backbone already learned; it doesn't create new algorithms. The steering mechanism is real, but its range is set during training.

This may be the most transferable finding of the four. Looping buys you the ability to run different operations per iteration, and the hidden state is the dial. But the dial's range depends entirely on the supervision you gave the model while it was being trained.

Diffusion wants loops, but not naive ones ​

Text-to-image scaling has traditionally meant bigger models or more denoising steps. The Looped Diffusion Transformer paper tries a third route: run shared transformer blocks repeatedly inside each denoising step, increasing computational depth at a fixed parameter count.

Naive looping doesn't work. Image quality fails to improve consistently, and the paper traces that to two causes. Weak supervision across intermediate loops, so the early iterations get no gradient signal about what they should compute. And unregulated attention updates that progressively erode local information.

Looped-DiT fixes both. Deep supervision across intermediate loops gives every iteration a target. Self-modulating attention keeps feature updates stable across loops. Under matched parameters and matched compute, Looped-DiT beats non-looped baselines. A 260M-parameter looped model surpasses a model 6.5x larger, roughly 1.7B parameters, across multiple text-to-image benchmarks, while requiring 4.9x lower inference compute.

The more interesting claim: under a fixed inference budget, increasing loop depth yields larger gains than adding more denoising steps. And the behavior shows signs of latent reasoning. Deeper loops progressively correct mistakes made in earlier loops, which is exactly the pattern the routing paper predicts: each pass gets to do something different, and later passes can fix earlier errors.

Common pitfalls ​

  • Don't tie weights and loop without intermediate supervision. Both the diffusion and routing papers show the same failure: unguided intermediate states get no useful gradient signal, and the loop underperforms a single pass. Add deep supervision over loop outputs, not just the final state.
  • Don't assume one loop equals one reasoning step. Graph walk trajectories advance by variable numbers of steps per loop, and can stall at a reached target. Design curricula and evaluations that tolerate variable per-loop progress, or you'll misread what the model is doing.
  • Don't use full backprop unrolls just because depth helps. The recursive reasoning study found that an intermediate gradient horizon is the stable choice. Full unrolls destabilize training, and single-step training leaves performance on the table.
  • Don't let attention run unregulated across loops. Attention updates that drift over iterations progressively erase local information, which is exactly what killed naive looped diffusion. Self-modulation or equivalent constraints belong in the design from the start.
  • Don't treat the shared block as a fixed algorithm. The entering hidden state routes what each loop computes. Change the state dynamics, normalization, or residual scaling between loops, and you silently change the behavior. Sometimes for the better. More often not.

Sources ​

All four preprints used in this article:

  1. Scaling Laws for Looped Mixture of Experts, http://arxiv.org/abs/2609.40316v1
  2. What Limits Recursive Reasoning Models: Optimization, Architecture and Test-Time Scaling, http://arxiv.org/abs/2609.39967v1
  3. Shared Weights, Selected Computations: How Looped Transformers Route What Each Loop Does, http://arxiv.org/abs/2609.39892v1
  4. Looped Diffusion Transformer, http://arxiv.org/abs/2609.40305v1

One thing to remember ​

In a looped transformer, the hidden state is the steering wheel. Weight sharing doesn't make loops repeat. The entering state, along with the intermediate supervision used during training, decides what each pass computes. The scaling law, the training recipe, and the diffusion fixes all hang off that single observation.

The Bottom Line ​

If you're compute-bound during training but have room to spend more time at inference, adopt looped MoE with law-derived recurrence. It matches a non-looped MoE about twice its size at equal training compute, and you can dial loop count up at test time without retraining.

If you're building a small reasoning module for an LLM to call on algorithmic subproblems, skip hierarchical architectures and use the simple recipe: intermediate gradient horizon, large physical batches, controlled recurrent state updates. A 13.6M model built that way doubled Arithmetic OOD accuracy over the strongest baseline.

If you're adding loops to a generation model, plan for supervision from day one. Naive weight tying erodes quality, and the fix, deep supervision plus attention stabilization, has to be in the training objective from the start. Expect deep supervision on intermediate loops to become standard in looped generators within the next year.