Appearance
Fixed Points Change the Economics of Looped Transformers
The loop tax
Looped transformers reuse one block across depth. One parameter set, applied repeatedly. The idea goes back to the Universal Transformer and weight-tied architectures, but Huginn was the proof that loops could carry multi-task learning and in-context meta-learning at scale. The frontier took notice. OpenAI's GPT-6 series reportedly runs this architecture, and Microsoft confirmed it on a public page, briefly, before the claim was edited out.
The appeal is real: you decouple capacity from compute. Parameters stay small, depth grows with however many passes you can afford, and serial reasoning scales with the loop count.
The tax is the problem. Every recurrence is charged in four separate places:
- Training backpropagates through each loop.
- Decoding re-enters the loop for every generated token.
- Prefill repeats the recurrence over the whole prompt.
- RL replays trajectories through the looped stack.
A model that runs 7 loops pays a 7x multiplier in each of those places, independently. That's the loop tax. It's why naive looped models stayed a research curiosity.
Fixed points change the cost equation
"Towards Looped Models Done Right, Part II" (arxiv 2610.06833) attacks the tax at the root. The core claim: as recurrent states converge toward fixed points, the path to them stops mattering. If the model lands in the same state whether you ran 4 loops or 8, you don't need to trace or differentiate the route.
That single idea produces four concrete savings:
| Stage | Naive approach | What fixed points enable | Reported gain |
|---|---|---|---|
| Training | Full backprop through every loop | Truncated backpropagation | Less compute per step, same target |
| Decoding | Recompute KV for each pass | Terminal KV sharing | Almost no accuracy loss |
| Prefill | Run the recurrence over the prompt | Distilled student at the fixed point | up to 1.79x faster |
| RL | Backprop through replayed trajectory | Gradients from saved rollout states | 2x faster updates |
Truncated backpropagation means you don't pay for loop depth in the backward pass. Terminal KV sharing means the converged key-value state is reused across decoding steps instead of recomputed. The 1.79x prefill speedup shows up directly as time-to-first-token. And the 2x RL figure means policy-gradient runs finish in half the time.
The KV result is the one to watch. At 1.6B parameters, a model trained with the learned depth prior and a 3x smaller KV cache matched the downstream average of fixed-depth training with the full cache. That's a memory cut to a third with no measurable downstream cost, on a model size where KV cache is a real operational line item.
Quick Take: The looped-model question isn't how many loops you run. It's whether the state converges to something reusable. If it does, the path to the fixed point is dead weight.
Shaping the fixed point
Convergence doesn't just happen. Two training choices determine where the fixed point lands: the depth prior and input injection.
Huginn's depth prior samples the number of loops from a broad distribution. Breadth is what makes terminal KV sharing possible, because the model learns to reach similar states across depths. But breadth dilutes supervision at the target depth more than sharing requires. Fixed-depth training gives clean supervision and breaks KV sharing entirely. The paper's middle path: learn the prior from prediction feedback, with an entropy term that keeps it broad. From 100M to 1.6B parameters, the learned prior lowers perplexity at every scale relative to Huginn's.
Injection is subtler. The recurrent state carries a component aligned with the input, and existing injection schemes let that component amplify or cancel the injected input depending on its sign. Orthogonal injection strips the input-aligned component before mixing. Again, perplexity drops at every scale tested relative to the existing schemes.
The two fixes are independent and complementary. The prior sets how many loops you run; the injection sets what enters the recurrence in the first place.
What Birkhoff geometry says about hyper-connections
The companion paper (arxiv 2610.06653) explains where these fixed points come from geometrically. Manifold-constrained hyper-connections (mHC) widen the residual stream to n parallel tracks and mix them at each layer with a doubly stochastic matrix, computed by Sinkhorn-normalizing exponentiated logits. The analysis lives on the Birkhoff polytope, the set of all doubly stochastic matrices.
The headline result reads like a decomposition theorem. A doubly stochastic mixer splits the stream into a mean channel, on which mHC behaves exactly like a residual network, and a difference channel, which each layer contracts by its second singular value σ2. The bound: σ2 ≤ 1 - n·min_ij H_ij. So the extra width is a fading memory with a horizon of 1/(1-σ2) layers. At σ2 = 0.99 you get roughly 100 layers of memory. At σ2 = 0.9, about 10. Among nonnegative mixers, only the permutations avoid collapse entirely.
The training story is just as clean. The Sinkhorn-logit map is a global chart on the polytope, and the logit gradient is exactly the Fisher-Rao gradient. Gradient flow on the logits follows the squared Fisher-Rao metric; the straight-through update everyone actually uses is exactly entropic mirror descent. The rates differ: gradient flow approaches the polytope's vertices at 1/t, mirror descent at an exponential rate. The bound on how fast each log-entry moves under gradient flow is 4n³||∇f||∞ ε, where ε is the distance to the nearest permutation.
There's a practical constraint buried in the last result. Sinkhorn's local convergence factor is σ2², so a fixed iteration budget limits the horizon you can realize. Want a mixer close to a permutation for long memory? That's precisely the slowest case for Sinkhorn to reach.
Key Numbers
- 1.79x: faster prefill from a distilled student at the fixed point
- 2x: faster RL updates, gradients from saved rollout states
- 3x: smaller KV cache at 1.6B, downstream parity with full-cache fixed-depth training
- 1/(1-σ2): memory horizon in layers for manifold-constrained hyper-connections
- σ2²: Sinkhorn's local convergence factor, which caps realizable horizon
The GPT-6 confirmation that vanished
GPT-6's looped architecture was an open secret before it was a confirmed one. The Information reported it. The rumor mill repeated it. Then Microsoft posted a page that said it plainly, and the internet archived it before the page was edited.
When the page was live, I went looking for the looped-transformer line before the thread could tell me what it said. The phrasing was matter-of-fact, like a spec sheet item: GPT-6.1 Sol runs two inference passes, "instead of three." The thread that followed spent most of its energy on "same base model weights as GPT-6 Sol." My read, after comparing the wording: GPT-6.0 and GPT-6.1 are both post-trained on the same pre-trained base model, with different post-training and one fewer loop in the 6.1 checkpoint. Same base. Different post-training. One less pass.
Then Microsoft edited the page and the claim disappeared. A retraction without a correction. The community read the sequence the same way I did: the earlier reporting was right, and the page slip was the confirmation.
json
{
"type": "bar",
"title": "Inference passes per response: GPT-6.0 Sol vs GPT-6.1 Sol",
"x_label": "Model",
"y_label": "Inference passes",
"caption": "Microsoft's page listed three inference passes for GPT-6.0 Sol and two for GPT-6.1 Sol, one fewer loop over the shared base model.",
"data": {
"labels": ["GPT-6.0 Sol", "GPT-6.1 Sol"],
"series": [
{
"name": "Inference passes",
"values": [3, 2]
}
]
}
}If the fixed-point results hold, the drop from three passes to two is exactly what the theory predicts. Once the terminal state converges, the extra pass contributes less than it costs. You remove it, and nobody can tell the difference.
Common pitfalls
Working with looped models, I keep seeing the same mistakes repeat. Here are the ones these papers seem designed to prevent.
Don't assume the loops converge. The entire fixed-point playbook depends on the state actually settling. If you choose hyper-connection mixers that contract slowly (σ2 near 1), a fixed Sinkhorn iteration budget leaves you far from the target mixer, and your horizon claims never materialize. Check σ2 before you trust any memory-horizon number.
Don't use plain input injection. The state's component along the input direction amplifies or cancels the injection depending on sign, which destabilizes the fixed point the model converges to. Orthogonalizing the injection before mixing fixes the interference and improved perplexity at every scale from 100M to 1.6B. It's a small change with a consistent win.
Don't treat the depth prior as a fixed hyperparameter. A broad prior enables KV sharing but dilutes supervision; fixed-depth training sharpens supervision but breaks sharing. The learned prior with an entropy term gets both. At 1.6B, that choice alone matches fixed-depth downstream quality with a third of the KV cache.
Don't mistake hyper-connection width for residual capacity. The mean channel is just a residual network; the difference channel is a fading memory. If you need long-range information to survive, only mixers near permutations do it without collapse. Random Sinkhorn mixers lose the difference channel in tens of layers.
Don't backpropagate through replayed RL trajectories. The rollout states have already converged; unfolding them again doubles your gradient cost for nothing. Compute gradients from the saved states instead, and the reported 2x speedup is yours.
One thing to remember
Looped transformers were expensive because every recurrence charged the full pipeline. The fixed-point framing dissolves most of that cost: converge the state, then truncate, share, distill, and differentiate at the terminal point. The design choices that shape the fixed point, the depth prior and the injection direction, decide whether the savings are real. Get those right, and the looped model stops being a research curiosity and becomes a cheaper way to buy depth.
The Bottom Line
If you're training a looped LM and you can verify convergence, swap full backprop for truncated backpropagation and terminal KV sharing. The paper shows the path to the fixed point is expendable, with KV memory cut to a third at 1.6B scale and no downstream loss.
If you're running RL against a looped policy, stop backpropagating through replayed trajectories. Gradients from saved rollout states are 2x faster because the rollout already converged, and unfolding it again adds nothing.
If you're designing hyper-connected architectures, treat σ2 as a memory knob with a cost. Want a long horizon? Push the mixer toward permutations, but budget enough Sinkhorn iterations, because the local convergence factor σ2² makes near-permutation mixers the slowest to reach.