Appearance
One model, many compute budgets
A deployed language model has to serve more than one compute budget. There's the free tier, the paid tier, the latency-sensitive API, the on-device build, and the batch job that can afford ten seconds per answer. The standard way to hit all of those points is to train or distill a separate model per budget. N budgets means N training runs, N artifacts to store, evaluate, and monitor. Add a tier and you rebuild the whole quality-cost curve from scratch.
Three papers posted in September 2026 push the other direction: keep the parameters, change the compute. Telescopic Language Models (TLM) train one transformer that is a valid language model at every depth. Foil works out how to loop a sparse mixture-of-experts (MoE) block so you spend more compute per token without moving the expert count. TaH2 attaches an iteration decider that gives extra loop passes only to the tokens that benefit. Different mechanisms, same intent: make capacity a dial you turn at inference instead of a point you lock in at training time.
If you haven't met looped transformers, the idea is simple. Run one block, or a short stack of layers, several times over the same hidden state. More passes, more compute, same parameters. The three papers are best read as three answers to what happens when you take that idea seriously.
The telescopic training objective
TLM is the cleanest statement of the capacity-as-a-dial idea. Take a normal transformer with L layers. At every training step, do two forward-backward passes. The first runs the full model and predicts the next token; call it the anchor. The second runs a truncated prefix that stops at a randomly sampled depth d, and it predicts the exact same full next-token target. Only the amount of visible model differs.
That second detail is the whole trick. A prefix never gets to lean on the layers above it. Over millions of steps, every prefix has to produce the full next-token prediction alone, so every prefix becomes a complete language model. The depth axis becomes a continuum of capacities inside one set of weights.
The design costs almost nothing. Two forward-backward passes per step, no architectural change, nothing extra at inference. You read out however many layers your budget allows and stop. That's the entire interface.
Why fixed exits leave the model at chance level
A lot of the paper is a takedown of the obvious alternative. Matryoshka Language Model Suites (MLMS) supervise a handful of fixed exits and call the result elastic. TLM's authors ran those suites as baselines and found a specific failure mode: the supervised exits are fine, but everywhere between them, the model has never been trained to predict anything. Their baselines show perplexity 10² to 10⁵ at unsupervized depths.
In concrete terms: perplexity around 20 is readable text on a modern tokenizer. Perplexity 100 is already bad. Perplexity 10⁵ means the model assigns near-uniform probability to every token in the vocabulary, so the output is effectively random. A fixed-exit suite is a Swiss cheese model. A few good points, nothing between them, and you only discover the holes when you actually serve an intermediate budget.
The sharper point is architectural. MLMS is one point in the design space TLM maps, and the failure isn't the nesting, it's the supervision. The paper puts it bluntly: the training objective, not the nesting itself, is what makes a model elastic.
The continuum is a training-time choice
The evidence comes from a 200M-parameter proxy trained on 20B FineWeb-Edu tokens, with an identical data stream for every method. Research-scale, but the comparison is clean because nothing else varies. The single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks. It cuts the area under the quality-budget curve by 43-44% relative to MLMS while matching it at full capacity, and it does that at about 12% lower GPU cost per run.
A 43-44% area reduction means almost half the territory that used to require separate training runs now comes from one artifact. The 12% cost saving is the part people gloss over: the elastic model isn't just more convenient, it's cheaper to train than the fixed-exit suite.
The prefix sampling density is a dial in the same sense. Concentrate sampling on a few depths and you recover fixed-exit quality there, at the price of the continuum. You can convert TLM into an MLMS by turning the dial. That's what it means for operating points to be a training-time choice rather than an architectural one.
Quick Take: If the training objective doesn't supervise every operating point, the model isn't elastic. It's just nested.
Looping a MoE without wrecking it
The second paper attacks the other axis of elasticity. Looping is compute-bounded, MoE is parameter-bounded. A looped MoE gets both: a fixed set of expert parameters applied more times per token. But nobody had worked out how to loop a MoE without degrading routing. Foil answers with two moves, holding expert parameters and expert compute per token fixed.
Flatten the experts. Halve the number of expert layers, double the experts per layer, double the passes. Every routing decision now draws from a larger pool, so each choice matters more.
Untie the attention. Each pass gets its own attention parameters while the experts and routers stay shared. The router picks experts from the big pool, and each pass interprets the result through its own attention.
The numbers back the design. At 20B tokens, every Foil variant has lower pretraining loss than the unflattened looped baseline. At 100B tokens, the loss improves monotonically with the degree of flattening, and the most flattened Foil ends 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better. Nats are natural-log units of cross-entropy; 0.012 at this scale is small on paper but consistent across a 100B-token run, and the monotonicity tells you it's a direction, not noise. Untying attention also yields more balanced and more confident routing at equal shape. Code and configs are up at github.com/SR-A-W/how-to-loop-moe.
Routing confidence beats load balance
Foil's ablations turn into direct design guidance, and this is where the field's conversation is currently stuck. Load balance has been the default MoE health metric for years. Most MoE training runs include an auxiliary loss that pushes experts toward uniform usage. Foil's evidence says that's the wrong target, because load balance measures uniformity, not usefulness. Forcing uniformity can silently compress the routing information the model needs.
What the community is saying: the live debate is about which metric should guide elasticity work. The TLM paper draws one line, arguing fixed-exit suites like MLMS shouldn't count as elastic when unsupervized depths sit at perplexity 10² to 10⁵. Foil draws another, arguing routing confidence tracks healthy expert use better than load balance. Neither argument has settled yet. Both will get re-run by other labs before the field moves on.
Reading the three ablations back to back, the line I keep coming back to is Foil's: "the returns of looping and of widening the expert layers amplify each other." That's a concrete recipe. A sparse looped MoE should use more experts per layer and more passes, not the same experts run more times.
Key numbers
- 43-44% lower area under the quality-budget curve for TLM vs the fixed-exit MLMS suite
- 10²-10⁵ perplexity for MLMS at depths that were never supervised, which is random-text territory
- 0.012 nat Foil's final loss gain at 100B tokens, at equal parameters and compute
- 2.74 vs 1.79 TaH2's accuracy-compute slope vs the non-looped baseline on AIME
Spending iterations where they work
The third paper asks whether looping pays off as outputs get longer. Prior work compared looped and non-looped models at matched parameters or per-token FLOPs, which sidesteps the actual question: what the accuracy-compute slope looks like as test-time decoding FLOPs double.
TaH2's authors found a frustrating pattern. Existing looped transformers often have steeper slopes than their non-looped baselines, yet they underperform at matched compute. And fixed-depth looping spends extra iterations on every token, while many tokens gain nothing from the second pass. So they built an iteration decider, post-trained jointly with the backbone through lookahead depth supervision. The decider gets online labels saying whether one more iteration would improve the prediction for that token. Extra compute goes where it moves the answer.
On AIME, TaH2 improves the accuracy-compute slope by 53% over the non-looped baseline: 2.74 vs 1.79. In practical terms, every doubling of test-time compute buys roughly half again as much accuracy as the baseline model would give you. It also exceeds the baseline's peak accuracy by about 3.4 points at matched test-time compute.
The depth behavior is the part to watch. As the maximum allowed iteration depth grows, existing looped models mostly plateau. TaH2 keeps gaining: +2.8 points over the baseline at depth 2, +3.9 at depth 8. That's adaptive allocation doing its job, and it's the strongest evidence yet that uniform depth is the wrong default. Code is at github.com/thu-nics/TaH.
The three dials
Put side by side, the papers cover three different levers on the same machine.
| Paper | Mechanism | What it costs | Headline result |
|---|---|---|---|
| TLM | Stochastic prefix supervision with a full anchor | 2 forward-backward passes per step, ~12% cheaper than MLMS | 43-44% less area under the quality-budget curve at 200M scale |
| Foil | Flattened expert layers, pass-specific attention | No extra training cost; more MoE passes per token | 0.012 nat lower loss at 100B tokens, on par downstream |
| TaH2 | Iteration decider with lookahead depth supervision | Post-training pass on a looped backbone | 53% steeper accuracy-compute slope on AIME, +3.4 over baseline peak |
The shared claim is what matters: a model can be trained once and used across compute levels. TLM makes depth a continuum, Foil makes MoE passes a meaningful scale, TaH2 makes loop depth a per-token decision. Any one of these gives you a softer quality-budget curve. All three together start to look like a model that doesn't care which budget it's served under.
Common pitfalls
The mistakes show up when people take shortcuts around the training objective.
Supervising a few fixed exits and calling the model elastic. Your unanchored depths will sit at perplexity 10² to 10⁵, which is random text. If you serve an intermediate budget from an MLMS-style model, you serve garbage and you won't know until users complain. Supervise the continuum or don't claim the continuum.
Looping the whole MoE block, attention included. Shared attention parameters across passes measurably hurt routing balance and confidence. Keep experts and routers shared, but give each pass its own attention.
Adding passes without widening the expert pool. The return on looping and the return on widening expert layers amplify each other. If you double the passes, flatten the expert stack to double the experts per layer. Running the same narrow expert layer more times leaves most of the gain on the table.
Tuning the router against load balance. Load balance tells you about uniformity, not usefulness. Foil's ablations show routing confidence is the healthier signal, and squeezing an auxiliary loss toward uniform usage can hide experts that route consistently but badly.
Spending extra loop iterations on every token at test time. Fixed-depth looping is the default and it plateaus as max depth grows. A large share of tokens gain nothing from a second pass. If you can afford a decider, gate depth per token.
One thing to remember
Elasticity is decided at training time, not at inference time. A model only serves a range of compute budgets if the training objective made every operating point competent. TLM supervises the whole depth axis. Foil reshapes the MoE block so every pass makes a better routing decision. TaH2 teaches the model when extra compute stops paying. If you evaluate a model at a budget it was never trained to serve, the fault is in the objective, not the architecture.
The bottom line
If you serve one model across multiple quality tiers, adopt TLM-style stochastic prefix supervision, because one training run replaces N compression runs and shaves 43-44% off the area under your quality-budget curve.
If you're putting MoE into a looped or test-time-compute setting, flatten the expert layers and untie attention per pass, because at equal parameters and compute that's where the loss gains live: 0.012 nat at 100B tokens, monotonic in the degree of flattening, with downstream accuracy on par or better.
If you're scaling compute at decoding time, skip fixed-depth looping and use an adaptive iteration scheme like TaH2's decider, because uniform depth plateaus as max depth grows while adaptive depth keeps gaining, +2.8 points over the baseline at depth 2 to +3.9 at depth 8 on AIME.
One thing to watch: these three designs are about to merge. Expect a combined model within six months, continuum depth training plus per-token adaptive loop control on a flattened MoE block. The ablations already point at the composite.