Appearance
The same reasoning, three different hiding places
Three recent papers ask the same question from different directions: when a model produces a reasoning trace, where does the behavior actually live? The answers all push back on the popular story that RL or chain-of-thought fine-tuning installs new skills.
Paper one shows that a base model with a fixed starting-token cue matches its RL-trained counterpart on math and coding. Paper two trains small transformers to multiply and finds that a space token between digits changes exact-match accuracy from 1% to 89%, then shows the model encodes the carry-in as an angle on a ring in the residual stream. Paper three routes reasoning through a continuous latent space and hits 97% of explicit chain-of-thought accuracy at a quarter of the latency.
The thread connecting them is that reasoning behaves like a routing problem. The skill is already in the weights. You're steering the model into the region where it lives, and the three papers turn different knobs: starting tokens, input formatting, and latent trajectories.
Key numbers
- 36 points: the MATH-500 pass@1 jump for Olmo-3-7B when you fix the starting cue ".\n\nOkay" (42% to 78%). No RL, no fine-tuning.
- 89%: exact-match accuracy on 4x4 multiplication when a space token separates digits. The same setup scores 1% without it.
- 97%: FLaRe's accuracy relative to explicit chain of thought, at roughly a quarter of the latency.
- 84%: of carry states that transfer to a target example by activation-patching one prediction slot after block 4.
A starting token worth thirty-six points
The first paper's setup is almost insultingly simple. Take a base model, don't fine-tune it, and pin the first tokens of its response. On MATH-500 pass@1, Olmo-3-7B goes from 42% to 78% just by starting with ".\n\nOkay". Qwen3-14B goes from 72% to 87% with "Alright,".
| Model | Cue | Without cue | With cue | What it means in practice |
|---|---|---|---|---|
| Olmo-3-7B | ".\n\nOkay" | 42% | 78% | Same weights, same latency, 36 extra points |
| Qwen3-14B | "Alright," | 72% | 87% | Stronger base, still a 15-point cue bonus |
This isn't ordinary prompt engineering. The prompt is identical in both conditions; only the model's own first output tokens are pinned. And the effect is selective. A cue that fires reasoning on math doesn't necessarily change refusal behavior, and the paper shows that cues steering compliance versus refusal trace back to different kinds of training documents.
The paper also explains part of why RL looks as powerful as it does. RL training makes these cues more likely to appear. When you fix the cue yourself, you recover much of the RL performance gap over the base model. A nontrivial slice of what gets credited to RL is really a shift in the start-of-response distribution.
Quick Take: Base models already know how to reason. The cue opens the hidden state that contains it, and most inference-time tricks work by making that state more likely.
The cue is a key into the training data
The strangest result is that the effect is causal, not merely associative. The authors perform data interventions: add or remove training documents that bind an arbitrary word to reasoning, and that word becomes, or stops being, an effective reasoning cue. They turn "chicken" into a cue that works. "Think duck duck goose" becomes as strong as "Think step by step" at eliciting a reasoning trace.
Hidden-state analysis backs this up. The representations induced by different cues correlate with different document types in the training set. The cue is a retrieval key. It opens the region of the residual stream that was shaped by math-solution documents during pretraining, and the reasoning behavior follows from whatever lives in that region.
The safety case study is the sharpest demonstration. Different cues elicit distinct refusal and compliance behaviors, and those behaviors correspond to the types of training data each cue associates with. Same model, same question, flipped safety behavior, all because the first token selected a different internal region.
Space tokens and the geometry of carry
The second paper comes at the problem from a different direction: where does arithmetic actually happen in the model, and what makes it learnable? The authors train small Llama-style transformers from scratch on 4x4 digit multiplication, no chain of thought, and find the input format decides everything.
Inserting a space token between digits raises exact-match accuracy from 1% to 89%. One character. The model wasn't failing because it lacked capacity; the default digit-token format made carry propagation nearly impossible to learn. Format is a learnability switch.
Two more findings stand out. Output positions are learned in carry-chain order, with the middle digits (the longest-range dependencies) learned last. And the model does explicit arithmetic in a geometric sense: the separator token that predicts each digit encodes the carry-in as an angle on a ring in the residual stream. Examples with more distinct carry values fill more of the ring.
Activation patching confirms the model actually uses this state. Match examples on the column sum, patch the prediction slot alone, and the source carry transfers in up to 84% of cases after block 4 for one middle column of the best model. Other columns assemble the carry at the neighboring answer slot before passing it along. When the model errs, it's almost always off by one, which is exactly what a small error on the carry or on the circular digit code would produce. A slightly rotated angle lands on the neighbor digit.
That off-by-one signature is the kind of thing you learn to look for in interpretability work. It's too structured for random failure.
Latent thoughts that meet five tests
The third paper steps back and asks what latent reasoning should do at all. It proposes five requirements for a continuous-space thought: useful (it produces the correct answer), diverse (resampling gives different trajectories), explainable (a decoded chain of thought reflects the reasoning the answer actually follows), refinable (more inference compute helps), and efficient (cheaper than explicit CoT at comparable accuracy).
Most current latent methods fail these tests in predictable ways. They learn shortcuts straight from the question, distill the explicit CoT into their weights, or imitate the CoT one token at a time. The result is a decoded chain that sounds right while the answer came from somewhere else.
FLaRe is the paper's fix: a recipe covering what the latent space encodes, how to shape it, where to train the flow, how to read out the answer, and a final stage trained on the model's own verified thoughts. A probe for each requirement shows FLaRe improves on prior latent methods in all five.
| Requirement | Common failure in prior methods | FLaRe approach |
|---|---|---|
| Useful | Shortcut from the question; answer decoupled from reasoning | Flow trained in a space shaped by verified thoughts |
| Diverse | One-token imitation of a fixed CoT | Resampling latent trajectories yields different paths |
| Explainable | Decoded CoT looks plausible but the answer ignores it | Decoded CoT reflects the trajectory the answer follows |
| Refinable | No compute/accuracy trade-off | Longer flow trajectories improve accuracy |
| Efficient | Often costs as much as explicit CoT | 97% of explicit CoT accuracy at roughly a quarter of the latency |
That last row matters operationally. If your serving cost is dominated by long CoT generation, explicit reasoning is effectively 4 to 8x more output tokens per request. Latent reasoning collapses that to a fixed-size trajectory plus a short verbalized answer, landing within about 3% of explicit CoT accuracy. The catch: you can only trust explainability if you measure it, and most teams measure final accuracy only.
Common pitfalls
Don't strip the start token. Serving stacks that truncate or reformat the beginning of a base model's response can silently destroy tens of points. The first tokens are a hidden-state selector.
Don't chase magic tokens on your own eval without checking training-data alignment. "Alright," works for Qwen3-14B because of what its pretraining data contains. If your domain differs, the cue won't fire. The paper's intervention method gives you a direct test: bind your word of choice to reasoning documents and check whether the effect appears.
For structured tasks, check the tokenizer before buying a bigger model. The 1% to 89% multiplication jump came from a single space character. Number formatting dominates architecture choice at this scale.
Don't evaluate latent reasoning on final accuracy alone. Distilled and imitation methods can hit decent numbers while their decoded chains are unfaithful. If the trace doesn't reflect the answer's actual path, you've built a model that lies plausibly under inspection.
When comparing base versus RL-tuned models, fix the starting tokens first. Part of the RL gain is cue frequency. Ablate it before you spend a training budget on something a free token can do.
One thing to remember
All three papers point the same way: measure the cheap knobs before investing in expensive ones. A starting token, a space character, and a latent trajectory each moved results more than many training budget increases would. The mechanisms differ, but the lesson is shared: the models already contain the behavior. The job is steering into it. Measure the cheap knobs first.
The bottom line
If you're serving a base model or benchmarking it against an RL-tuned one, fix the starting-token cue before touching anything else. It's free: 36 points on MATH-500 for Olmo-3-7B, and fixing the cue recovers much of the apparent RL gap.
If you're building arithmetic or structured-data pipelines, add separator tokens and audit your number formatting. A space token turned 4x4 multiplication from a 1% task into an 89% task, and the resulting model computes carries in a geometric code you can inspect.
If latency is your constraint, FLaRe-style latent reasoning gets you 97% of explicit chain-of-thought accuracy at a quarter of the latency. One thing to watch: shortcut-trained methods fail the explainability test, so expect evaluation suites that probe whether decoded thoughts match the answer's true path to become standard within the next year.