Appearance
The claim that should make you question your RL budget
Reinforcement learning is the default answer for improving reasoning in LLMs. Every frontier lab runs GRPO or PPO over millions of rollouts, with reward models, verifiers, and weight-sync infrastructure humming in the background. A paper posted on arXiv in early May 2026 argues that most of that machinery is aimed at the wrong target.
"Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning" makes a sharp claim: RL does not teach reasoning models new strategies. It redistributes probability mass over solutions the base model already contains. The beneficial footprint is a sparse correction at 1-3% of token positions, and the authors replicate the gains without RL at roughly 1000x less compute.
If that holds, the reasoning post-training playbook changes. You might not need the RL loop at all. You might need a few hundred rollouts and a single GPU.
Key Numbers1-3%: token positions RL actually changes in a reasoning trace. On a 2,000-token trace, that's 20 to 60 positions. Top-5: the promoted token is always within the base model's top-5 alternatives. The model already knows the answer. ~1000x: reported training cost reduction with ReasonMaxxer, from thousands of GPU-hours to minutes on one GPU. Tens of problems: the entire training set, small enough to curate by hand in an afternoon.
What the token-level analysis found
Across multiple model families and RL algorithms, the pattern is consistent. RL's corrections concentrate at high-entropy decision points, the positions where the model is unsure which branch to take. Only 1-3% of token positions are affected. Everything else in the rollout is already fine.
The promoted token always lies within the base model's top-5 alternatives. The model knows the right answer. It just isn't confident enough to commit. RL bumps it up the ranking.
Two control results make the point. Targeted corrections at those few positions recover most of RL's accuracy gain. Random corrections fail. And the base model's own entropy identifies the positions without any RL-trained model. You don't need to run RL to find out where RL would help. You need a forward pass and an entropy threshold.
The whole correction is low-dimensional. You can represent it in a tiny fraction of model parameters. That's a strange thing to say about a process that looks like it's reshaping the entire policy.
ReasonMaxxer: RL-free post-training
The authors turn the finding into a method. ReasonMaxxer applies a contrastive loss only at entropy-gated decision points, using a few hundred base-model rollouts and no online generation. No policy rollouts during training. No reward model in the loop. No weight synchronization between a training cluster and a serving cluster.
Across three model families, six scales, and six math reasoning benchmarks, it matches or exceeds full RL performance. The training budget: tens of problems and minutes of single-GPU training.
Translate that into what you need. A few hundred rollouts fits on one node. Tens of problems is a dataset you can assemble in an afternoon. Minutes on a single GPU means a workstation with one RTX 4090, not a cluster allocation.
Quick Take: If the sparsity finding holds beyond math benchmarks, the reasoning post-training playbook shifts from "spend more on RL" to "find the few positions where the model is unsure, and nudge them."
Why the RL loop exists anyway
None of this means the RL loop is pointless. ReasonMaxxer still needs rollouts, just a few hundred instead of millions. And the evidence here is math reasoning, where the reward is a verifiable answer. The harder the reward signal gets, the less the sparsity story applies.
RL became the standard recipe for reasoning after DeepSeek-R1 showed that RL alone could elicit chain-of-thought behavior. The loop exists because reasoning improvements compound when the reward is rich: code that compiles, agent trajectories that finish tasks, tool calls that return the right result. Those tasks don't decompose into a few uncertain token positions as cleanly as math does. A full RL run at frontier scale means millions of rollouts, a verifier scoring every one, and a policy update that touches every token in every response. Most of that update, the authors argue, is wasted motion.
The table below sums up where each approach stands.
| Dimension | Full RL (GRPO/PPO) | ReasonMaxxer |
|---|---|---|
| What the update touches | All rollout tokens | 1-3% of positions, entropy-gated |
| Rollouts needed | Millions, online | A few hundred, offline |
| Training budget | Thousands of GPU-hours | Minutes on one GPU |
| Relative cost | ~1000x | 1x |
| Reward signal | Verifier or reward model in the loop | None, uses base-model entropy |
| Result on math benchmarks | Strong | Matches or exceeds RL |
| Core assumption | RL teaches new strategies | Base model already contains the answer |
The infrastructure reality: slime
If you're running RL at frontier scale, the loop is the product. THUDM's slime framework is built for exactly this. It connects Megatron for training with SGLang for rollout, and routes training, generation, reward computation, and environment interaction through the same data buffer. This is the loop that produced GLM-5.2, GLM-5.1, GLM-5, GLM-4.7, GLM-4.6, and GLM-4.5. That's release-grade post-training.
The design choices are telling. slime passes Megatron arguments through directly and exposes SGLang arguments with a --sglang- prefix, so upstream engine improvements stay usable without a wrapper layer. It supports delta weight sync for training/inference disaggregation, PD disaggregation for agentic workloads, and session-affinity routing for multi-turn agents. RL bugs are silent, so the project treats reproducibility, fault tolerance, tracing, and CI as first-class concerns. Beyond GLM, slime supports Qwen3.x, DeepSeek V3 and R1, and Llama 3.
The projects built on slime show where RL post-training is heading. Dressage from Alibaba Accio does agentic RL for blackbox agents. Relax from RedAI does omni-modal RL across text, vision, and audio. P1 is a physics reasoning model trained entirely through RL. RLVE scales RL across 400 verifiable environments. TritonForge trains models that write GPU kernels. APRIL attacks the rollout bottleneck, which is where over 90% of RL training time goes.
| Project | Builder | What it adds |
|---|---|---|
| Dressage | Alibaba Accio | Agentic RL for blackbox agents, token-level trajectory capture |
| Miles | RadixArk | Enterprise features: LoRA, TITO, low-precision training |
| vime | vLLM project | slime with a vLLM rollout backend |
| Relax | RedAI Infra | Omni-modal RL (text, vision, audio), fully-async training |
| P1 | Open source | Physics reasoning models trained entirely through RL |
| RLVE | Open source | 400 verifiable environments with adaptive difficulty |
This work asks whether the RL optimization loop is necessary. slime answers a different question: if you're going to run RL at scale, how do you keep the loop stable, debuggable, and fast for weeks? The teams building on slime are betting that rich reward signals and agentic tasks keep the loop alive.
What the community is saying
When this landed on Reddit, the reactions split along a predictable line. People doing math reasoning saw the 1-3% number and felt vindicated. I've burned more GPU-hours than I want to admit on RL runs that mostly re-ranked answers the base model already had. The entropy-gating result matches what I've seen in my own rollout logs: the model flips a coin at a handful of forks, and RL just biases the coin.
The agentic RL crowd was more skeptical, and I'm with them on this. My team ran into the rollout bottleneck the approach sidesteps. We spent more time on rollout stability, sandbox timeouts, and reward sparsity than on the policy update itself. For long-horizon agent tasks, the problem isn't which token to promote. It's that the reward arrives once at the end of a 5,000-token trajectory, and you can't decompose that into a few entropy-gated positions.
The most useful criticism I saw: the benchmarks are all math. Six math reasoning benchmarks, three model families, six scales. That's a solid base, but it's one domain. Until someone shows entropy-gated selection working for code generation or agentic control, the general claim stays provisional.
What trips people up
- Treating 1-3% as a universal constant. It's measured on math reasoning with verifiable rewards. Code and agentic RL likely have different sparsity profiles. Measure your own token-level footprint before restructuring your pipeline.
- Assuming "no RL" means "no rollouts." ReasonMaxxer still needs a few hundred base-model rollouts to find the entropy-gated positions. You still need a serving backend. What you don't need is the online loop.
- Using the wrong entropy signal. The gating works off the base model's own entropy. If you compute entropy from a fine-tuned or RL-trained policy, the positions shift and you correct the wrong forks. Use the base model.
- Ignoring the reward problem. For math, the reward is a string match against a known answer. For agentic tasks, the reward is sparse, delayed, or learned. Entropy gating can't fix a missing reward signal.
- Skipping the debug paths in RL frameworks. RL bugs are silent. Weight sync issues and numerical precision problems don't crash; they quietly corrupt the policy. Run the rollout-then-train replay path before a long run. It's boring, and it catches the failures that waste weeks.
One thing to remember
RL for reasoning is looking less like a training algorithm and more like a selection mechanism. The compute is in the loop, but the value is in a few token positions. Whether you exploit that with ReasonMaxxer or fold entropy-gated loss masking into your existing RL framework, the cheapest reasoning gains are probably sitting in the uncertainty of the model you already have.
The Bottom Line
If you're doing math reasoning post-training on a tight budget, try entropy-gated contrastive loss before spinning up an RL stack. ReasonMaxxer matches full RL in minutes on one GPU, and the worst case is you've spent an afternoon before falling back to RL.
If you're training frontier-scale models or doing agentic and code RL, keep the heavy infrastructure. slime-style loops exist because rollout serving, weight sync, and reward computation are the hard parts at scale, and the sparsity finding doesn't remove them.
One thing to watch: entropy-gated selection getting folded into RL frameworks as a targeted loss mask inside the loop. That gives you RL's exploration benefits with ReasonMaxxer's compute profile. Expect it within six months, probably as a flag in slime, veRL, or OpenRLHF.