Skip to content

The RL Post-Training Stack: Five Papers, One Pipeline

#reinforcement-learning #reasoning #post-training #agents #credit-assignment #distributed-training

RL post-training is where frontier reasoning ability gets made. Base checkpoints keep getting cheaper and more capable, but the last few points on MATH500 or GPQA come from reinforcement learning loops, not from pretraining. Run one of those loops and the pain shows up at every layer of the stack.

Credit assignment is murky. Sampling policy is guesswork. Checkpoints are too big to move between clusters. And whatever you learn rarely transfers to the next model family you fine-tune.

Five papers from the same October window attack that stack from five different layers. One proves, formally, why on-policy exploration beats offline reward modeling. Another fixes credit assignment for multi-turn agents. A third makes training-free sharpening adaptive per query. A fourth shows which parts of an RL update survive transfer across model families. And the last one makes trillion-parameter weight sync fast enough that agentic RL can scale at all. They read like five chapters of one story.

Why on-policy exploration wins ​

Start with the theory, because it changes how you read the other four results. The paper (2610.08561) models the reward as a hierarchical function on the response space: infinitely many local components, each of which only matters once everything before it is resolved. That is a good model of reasoning. You cannot earn the final answer reward on a geometry problem until you've set up the diagram, chosen the approach, and executed the algebra. Each gate only matters after the previous one succeeded.

Against that reward, a natural Transformer actor-critic loop, which alternates between sampling from the current KL-regularized policy, fitting a Transformer critic to observed rewards, and updating the policy, achieves minimax-optimal rates in the query budget and the regularization strength, up to log factors. In plain terms, no algorithm does better given the same number of queries. The KL regularization keeps the updated policy close to the base model, which is the technical version of saying the loop doesn't collapse into reward hacking.

The sharp contrast is offline reward modeling: sample from the fixed reference distribution, score, train. The paper proves regret decay there is stuck at a logarithmic rate. That matches what practitioners keep rediscovering empirically. Re-scoring old rollouts is a cheap RL loop, and it stalls after a few rounds. Fresh rollouts from the current policy progressively zoom in on the region where the reward is concentrated.

One loop, five papers ​

Here's the RL post-training loop, with each paper mapped to the stage it changes.

PaperLayerCore moveHeadline result
Hierarchical Reasoning Rewards (2610.08561)TheoryModels reward as hierarchical components, proves actor-critic ratesMinimax-optimal regret; fixed-reference sampling caps at logarithmic decay
VETTA (2610.08402)Agent credit assignmentTurn-level and token-level value heads on one shallow critic+37.5% ALFWorld, +22.3% WebShop over PPO with a 1.5B model
Adaptive Power Sampling (2610.08563)Inference samplingPer-query sharpening exponent from the self-reward gapBeats fixed-exponent power sampling on MATH500, HumanEval, GPQA
Selective-RL (2610.08659)Cross-model transferTransfers dominant directions of the RL-stage update onlyWins 12 of 15 comparisons vs full-update interpolation
NeMo-DCR (2610.08430)Systems / refitBit-exact XOR-mask delta refit with affine placement1T refit in 150 s instead of 87.5 min

NeMo-DCR sits on the weight-sync edge, the one people forget until a rollout cluster goes idle. Selective-RL changes what you do with the policy update after training. APS changes what you sample at rollout. VETTA changes how you credit the samples you got. The theory paper explains why the whole loop works at all.

Quick Take: On-policy exploration isn't a habit, it's a statistical guarantee. With hierarchically structured rewards, sampling from the current policy is the difference between minimax-optimal regret and logarithmic decay.

Credit assignment across turns and tokens ​

Multi-turn agents get sparse feedback: one final outcome for a whole episode of tool calls, web searches, and partial answers. But the model generates that episode one token at a time. So there are two credit-assignment questions stacked on top of each other. Which response in the trajectory earned the reward? And which generation decisions inside that response mattered?

Turn-level methods evaluate whole responses but can't tell a good response with one bad sentence from a bad response with one good sentence. Token-level methods can propagate feedback across turns but don't explicitly model which response deserves the credit. VETTA (2610.08402) learns both at once: separate value heads for turn-level and token-level credit on a shared lightweight critic, then combines each turn advantage with a within-response-centered token residual in the PPO update.

The critic only keeps the early Transformer blocks of the checkpoint that initialized the actor. That keeps value learning cheap. In my experience, critic efficiency is where agent RL projects quietly die; a full-size critic doubles the memory and compute of every rollout batch. A compact one changes the economics of the whole run.

The numbers are worth unpacking. With Qwen2.5-1.5B-Instruct, VETTA beats PPO by 37.5% on ALFWorld and 22.3% on WebShop, both relative success-rate improvements. A 1.5B model is small enough to fine-tune on a single node, so this is reproducible without a cluster. With the 7B model, it hits 95.5% on ALFWorld and 76.0% on WebShop. ALFWorld at 95.5% is close to saturated. The WebShop number is the one I'd watch; that benchmark still has headroom, and 76% with a 7B model is strong for a text-based shopping agent.

Sampling without training ​

Power sampling is the training-free trick for reasoning: sharpen the base model's output distribution and sample from the sharpened version instead of the raw softmax. The catch is that existing methods sharpen uniformly across queries. An easy arithmetic query and a hard multi-step proof get the same exponent.

That's wrong in both directions. Over-sharpen an easy query and you collapse the small probability mass that would still have produced the right answer. Under-sharpen a hard query and you never concentrate your samples where they matter.

Adaptive Power Sampling (2610.08563) adjusts the sharpening exponent per query at test time, using the relationship between answer agreement and the model's self-reward. The theory says the benefit of further sharpening is determined by the self-reward gap between correct and incorrect responses. Plainly: if the model scores the right answer clearly higher than the wrong ones, sharpening amplifies a real signal. If the gap is small, sharpening just concentrates noise.

When I've run fixed-exponent sharpening myself, the failure mode is predictable. Crank the exponent to save the hard problems and the easy ones start latching onto a single confidently wrong answer. APS consistently beats fixed-exponent power sampling across MATH500, HumanEval, and GPQA, and it costs nothing at train time.

What actually transfers when you merge ​

Model merging is the training-free way to port reasoning into a vision-language model: interpolate a reasoning-tuned LM checkpoint with a VLM and hope the skills carry over. The problem is that endpoint-based transfer conflates two different things. Differences between the two base models that existed before any RL, and changes actually acquired during reasoning post-training.

Selective-RL (2610.08659) isolates the RL-stage update and shows its components differ substantially in cross-model transferability. The dominant matrix-wise directions transfer better than the complete update. So the method keeps those dominant directions, preserves their magnitude, and transfers only them to the language modules of a VLM.

Across three model families and five visual-reasoning benchmarks, Selective-RL beats full-update interpolation in 12 of 15 comparisons, including an 8.55 percentage-point MathVision gain on the Qwen recipient. A jump that size is the difference between a model that fumbles visual geometry and one that consistently gets it. The controls matter too: update magnitude alone, or arbitrary low rank, does not reproduce the gains. The dominant directions are doing real work, and the distinction between what post-training acquires and what actually transfers is the paper's lasting contribution.

Trillion-parameter refits in two and a half minutes ​

Agentic RL disaggregates training from rollout. The training cluster updates the policy; the rollout cluster generates experience with the current weights. Every policy update has to reach the rollout cluster before the next batch can start. At frontier scale, that sync is the hidden tax.

A full 1T checkpoint between two AWS regions takes 87.5 minutes. That's an hour and a half of the rollout cluster idling, per policy step. Measurements of BF16 training show only about 1% of weights change their stored values per step. You're moving 100% of the weights to deliver 1% of them.

NeMo-DCR (2610.08430) sends only the changes and stays bit-exact: receivers get the same parameter and buffer bits as a dense refit. Prior delta systems cheated somewhere, on placement, exactness, or failure recovery. Placement is handled by fixed affine mappings that project changes from training shards into the checkpoint's canonical coordinates, with residual conversion covering the rest, and the serving runtime's native loader doing placement in receiver storage. Exactness is the interesting part: compressible XOR masks carry the affine changes whose projection and loader preserve stored bits, and overwrites carry the rest. Receivers apply both in place, retries overwrite partial writes, and a joint commit binds the policy to the baseline for the next delta. No cross-cluster collective; payloads stream over object storage or a relay tree while the delta is being built.

The numbers are the point. At 3% and 5% change rates, refits on 30B to 1T models run 12 to 40 times faster than a transport-only full-checkpoint reference. A 1T relay-tree refit at 3% takes 150 seconds instead of 87.5 minutes. That's the difference between the rollout cluster waiting through a long lunch break and waiting out a quick coffee.

Key numbers

  • 87.5 min: dense 1T checkpoint refit between two AWS regions, per policy step.
  • 150 s: NeMo-DCR relay-tree refit on a 1T model at a 3% change rate.
  • ~1%: fraction of BF16 weights whose stored values change per training step.
  • 37.5% / 22.3%: VETTA's success-rate improvement over PPO on ALFWorld and WebShop with Qwen2.5-1.5B.
  • 8.55 pp: Selective-RL's MathVision gain on the Qwen recipient over full-update interpolation.

Common Pitfalls ​

These papers are new, but the failure modes they fix are old. Here are the ones I keep hitting, now with names.

Merging endpoint checkpoints instead of the RL-stage update. I've interpolated final reasoning-tuned checkpoints into a VLM base and watched it inherit the base model's weak spots along with the new skills. Selective-RL says that happens because most of the endpoint delta is pre-existing model difference, not transferable acquisition. Isolate the training-stage update, then keep only the dominant directions.

Picking one credit-assignment level. I've run turn-level rewards on a shopping agent and watched it nail the right actions while executing them sloppily, because nothing rewarded token choices inside the response. Token-only had the reverse failure: local phrasing improved while whole turns drifted off-task. VETTA's two heads aren't an optimization. They're the structure of the problem.

Reconstructing deltas with arithmetic. Don't send raw delta tensors and have the receiver add them to the baseline. Floating-point addition is not associative, so after a few refits the training cluster and rollout cluster silently disagree on the weights. That's why NeMo-DCR's XOR masks preserve the exact stored bits. Bit-exactness is the design goal, not a nicety.

One sharpening exponent for every query. Fixed-exponent power sampling looks fine on a dev set of medium-difficulty prompts, then wrecks the easy and hard tails in production. The self-reward gap tells you per query whether sharpening will help. If the model can't distinguish its own correct and incorrect answers, exponent 2 and exponent 8 both just sharpen noise.

Re-scoring stale rollouts. Scoring old responses from the frozen base policy sounds like a free speedup. The theory says regret decay collapses to a logarithmic rate, and my experience matches: the loop stalls after a few rounds. Fresh on-policy samples cost more, and that cost is the point. You're buying the zoom-in property.

One Thing to Remember ​

The thread across all five papers is that post-training is where reasoning is actually earned, and every layer of the loop is currently wasting part of your budget. Sampling, credit assignment, weight sync, and transfer each have a measurable gap between what happens and what should happen. The fixes share a shape too: isolate the stage you're working on, exploit its structure, and stop moving information that doesn't need moving. Sample from the current policy, credit both turns and tokens, sharpen per query, transfer only dominant update directions, and sync only what changed, bit-exactly.

The Bottom Line ​

If you're training multi-turn agents, adopt dual-level credit assignment. VETTA-style turn and token heads on a shared shallow critic beat plain PPO by 22 to 37 percent on standard agent benchmarks, and the small critic keeps the run affordable. Single-level credit leaves a whole error class uncorrected.

If you're running agentic RL across regions with a model at or above ~100B parameters, switch to delta refits now. Dense sync at 87.5 minutes per step means your rollout cluster idles, your iteration cycle stretches, and every experiment takes twice as long as it should. A 150-second trillion-parameter refit makes weight sync the boring part it should be.

If you're doing training-free reasoning, move from fixed-exponent power sampling to per-query adaptivity. And watch the transfer results: within six months, dominant-direction RL update transfer will replace naive checkpoint merging for moving reasoning skills between model families.