Appearance
The post-training bottleneck
Pre-training gets the headlines, but post-training is where the model becomes usable. It's also where the compute bill quietly doubles: data selection over millions of candidate examples, supervised fine-tuning, distillation runs, and reinforcement fine-tuning (RFT) on agentic tasks. Each step has its own hidden tax.
Five papers posted to arXiv this month attack those taxes from different angles. One cuts the FLOPs of gradient-based data selection by an order of magnitude. Another proves that your choice of divergence is secretly your entropy policy in distillation. A third replaces the repeated rollouts of GRPO-style RL with an ordinal filter over a replay buffer. The fourth shows that a Metropolis-Hastings correction dominates importance resampling when composing reward-specific policies at decoding time. The fifth checks whether efficient reasoning training destroys the things that make chain-of-thought useful for oversight.
The thread connecting all five: post-training costs are not fixed properties of your model. They're choices. And the papers give you better choices.
Data selection without the backward pass
Gradient-based data selection is the standard way to pick post-training examples. You rank each candidate by how well its gradient aligns with a small validation set, then keep the top fraction. It works. It also requires a full backward pass on every candidate sample, and when your candidate pool is in the millions, that's a bill you feel.
LESSER is a drop-in wrapper around these selection methods that uses only output-layer gradients. The trick: you can extract those from the forward pass alone, no backprop needed. Feature-extraction FLOPs drop by 9.7x for SFT benchmarks and 3.0x for RL benchmarks. If a selection pass over your SFT pool used to take a day, that becomes roughly two and a half hours.
The paper reports that even when output-layer and full gradients rank individual samples differently, the batches they select stay aligned on downstream task performance. The ranking disagrees; the selection agrees. For a method that costs a tenth of the compute, that's the result that matters.
The practical read: if you had avoided gradient-based selection because of cost, the cost argument just weakened considerably. The wrapper is drop-in, so you can bolt it onto existing selection pipelines without changing the downstream flow.
Key numbers
- 9.7x: FLOP cut for SFT data selection; 3.0x for RL (LESSER).
- Every convex f-divergence: MH output is closer to the target than SIR at the same rollout budget.
- Sokoban and Search-R1: FTW matches GRPO and PPO with no value model and no group rollouts.
- Forward KL inflation of student entropy: proven, and verified in pretraining and SFT.
Divergence is the entropy dial
Distillation is a core primitive now, but its dynamics are poorly understood. Most people treat the divergence as a fixed implementation detail. The paper "Divergence controls entropy in distillation" argues it's actually the main control knob for student behavior.
The headline result: forward KL inflates the student's entropy above the teacher's. Cross-entropy training is a special case of this, so the identity shows up everywhere, and the authors verify it quantitatively in pretraining and supervised fine-tuning. The practical effect: train with forward KL and the student stays more uncertain than its teacher. For calibration-sensitive applications, that's a feature. If you wanted a sharp, confident student, you've picked the wrong divergence.
Reverse KL carries no such guarantee. It deflates student entropy, until the gap between student and teacher gets large enough that the dynamics flip. And interpolating between the two changes entropy smoothly early in training but abruptly at convergence.
One detail is easy to miss: the lower entropy of on-policy distillation comes from token-level reverse KL, not from on-policy sampling. If you've been explaining your student's sharpness as a sampling artifact, the divergence is doing that work. The clearest case is self-distillation, where conditioning on privileged information deflates entropy, and the divergence hyperparameters that work best are the ones that compensate for it.
So when you set up your next distillation run, ask what entropy behavior you want, then pick the divergence accordingly. It's a temperature knob that most pipelines leave on a default.
Quick Take: All five papers trade a familiar post-training expense for a cheaper one, and the guarantees survive the swap.
Critic-free RL for environments you can't re-roll
GRPO-style reinforcement fine-tuning is the default for critic-free RFT in agentic LLMs. It computes a group baseline over repeated rollouts of the same prompt to reduce target variance. That setup breaks in stateful environments. When your policy is acting in a live service or a security sandbox, you can't reset the world and roll out again. And with long, sparsely verified trajectories, aggressive updates entrench the noise from the few rewards you actually receive.
I've hit exactly this wall: the first thing to go when you move from benchmarks to a real environment is the ability to re-roll, and with it the variance reduction you were relying on.
Follow the Winners (FTW) replaces group rollouts with an ordinal filter on replay-buffer samples, adapted from the cross-entropy method. Select the top fraction of buffered trajectories by return, update toward those, and you get polynomial concentration in the order statistic of returns. The derivation goes through a control-as-inference lens that recovers GRPO and DPO as specific modeling choices, which is useful for positioning: GRPO is risk-neutral, while DPO and FTW share a bounded risk-seeking offset. FTW exposes that offset as a knob, and the paper identifies it as the inherent price of variance reduction through ordinal filters. A critic model buys you a different trade along the bias-variance axis, but only if you can afford training one.
| Method | Baseline | Critic model | Risk attitude | Rollout requirement |
|---|---|---|---|---|
| GRPO | group mean over rollouts | no | risk-neutral | repeated rollouts per prompt |
| PPO | learned value function | yes | set by the critic's bias-variance trade | on-policy rollouts |
| DPO | implicit, offline | no | bounded risk-seeking | none (offline preference data) |
| FTW | ordinal filter on replay buffer | no | bounded, tunable risk-seeking | single-pass rollouts into a buffer |
Scaled to agentic LLM post-training, FTW matches GRPO and PPO on Sokoban and Search-R1. That's a strong result: parity with both a group-baseline method and a value-model method, achieved with CPU memory instead of a critic or repeated rollouts. For agentic training in production-ish settings, that trade is the one you actually want. Your GPU budget goes to rollouts of the real policy, and the replay buffer does the variance reduction.
Composing policies at decoding time
Multi-reward post-training has a cost problem: every trade-off between rewards wants a retrain. Decoding-time policy composition is the alternative. You train one policy per reward, then combine them at inference. The target distribution is a weighted product of the policies' probabilities over complete responses. The problem is that standard implementations combine next-token probabilities, and that introduces sampling bias, because a product over tokens is not the same as a product over complete responses.
An iterative correction based on independence Metropolis-Hastings (MH) has been floating around for this. The paper's main result is a dominance theorem: for every rollout budget, MH produces an output distribution at least as close to the target as sampling-importance-resampling (SIR) with the same budget, measured by every convex f-divergence. SIR is the natural baseline here, the thing most people reach for when they want to correct composition bias. The theorem says don't.
There are two asymptotic regimes worth knowing. When the reward-specific policies approach agreement, the correction's sampling error shrinks. When the log ratio between the target and the uncorrected decoder fluctuates widely, which happens for long responses, the error grows. If you're composing over long generations, the correction isn't optional; it's the difference between sampling from something near the target and something far from it. The paper also derives a lower bound on MH's improvement over the uncorrected decoder in a consensus objective measuring agreement with the supplied policies, so you get a guarantee in the regime where you'd previously have crossed your fingers.
The decoding-time overhead is per-sample and cheap. If you already have reward-specific policies in production, swapping SIR for MH is closer to a drop-in change than a redesign.
Length pressure and chain-of-thought faithfulness
Efficient reasoning training is appealing for one simple reason: CoT inference is expensive, and shorter reasoning costs less. But every efficiency method applies length pressure differently, and people worry the model compensates by skipping steps, making the CoT a post-hoc story rather than a trace of the decision.
The paper fine-tunes models with three methods: a fixed generation budget, a per-example length target, and a group-relative length reward. Then it measures two properties separately, which turns out to matter.
Faithfulness, how well the CoT reflects the model's decision on related inputs, falls in most settings. The paper attributes this mainly to consistency: the trained models are less consistent, so the reasoning is a worse predictor of behavior. Monitorability, whether the CoT reveals when input interventions alter the output, is more robust. The models keep acknowledging the influence on their answer even when the CoT is much shorter.
Those two findings should be kept separate, because they support different operational decisions. If you're using CoT for debugging individual wrong answers, the faithfulness drop is a real cost; the reasoning trace will mislead you sometimes. If you're using interventions to audit behavior, monitorability surviving means that oversight signal mostly persists even with aggressive length pressure. Efficient reasoning training doesn't uniformly destroy what makes CoT useful. It degrades one property and preserves another.
Common pitfalls
Full-parameter gradients when you don't need them. If your data-selection pool is large, output-layer gradients do the same batch-level job at 9.7x lower FLOP cost. Running the expensive backward pass on every candidate is the default, and it shouldn't be.
Picking a distillation divergence without an entropy target. Forward KL inflates student entropy; reverse KL deflates it. If you picked forward KL because cross-entropy uses it, you've decided your student should be more uncertain than the teacher. Make that call explicitly.
Group baselines in stateful environments. GRPO-style repeated rollouts don't exist in live services or security sandboxes. If you can't re-roll, an ordinal filter on a replay buffer is the variance-reduction mechanism that still works. I learned this one the hard way. In a live rollout you get one shot, and variance reduction has to work with what you've already collected.
Averaging next-token probabilities for composition. The target is a product over complete responses, not per-token. Uncorrected composition biases your samples, and the bias grows with response length. The MH correction dominates SIR at equal budget, so there's no good reason to settle for the biased version.
Assuming faithfulness and monitorability fall together. They don't. Length pressure drops faithfulness while monitorability mostly holds. If you audit through interventions, the signal survives; if you debug through reasoning traces, it won't be as reliable. Conflating the two leads to either paranoia or complacency.
One thing to remember
Each of these papers attacks a cost that practitioners have been treating as fixed. Backward passes for data selection, group rollouts for RL variance, retraining per reward trade-off, and the divergence hyperparameter you set once and forget. The recurring pattern is that a cheaper substitute exists, and in most cases it comes with a theorem or a benchmark tying it to the expensive version. Post-training cost is not destiny. It's a set of choices, and the field is rapidly converting each choice into a well-characterized trade-off.
The bottom line
If you're selecting post-training data at scale, adopt output-layer gradient selection. It's a drop-in wrapper, cuts feature-extraction FLOPs 9.7x on SFT and 3.0x on RL, and the batches it selects stay aligned with full-gradient selection on downstream tasks.
If you're distilling or composing policies, treat divergence and sampling correction as your primary controls. Forward KL for a calibrated student, reverse KL for a sharp one, and when you're composing reward policies at decode time, use the MH correction over SIR: it's provably closer to the target at every rollout budget.
If you're doing critic-free RL in an environment where repeated rollouts are impossible, switch to an ordinal filter on a replay buffer. FTW matches GRPO and PPO on Sokoban and Search-R1 without a value model, and that's the trade you want when irreversibility is the constraint.