Appearance
The post-training efficiency problem
Full-parameter fine-tuning of a 13B model with AdamW spends more memory on optimizer state than on the weights themselves. The standard fix, 8-bit quantization, gets you to 27.7 GB of persistent state. That's still enough to force smaller batches, shorter sequences, or a smaller model than you'd like. And memory is only half the story. Reinforcement learning post-training is producing models that drift into strange languages inside their own chain of thought, and the masking rules meant to keep off-policy rollouts honest can cancel out real signal.
Five papers posted this month attack different stages of the same problem. A new optimizer that cuts state to 0.16 GB. A proof that language drift is baked into RLVR optimization. A masking rule that stops opposing probability changes from erasing each other. A post-training normalization pass for LoRA. And a sampling method that makes SFT behave like on-policy RL. Together they map where post-training actually bleeds resources.
TACO: optimizer state down to 0.16 GB
The TACO paper starts with a specific complaint about Muon. Muon cut optimizer memory by using matrix-valued updates, but its geometry doesn't match AdamW. Fine-tuning AdamW-pretrained models with it degrades performance. TACO keeps Muon's operator-norm steepest-descent view and pushes the geometry further: for each column of a 2D weight matrix, it keeps the sign of the largest-magnitude entry. That single value is the exact steepest-descent direction under a dimension-normalized 1-to-1 operator norm, and the persistent state it needs is nearly nothing.
On OPT-13B, optimizer state drops from 27.7 GB with AdamW-8bit to 0.16 GB, and peak training memory from 80.6 GB to 27.5 GB. Accuracy and runtime stay comparable. 27.5 GB of peak memory means the optimizer stops being the bottleneck on an 80 GB H100; activations take over from there, and those have known fixes like gradient checkpointing and sequence packing.
TACO in numbers: 174x less persistent optimizer state than AdamW-8bit on OPT-13B: 27.7 GB down to 0.16 GB 2.9x lower peak training memory: 80.6 GB down to 27.5 GB Full-parameter fine-tuning of 30-32B models on a single 80 GB H100
The practical consequence: TACO's memory savings are only useful if you spend them. Keep the same batch size and you've just made your training run cheaper for no reason. Push batch size or sequence length instead.
Language drift is a reward problem
The drift paper is the one to read if you're responsible for monitoring chains of thought. Language drift in RLVR-trained models has been observed for a while but not explained. This paper pins down three things. First, a proof that RLVR's optimization pressure permits unbounded language drift, while SFT does not. Second, empirical evidence that drift appears specifically when the target behavior can't be drawn out of the base model. Third, a proof of impossibility: you cannot constrain drift without constraining expected reward.
Read that last result twice. If you're RLVR-training at the frontier and need monitorable CoTs for safety review, the two goals are in direct tension. Not a tuning problem. Not a prompt problem. A geometric fact about the objective. If drift is unacceptable, either accept a reward ceiling or use a different post-training objective, and decide which before training starts, not after.
CARM: when masks cancel real signal
Sequence-level masking decides whether an entire sampled response contributes to optimization when rollout and training engines drift apart. The common rule takes the length-normalized geometric mean of sampled token probability ratios. The problem is in the signs. A token that becomes twice as likely paired with a token that becomes half as likely sums to zero. The policy has drifted substantially in both directions and the mask reports "fine."
CARM changes one thing: take the absolute value of each token log-ratio before averaging. The paper proves that accepted responses satisfy a joint bound on the fraction of token ratios outside a prescribed band and their mean log-distance beyond it. Empirically that's worth up to 3.13 percentage points on mean@16 across AIME 2024/2025/2026 and BeyondAIME, and 2.88 points on pass@1 across four code benchmarks. For a one-line change to an existing mask formula, that's a lot of recovered signal.
Quick Take: across these five papers, post-training efficiency comes less from clever loss functions and more from how memory, data, and credit are allocated.
LoRA-Norm: rebalance adapters after training
LoRA lets you specialize a model cheaply, but the learned update isn't balanced. The LoRA-Norm paper calls this adaptation imbalance: a few singular directions dominate the trained update, so final performance is sensitive to how gains are allocated across directions. Learning where to adapt doesn't guarantee the gains are well balanced once you get there.
I've watched a fine-tuned assistant lose its coding ability after a LoRA run. The adapter looked great on the target task and silently degraded everything else. LoRA-Norm is the kind of fix you'd want in that situation: keep the learned directions, rebalance their singular values with a fixed nonlinear transformation, then restore the original total spectral mass with a nuclear-norm restoration step. No calibration data, no extra training, no inference overhead. Across two backbones and three adaptation tasks it improves both specialization and capability retention, beating post-hoc spectral pruning and gradient-guided editing.
This is a one-time pass you run after training the adapter and before deploying. It's cheap. Use it whenever a fine-tuned model loses general capability the base model had.
MCMC sampling: SFT learns better than you think
The SFT vs RL debate gets a useful twist. Conventional wisdom says RL generalizes on new tasks without losing old capabilities, while SFT overfits narrow distributions and forgets. But RL has a hidden cost: the model has to repeatedly sample successful trajectories, and if it can't find them, you're stuck. This paper asks what happens if you keep SFT's off-policy data but reshape its distribution to be more on-policy.
The answer is an MCMC sampling algorithm that progressively transforms off-policy traces into on-policy-like examples, given a reference model. SFT trained on this reshaped data rivals RL across scientific skill acquisition, mathematical reasoning, and open-ended expertise, while generalizing better and forgetting less than strong on-policy baselines. The framing shift is the interesting part: sampling as a data-shaping operator rather than an exploration mechanism.
One stack, five interventions
Each paper sits at a different point in the post-training stack, and none of them require a new loss function.
| Method | Stage | Problem it attacks | Extra cost | Headline result |
|---|---|---|---|---|
| TACO | Optimizer | Optimizer state memory | Near-zero state, same runtime | 174x state reduction, full FT of 30-32B on one H100 |
| CARM | RL response mask | Off-policy signal cancellation | No extra training | Up to +3.13 mean@16, +2.88 pass@1 |
| MCMC sampling | SFT data pipeline | Off-policy data mismatch | One sampling pass over traces | SFT rivals on-policy RL |
| LoRA-Norm | Post-LoRA | Adapter gain imbalance | One-time post-processing | Better specialization and retention |
| Drift analysis | RLVR | CoT monitorability | Theory, no compute | Drift unbounded under RLVR, unfixable without reward loss |
Common pitfalls
Swapping in a memory-efficient optimizer but leaving batch and sequence settings unchanged. The 2.9x peak memory reduction only matters if you spend the headroom. Keep the old settings and you've optimized for nothing; the next wall is activations, and that needs checkpointing or sequence packing, not another optimizer.
Trusting geometric-mean masks when rollout and training engines differ. I found that a response can pass the mask while containing tokens far outside the acceptable probability band, because opposing signed log-ratios cancel inside the mean. The absolute-value rule in CARM is a one-line change that removes a whole class of silent failures.
Treating language drift as a decoding problem. If you're seeing drifted CoTs, a temperature tweak or a blocked-words list won't fix it. The drift result shows this is an optimization-side constraint, so any mitigation has to enter the objective itself.
Pruning or editing LoRA matrices blindly. Post-hoc spectral pruning hurts more than LoRA-Norm's rebalancing, because the learned directions aren't the problem, the gain allocation is. Inspect the singular value spectrum before you touch anything.
Assuming off-policy data has to stay off-policy. The MCMC sampling paper shows a distribution-alignment pass gets most of RL's generalization benefit without the rollout infrastructure, which changes the cost calculation for adding new capabilities.
One thing to remember
The loss function is the least interesting part of post-training. TACO changes the optimizer geometry. CARM changes the mask. MCMC sampling changes the data distribution. LoRA-Norm changes the adapter after training. The drift paper changes what you should measure. Every one of them moves the same needle as a training trick, with less memory, less compute, or more stability.
The bottom line
If you're fine-tuning models that barely fit on your GPU, replace AdamW-8bit with TACO. The 2.9x peak memory drop buys batch size or sequence length, and you don't pay an accuracy or runtime penalty for it.
If you need monitorable chain-of-thought, budget for the tradeoff before RLVR training. The drift paper proves you can't constrain drift without constraining expected reward, so decide which one you're optimizing first.
If you're adding a new capability and don't have rollout infrastructure, run the MCMC sampling approach on your SFT data before building an RL pipeline. You'll get most of the on-policy generalization at SFT compute cost, and you can keep the RL loop for the cases where sampling alone doesn't reach the target behavior.