Appearance
The learning signal problem
Most RL failures aren't failures of optimization. They're failures of the learning signal itself.
Policy-gradient methods sit at the center of modern RL, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment, and action-sampling noise. Notice what those three have in common. Classification has none of them. A classifier is a policy whose expected reward, its expected accuracy, is the probability it assigns to the correct label. And because the label is known, the policy gradient is exact and smooth. No exploration, no credit assignment, no sampling noise.
Yet exact policy gradient still loses to cross-entropy, even when both optimize expected accuracy. That asymmetry is the thread that ties together five new arXiv papers I've been reading. Each attacks a different slice of the same underlying problem: how do you get a clean learning signal into a sequential decision-making agent? One fixes the myopia of policy gradients. One traces real data dependencies through terminal sessions. One puts convergence rates on success conditioning. One asks whether phase-structured problems want one policy or many. One grows exploration from both ends of the task space.
Read together, they're less like five separate results and more like a coordinated assault on sample efficiency.
Cross-entropy is patient accuracy
Planning to Learn opens with an almost cheeky move: it uses classification to explain why policy gradients underperform, then fixes them with a one-line change.
The argument goes like this. The exact policy gradient values an update only by what it buys right now. But every update also sets where the next one starts. So an update's true value depends on how much learning remains. The exact gradient is the zero-horizon limit of that value. It's myopic. Cross-entropy, viewed this way, is patient accuracy: the total error an example would pay if its log-odds rose at unit speed forever.
The contribution is the horizon loss, which truncates that total at the learning that remains, a one-line change that slides from cross-entropy toward exact policy gradient as training runs out of steam. In a simple allocation model, the authors prove it escapes the trap that catches each endpoint. Empirically, it improves top-1 accuracy over cross-entropy on MNIST and on ImageNet with ResNet-50, ResNet-101, and ViT-S/16 at a flat learning rate. Flat is the operative word: no schedule tuning, no warmup gymnastics. The gain grows as label noise increases, which is how real data behaves, not how benchmarks pretend it looks.
The implication for RL is uncomfortable. If your policy gradient is exact but training stalls, the problem may not be exploration or variance. The update is optimizing the current step while ignoring how much learning the rest of the run has left to do. That's not something more compute fixes.
Credit assignment that follows the data flow
Credit assignment gets harder when your agent lives in a terminal. In coding and debugging tasks, later commands usually depend on intermediate results produced by earlier ones. A compile error early in a session poisons everything that follows. Existing trajectory-level and step-level credit assignment methods don't trace the read-write dependencies through which commands actually affect the final outcome. The training signal still flows to irrelevant operations.
This is the failure I keep hitting when I train terminal agents: the agent does real work, but the gradient washes over every command roughly equally, and the operations that mattered get diluted by the noise. DepGPO, short for Dependency-Aware Group Policy Optimization, is built specifically for this. It constructs a command dependency graph from execution traces, then traces backward from the resources inspected by the task verifier. Credit goes to the relevant writes and their supporting reads along those paths, and trajectory advantages get redistributed across steps on that basis.
The results on complex terminal tasks are better task performance and measurably steadier training. The principle is almost embarrassingly simple: if a command never contributed to the output the verifier checked, it shouldn't share the reward.
Quick Take: Every result here improves RL by reshaping the learning signal, not by throwing compute at the symptoms.
Success conditioning converges, and now we know how fast
Success conditioning is one of those techniques everyone uses and nobody can fully justify. The idea is simple: update the policy by increasing the probability of actions that yielded successful outcomes. It shows up all over applied RL, often as the implicit filter behind "only train on the good rollouts." But its limiting behavior and convergence rates were not well understood. That's awkward for a strategy people stake training runs on.
This paper closes the gap. It proves success conditioning converges to an optimal policy on a broad class of Markov decision processes. Then it gets specific. For discounted MDPs, you get an ε-optimal policy within O(1/ε^p) iterations, where the exponent p depends on problem data. That's polynomial in 1/ε, not exponential. If you want error 0.1 instead of 0.2, you pay a constant factor in iterations, not ten times as many. For single-period MDPs, the rate is O(log(1/ε)), which means you can slice an order of magnitude off ε for a constant number of extra iterations.
The practical read: if you're using success as a filter in goal-conditioned manipulation or anywhere rewards are sparse, there's now a convergence guarantee behind that heuristic, and the rates tell you what to expect before you start a long run.
One policy or many for phase-structured worlds?
Many RL problems are non-stationary but structured. A robot that walks, then climbs, then grips has three phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the standard move is to augment the state with phase information and run one policy across everything. But prior work kept finding that separate policies per phase can beat that shared policy, and nobody had a clean explanation.
This paper gives you the theory and the decision rule. First, the theory: a shared, state-augmented policy can theoretically achieve the performance of any multi-policy solution. So the gap isn't representational. It comes from function approximation, the learning and optimization processes, and, for multi-policy approaches, sample efficiency and loss of continuity between policies. In practice, those costs land unevenly.
The authors propose a regime-based phase decomposition: compare the duration of transient system dynamics, the settling time after a phase switch, against the duration of the quasi-stationary period. That ratio tells you which policy structure will win. Their experiments validate four hypotheses. Longer phase durations favor multi-policies. Heterogeneity between phases burdens a single policy. Multi-policies need enough data per phase to be worth the split. And the transition dynamics between phases can flip the answer.
If you're deciding between one agent and a mixture, measure how long your system takes to settle after a phase change. If the transient is short relative to the phase length, separate policies earn their keep. If phases flicker, a shared policy amortizes better.
Explore from both ends
Sparse-reward, long-horizon tasks have an exploration problem with a geometry. A policy starting from the initial state rarely reaches the goal, so it receives no learning signal at all. The standard escapes are reference motions, hand-designed curricula, and shaped rewards, which all require demonstrations or task-specific engineering. Automatic start-state and goal curricula avoid that requirement, but they expand from one side only. The policy has to cover the full distance from that side, and the middle stays dark far too long.
BVER, the Bidirectional Voronoi-biased Exploration curriculum, borrows the trick from bidirectional RRT planning. It grows start states outward from the goal and goals outward from the initial-state distribution. Both get biased toward unexplored task space, and they get steered toward each other, while one goal-conditioned policy trains on both sets. On point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer, BVER learns faster than every reference-free curriculum it was compared against.
The numbers on box climbing are the ones that stick. To reach 95% success on a 0.4 m box, BVER needs roughly 65% fewer iterations than the best alternative. That's the difference between a training run that finishes overnight and one that eats the whole week. It's also the only reference-free method that learns to climb a 0.7 m box at all; the other curricula just stall at the base. And it does this without a single demonstration, approaching the sample efficiency of reference-based curricula. The ablation is decisive: growing from both ends outperforms either direction alone.
The five papers at a glance
Here's how the five papers split the problem.
| Paper | Slice of the problem | Key idea | Headline result |
|---|---|---|---|
| Planning to Learn | Myopic policy gradients | Horizon loss truncates total error at remaining learning | Beats cross-entropy on ImageNet; gain grows with label noise |
| DepGPO | Credit assignment in terminal agents | Dependency graph from execution traces; credit follows read-write paths | Better task performance and training stability |
| Success conditioning convergence | Missing theory for a common trick | Proves convergence to an optimal policy | O(log(1/ε)) for single-period MDPs, O(1/ε^p) for discounted |
| Phase-structured RL | One policy vs many | Regime-based decomposition by transient vs quasi-stationary duration | Decision rule for when to split |
| BVER | Exploration curricula | Bidirectional start-goal growth, Voronoi-biased | 65% fewer iterations to 95% success; climbs 0.7 m box |
Key numbers
- 65% fewer iterations: what BVER needs to hit 95% success on a 0.4 m box versus the best reference-free curriculum
- 0.7 m: the box height only BVER can climb among reference-free methods
- O(log(1/ε)): iterations to an ε-optimal policy for single-period MDPs under success conditioning
- O(1/ε^p): iterations for discounted MDPs, with p set by problem data
Common pitfalls
If you're going to apply any of this, these are the mistakes I've seen people make.
Use exact policy gradients just because you can compute them. Exactness doesn't buy you patience. If the target is known and training stalls, the myopic gradient is the suspect. Add the horizon-style truncation and watch the same compute budget go further.
Hand out trajectory-level advantage to every command in a terminal session. Unless you trace which commands wrote what the verifier read, irrelevant operations absorb credit and training destabilizes. DepGPO's dependency graph is the fix, and it doesn't require changing the environment, just instrumenting the trace.
Treat success conditioning as a hack with no theory. It has a convergence guarantee now. But check your regime: discounted MDPs cost O(1/ε^p) iterations, single-period problems cost O(log(1/ε)). If you're in the first regime and p is large, budget accordingly.
Force one shared policy across phases because it's simpler. The shared policy can represent anything a multi-policy solution can, in theory. In practice, heterogeneity across phases and thin data per phase break it. Measure the transient duration against the quasi-stationary period before you commit.
Grow your exploration curriculum from one side. Start-state-only and goal-only curricula leave the middle dark. BVER's ablation shows bidirectional growth is strictly better, so if your curriculum always expands from the start state, try growing from the goal too.
One thing to remember
The through-line is simple. All five papers get their wins by changing what the learning signal looks like, not by adding compute to compensate. The horizon loss makes the gradient patient. DepGPO routes credit through actual data dependencies. Success conditioning gets rates instead of vibes. The phase paper matches policy structure to dynamics. BVER shortens the path exploration has to travel. Each is a reminder that the fastest way to make RL learn more efficiently is to look at the signal itself.
The bottom line
Match the learning signal to the structure you can observe. That's the through-line across all five results.
If you're training any policy where the target is known, from classifiers to filtered rollout data, adopt the horizon-loss style truncation, because exact policy gradients optimize only the current update and leave the rest of the training budget on the table.
If you're building agents that operate in terminals or through tool calls, trace command dependencies before redistributing advantages, because trajectory-level credit washes out the steps that actually produced the verifier's answer.
One thing to watch: if you're working on sparse-reward tasks, bidirectional curricula and dependency-aware credit assignment both exploit structure you already have access to, and they'll likely become table stakes in standard RL recipes within six months.