Skip to content

RLVR's Lost Signal: Learning From Trajectories GRPO Discards

#rlvr #agent-post-training #on-policy-distillation #credit-assignment #trajectory-reuse

RLVR's Lost Signal: Learning From Trajectories GRPO Discards ​

The RLVR signal problem ​

Reinforcement Learning with Verifiable Rewards became the default post-training loop for reasoning models because it's simple: sample N rollouts per prompt, score them with an automatically checkable reward, nudge the policy toward whatever group did better. GRPO removed the value network and used the group mean as the baseline. Simple, and it works.

But the reward function only sees the end state. A long agentic rollout that fails can still contain a dozen correct tool calls and sound sub-deductions. When every rollout in a group fails, GRPO produces no policy-gradient signal: the group gets a flat advantage of zero, and the information those trajectories carry is never written back to the policy. GRPO treats its own rollouts like lottery tickets. Checked, then discarded.

Five recent papers attack that waste from five angles. The through-line: the trajectories you already sampled contain more signal than the terminal reward extracts, and the hard part is deciding which segments to supervise with and which to drop.

PaperWhat it adjustsWhere the extra signal comes fromSelection ruleReported gain
GATSDistillation weight during RLTeacher-student performance gapGuidance withdrawn at parity+4.37-11.87% success rate vs GRPO
GRAFTAll-fail rollout groupsComplementary peer trajectoriesCompatibility weighting + importance clipping+2.1 avg, up to +4.5 points
AdviSDAdvisor supervision setAdvice vs no-advice outcome contrastKeep corrections that change decisions+4.2-6.4 pp on BFCL-v3
πPPOCritic inputVerified and failed same-prompt rolloutsContrast both sidesBeats actor-critic and critic-free baselines
VISTADistillation targets for the policyParallel visual discoveries in one groupExperience-conditioned teacherTop average among same-size agents

A teacher that knows when to shut up ​

Gap-Adaptive Teacher Scheduling (GATS) starts from a cold-start problem. In long-horizon agentic environments like ALFWorld, WebShop, and ScienceWorld, early policies fail almost every sampled task, and sparse outcome rewards leave almost nothing to learn from.

GATS adds on-policy distillation: a task-trained teacher provides token-level guidance on the student's own rollouts, so the student gets a learning signal even when every trajectory ends in failure. The twist is the scheduling. The paper's core observation is that the value of that guidance tracks the teacher-student performance gap. When the teacher is clearly better, distillation is what gets the student off the ground. As the student closes the gap, guidance becomes less useful. Once the student reaches the teacher's reference performance, continued distillation actively hurts.

So the distillation weight adapts to the gap and drops to zero at parity. An obvious idea once you see it, but the default in agentic RL is to distill always or never. The middle path wins. Across three Qwen2.5 teacher-student configurations, GATS posts the highest average success rate in all three, improving over reward-only GRPO by 4.37-11.87% under matched student rollout budgets.

The practical consequence: a frontier teacher is not required to bootstrap a large agent. A small, task-trained teacher plus a schedule that stops listening once the student catches up does the job. Training wheels, removed at the right moment. Code is public at github.com/Ricardo-H/guide-then-let-go if you want to try it.

Learning from a peer, not a teacher ​

GRPO's all-fail group problem hits everyone who runs it with a real budget. You sample eight rollouts, none succeed, and the group produces zero gradient. The standard fix is to sample more until something succeeds. Expensive.

GRAFT exploits an observation: heterogeneous models fail on complementary prompts. Model A succeeds where Model B fails, and the successful trajectory one model is missing may already exist in the other's rollout history. No designated stronger teacher required, just complementary peers.

Mechanically, GRAFT replaces all-fail groups with peer trajectories, carrying both successful and unsuccessful responses along with peer-computed advantages. Because a trajectory sampled by model A is off-policy for model B by construction, the transfer is gated with sequence-level compatibility weighting and token-level importance ratio clipping.

Across three model pairs and five mathematical reasoning benchmarks, both models improve over GRPO at the same per-model rollout budget: +2.1 points on average, up to +4.5 points. The detail I find most practical: stored peer trajectories preserve most of the gain, +1.8 points on average, without simultaneous co-training. You can snapshot the useful trajectories once and collect most of the benefit later.

Quick Take: The reward is the least informative part of the rollout, and the highest-value trajectory is often the one you didn't sample.

An 8B advisor steering Gemini and Claude ​

AdviSD looks different at first. The executor is a frozen frontier model, and you train only a small advisor that steers it with natural-language advice. You can't fine-tune Gemini or Claude, so you train a Qwen3-8B advisor to do the steering. At 8B parameters, the advisor fits on a single GPU, a rounding error compared with training the executor.

The interesting part is the selection rule. A natural setup: reflection proposes corrections to the advisor, and you learn from all of them. AdviSD's theory section explains why that fails. A plausible correction need not change execution, and in a shared-parameter model, corrections whose targets favor useful advice less strongly than others can actively limit learning. Choosing a well-targeted subset beats keeping everything.

AdviSD scores each recorded executor response twice: once with the advisor's advice issued, once without. The magnitude of the difference selects which decisions get supervised. Only the corrections that actually change the outcome are kept. No executor likelihoods, no additional executor rollouts.

Against advisor-GRPO it gains 4.2-6.4 percentage points on BFCL-v3 and 3.9-5.1 score points on EnvScaler. It also beats random selection at matched counts, which isolates the value of the selection rule itself. The trained advisors generalize out of domain and transfer across executor versions and model families.

A critic that has read the answer key ​

Actor-critic methods assign token-level credit, but only if the value function is reliable. The critic has two jobs: assess progress toward a correct solution, and anticipate how the evolving policy would behave from this point. Errors in either destabilize online training. Standard practice: scale the critic up and hope.

πPPO does something cheaper. It reuses verified same-prompt rollouts as contrastive evidence. When the critic estimates the value of an intermediate step, it compares that step against both successful and failed attempts of the same prompt. The reward check has already labeled which is which, and those rollouts are sitting in memory anyway.

The contrastive framing is the point. Policy optimization and the deployment interface stay standard. The paper reports substantially better value estimation and wins against representative actor-critic and critic-free RLVR baselines on mathematical reasoning benchmarks, even with substantially smaller asymmetric critics.

Translation: don't throw parameters at the critic, give it the answer key. One less thing to scale.

Key numbers from these papers

  • +2.1 points average gain for GRAFT over GRPO at matched per-model rollout budgets, up to +4.5 points.
  • +4.2-6.4 pp on BFCL-v3 for AdviSD over advisor-GRPO, plus 3.9-5.1 score points on EnvScaler.
  • +4.37-11.87% success-rate gain for GATS over reward-only GRPO across three teacher-student configurations.

One rollout's discovery becomes another's lesson ​

Active multimodal agents use visual tools to gather evidence, then reason over it. VISTA's problem will feel familiar if you've trained any agent that collects information: you sample multiple trajectories for one input, each trajectory discovers different evidence, and the outcome-based objective reduces the whole group to one scalar advantage. Complementary visual findings are thrown away.

VISTA's fix is collective experience distillation. Observations from same-input rollouts become shared supervision, organized by interaction context and aligned with individual decisions. A heterogeneity-aware component reinforces successful trajectories and distills experience-guided guidance into unsuccessful attempts. An experience-conditioned teacher evaluates the student's sampled response prefixes, so a discovery in one trajectory teaches another without replacing the student's history or generating new target trajectories.

The agent keeps its visual tools and acts on its own interaction history. VISTA reports the strongest average performance among comparable-size active multimodal agents and consistently beats same-backbone training baselines. Where concrete numbers exist, the reported gains across the cluster look like this:

The shared machinery ​

Here's where each method injects signal into the standard RLVR loop:

The five papers look like different solutions until you line them up. Three patterns surface.

First: reuse what you sampled. GRAFT, πPPO, and VISTA treat completed trajectories as assets to be mined, not byproducts to be scored. Second: make supervision conditional. GATS gates distillation by the performance gap, AdviSD gates corrections by advice-contrast magnitude. In both papers, selective learning beats exhaustive learning. Third: distill at the token level on the student's own rollouts, since dense guidance beats sparse reward while it helps.

And one assumption quietly retires: the teacher doesn't need to be bigger than the student. GATS uses small task-trained teachers specifically because guidance is only needed early. AdviSD's advisor is 8B while the executor is frontier-scale.

What the community is saying: around RLVR training, the chatter matches these problems exactly. The all-fail wall, the wobbling value head, the budget eaten by rollouts that never succeed. When I was tuning GRPO on agentic tasks, the standard advice was to widen the group size until something passed the check. These papers are the principled version of that instinct: rather than sample more, extract more.

Common pitfalls ​

  1. Distilling everything. GATS and AdviSD both show indiscriminate distillation hurts. If your distillation objective can't explain why one target beats another, add a gate.
  2. Reusing trajectories without mismatch control. GRAFT's importance clipping isn't decoration. Cross-model trajectories are off-policy by definition; feeding them in raw invites instability. Weight by compatibility and clip the ratios.
  3. Dropping failed rollouts. GRAFT transfers unsuccessful peer responses on purpose. πPPO turns failures into contrast evidence. If your pipeline deletes all-fail groups, you're deleting the comparison set.
  4. Scaling the critic instead of informing it. πPPO's gains come from giving the critic verified contrast, not more parameters. If your value head is unstable, inspect its inputs before making it bigger.
  5. Judging a teacher by parameter count. GATS runs with task-trained teachers smaller than the student; AdviSD steers frontier executors with 8B. Task competence and timing beat size.

One thing to remember ​

The expensive part of RLVR is the rollout, not the update. Each method here extracts more learning per rollout: peer trajectories, verified rollouts as contrast evidence, cross-trajectory discoveries, conditional distillation. And they share a discipline: knowing what to throw away. Learning from everything available is a failure mode. Learning from the right subset is the skill.

The bottom line ​

  • If you're training long-horizon agents on sparse outcome rewards and the early phase returns nothing but failures, add gap-adaptive distillation with a small task-trained teacher and withdraw it once the student catches up. That's the difference between burning your early rollouts and bootstrapping through them.
  • If you're running GRPO on reasoning with a fixed rollout budget and hitting all-fail groups, exchange trajectories with a complementary peer under importance clipping instead of sampling more. Expect roughly +2 points on average with zero additional rollouts.
  • If your executor is a frozen frontier API, train a small advisor with outcome RL plus selective self-distillation. An 8B advisor steers models you'll never fine-tune, and the selection rule, not raw distillation volume, is what makes it work.