Skip to content

The Reward Bottleneck: Seven New RL Papers, One Shared Problem

#reinforcement-learning #reward-design #policy-optimization #intrinsic-motivation #exploration #fully-homomorphic-encryption

The reward bottleneck ​

Every paper in this batch is wrestling with the same bottleneck: the objective. FERPO reworks the policy improvement step so it stops differentiating the critic with respect to actions. F-CIP derives intrinsic rewards from system dynamics alone, no hand-picked information variables required. InterEvolve turns rewards into programs that an LLM evolves at test time. ATPO makes reward shaping controllable for video safety moderation. HAO keeps value learning stable when it has to run under fully homomorphic encryption. And a theory paper shows when exploration rewards quietly send agents somewhere useless. Only Faynt, the Super Smash Bros. Melee agent, is mostly a scale story, and even it leans on curricula and distillation rather than reward design.

The RL loop hasn't changed. What's shifting is where researchers intervene in it.

Four of the five dotted connections sit on the reward and value-learning stage. That's not a coincidence. Reward design is the last great hand-crafted artifact in RL, and these papers treat it as the thing to fix.

PaperProblem it attacksCore mechanismHeadline result
FERPOUnreliable policy updates from critic derivativesForward-KL target fitting with self-normalized importance samplingCompetitive results, faster actor updates than REPPO
F-CIPHand-selected information variables in intrinsic rewardsControllable Information Production defined by dynamics aloneUnsupervised discovery of balancing and running gaits
Exploration theoryIntrinsic rewards that don't buy informationCounterfactual information criterionCount-based and similar objectives shown Pareto-suboptimal
InterEvolveNew tasks with a frozen controllerLLM-evolved reward programs plus numerical tuningReleases competence human rewards leave untapped
ATPOBinary framing and static objectives in video safetyAdaptive Tversky RewardJaccard Index 40.66 to 75.44 on SafeWatch-Bench-Real
HAOValue divergence under FHE polynomial approximationZero-mean centered TD targets0% boundary breaches, baseline hit 83.8%
FayntScaling competitive Melee play10M/75M Transformers, curricula, distillation98.4% win rate vs specialists, 5.2 ms per decision

FERPO: stop differentiating the critic ​

Most continuous-control policies improve by following the gradient of a learned critic with respect to actions. You sample an action, ask the critic how good it is, and nudge the policy toward higher-scoring actions. DDPG-style methods have run on this for a decade. It rests on an assumption that quietly fails in practice: a critic that predicts returns accurately also has accurate derivatives. Those are different properties. A value function can be right at every point you evaluate and still have a misleading local slope, and policy updates inherit that error.

FERPO removes the differentiation step. Instead of pushing actions along the critic's gradient, it derives an optimal target action distribution from a policy-improvement objective regularized by entropy and KL divergence, then fits the actor to that target with a forward-KL objective, estimated using self-normalized importance sampling with actions from the rollout policy. The KL regularization keeps the target close to the rollout policy, which keeps the importance weights from exploding.

The forward-KL choice is the interesting part. Reverse-KL, the usual KL objective in policy optimization, is mode-seeking: it collapses onto a single high-value region and ignores the rest. Forward-KL encourages coverage of multiple high-value modes, and that mode coverage acts as exploration.

Experiments on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains, and the actor updates run faster than REPPO's. None of this lands like a win-rate statistic. It's a reliability fix, and the field needs more of those. If you've ever watched a policy unravel after a critic update that looked fine on paper, you already know the failure mode.

The intrinsic reward puzzle ​

Two papers in this batch attack intrinsic rewards from opposite ends.

F-CIP starts from an uncomfortable observation: every intrinsic motivation objective the field has produced asks you to pick information variables first. Which dimensions of the observation carry signal? Which features should the agent track? That choice reintroduces the domain expertise that unsupervised RL was supposed to eliminate. Forward CIP is a formulation of Controllable Information Production defined by the system's dynamics alone. No variables to select. The paper proves the objective is RL-compatible, and agents trained with it discover primitive behaviors like balancing and controllability maintenance without any task reward. Add a trivial forward-velocity reward and you get hopping and running gaits that normally take hand-built reward shaping to produce.

Then there's the other paper, which is a bucket of cold water. It proposes that exploration be judged by counterfactual information (how well a policy's histories can substitute for experience under alternative policies). Under that formal criterion, count-based, prediction-error, empowerment, and information-gain objectives can all have maximizing policies that are Pareto-suboptimal at acquiring information. The authors build a single, simple environment where every one of those objective families fails. The practical reading is blunt: maximizing an intrinsic reward is not the same as becoming informed.

It's a failure mode I recognize from my own training runs. Curiosity bonuses that keep climbing while the agent circles through high-novelty states and never touches the transition that actually matters. The paper turns that intuition into a formal criterion and shows the failure is structural, not a tuning problem. F-CIP's contribution is showing that at least one intrinsic objective escapes the hand-selection trap.

Quick Take: the strongest results in this batch come from changing what gets optimized, not from adding compute, and the pattern shows up from exploration to policy improvement to video safety.

InterEvolve: rewards you evolve at test time ​

InterEvolve attacks a slice of the reward problem that gets little attention: the task arrives after the controller is frozen. You have a broad humanoid loco-manipulation controller, trained once, and a new contact-rich, multi-stage task shows up. Retraining is off the table. The paper's insight is that the controller already holds most of the competence the new task needs. What's missing is an interface expressive enough to specify contact-rich sequences, and measurable enough that execution feedback can guide planning.

The interface has two parts. An object-aware forward-backward foundation model turns a new reward about the body or objects into behavior at test time, using object residuals on a frozen body prior. And tasks are specified as reward programs: staged rewards with completion conditions and tunable constants. At test time, an LLM agent revises the program structure in context, drawing on execution feedback and a library of verified programs, while a numerical optimizer tunes the constants. Every candidate gets verified across parallel simulation scenarios, so the loop is try, evaluate, revise, repeat. The program improves from the controller's own attempts, and what it learns sticks around in the skill library.

The results are an indictment of hand-designed rewards. Human-authored rewards left much of the forward-backward model's loco-manipulation competence untapped. The evolved programs released it, sometimes through strategies a human wouldn't have written. The system produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run on a physical Unitree G1 from egocentric onboard perception. Frozen controller, evolved reward, real robot. That's the loop closing.

ATPO: a precision-recall dial for video safety ​

Video safety detection keeps getting framed as binary: is this harmful or not? Real moderation is multi-label. One video can tick several unsafe categories at once, and a pipeline that has to act per category needs more than a single harmful/not-harmful knob. It also needs to choose its operating point per category. One category might tolerate false positives, another absolutely can't.

ATPO is a reinforcement learning framework that makes that choice part of the training objective. The Adaptive Tversky Reward adjusts false-positive and false-negative penalties during training, so the precision-recall trade-off becomes steerable. You train once, then move the operating point to match the moderation policy.

The numbers justify the machinery. Jaccard Index on SafeWatch-Bench-Real jumps from 40.66 to 75.44. For multi-label detection, that's close to doubling the number of correct label sets, and it's the difference between a tool moderators can trust and one they can't. The same mechanism steers the precision-recall operating point reliably, which matters for deployment: heterogeneous moderation pipelines can tune the same model per category instead of maintaining separate classifiers.

HAO: when your critic lives under encryption ​

When I first saw "reinforcement learning under fully homomorphic encryption," I assumed it was a niche systems problem with a boring answer. It's niche, but the failure mode is anything but boring, and it's a warning about treating approximation as an afterthought. In FHE, every nonlinear operation has to be replaced with a polynomial approximation. ReLU becomes a polynomial. Sigmoid becomes a polynomial. That substitution is tolerable in a feedforward network where errors stay local. It is not tolerable in RL, because TD learning is recursive: the approximated value function feeds its own next estimate. The approximation error doesn't stay bounded, it compounds through the recursion. The paper calls this Bellman drift, and it's a specific, preventable disease.

Standard stabilization doesn't cure it. L2 weight decay plus gradient clipping, the usual medicine, breached the safe polynomial approximation domain on 3 of 5 seeds. The unstabilized baseline breached it in 83.8% of episodes. HAO's fix is structural: it adapts the zero-mean centering projection from advantage-based value estimation to TD targets. This linear projection annihilates the uniform state-value baseline, which is what drives the drift, while preserving per-state action rankings. It costs zero additional nonlinear multiplicative depth and avoids expensive ciphertext bootstrapping. Across all random seeds, HAO agents had 0% boundary breaches, and in tabular domains they improved optimal policy accuracy by 18 percentage points. It also stays stable when DP-SGD-style Gaussian noise lands on the clipped gradients.

The lesson reaches beyond FHE: in a recursive learning loop, approximation error is a stability problem wearing an accuracy costume.

Key numbers from this batch:

  • 98.4%: Faynt's same-character win rate against prior Melee bots, 240 of 244 games
  • 40.66 to 75.44: Jaccard Index without and with ATPO on SafeWatch-Bench-Real
  • 83.8%: share of episodes where the unstabilized FHE baseline breached the safe approximation domain
  • 5.2 ms: Faynt's per-decision inference time on an NVIDIA T4, about a third of a 60 fps frame
  • 18 points: HAO's gain in tabular optimal-policy accuracy

Faynt: scale beats imitation loss ​

Faynt is the outlier in this batch. It's a scale and systems paper, and its results are unusually clean. Two Transformer policies, 10M and 75M parameters, each control all 26 Super Smash Bros. Melee characters from a single checkpoint. The 10M model wins 240 of 244 same-character games, 98.4%, against fourteen prior specialist and multi-character releases, with a winning record against every one of them. One caveat the authors state plainly: those opponents carry a 21- or 24-frame action delay and Faynt uses no added delay, and the effect of that difference hasn't been isolated. Against a privately supplied zero-delay Slippi-AI model, the 10M still wins all 68 games across two conditioning settings.

The supervised 10M wins 69.7% of games on the initial benchmark, compared with 45.4% for the pretrained 75M, despite having higher overall held-out controller-prediction loss. If you picked checkpoints by imitation loss, you'd pick the worse model. The weighted validation loss the authors use for selection does agree with win-rate ordering, but only because it was engineered to. Offline proxy metrics in competitive RL can contradict the thing you actually care about.

ModelParametersInference per decision (T4)Win rate, initial benchmark
Supervised 10M10M5.2 ms69.7%
Pretrained 75M75M8.7 ms45.4%

Training itself is a pipeline to copy: pretraining on roughly 840,000 human replays, then post-training with rank- and outcome-based curricula, distillation from the 75M down to 10M, and RL restricted to Fox mirror matches. The inference numbers matter too. 5.2 ms per decision on a T4 fits inside a single 60 fps frame with room to spare, which is what makes real-time play on commodity hardware plausible. Weights, benchmark suites, and a tournament platform are all open-sourced.

Common pitfalls ​

Some failure modes these papers make obvious.

Don't assume a critic with low value error is safe to differentiate. Value accuracy and action-derivative accuracy are different properties, and FERPO's premise is that the first doesn't buy you the second. If your policy updates are erratic, try fitting to a value-derived target distribution before you blame the learning rate.

Don't validate intrinsic rewards by watching the bonus climb. Under the counterfactual information criterion, count-based, prediction-error, empowerment, and information-gain objectives can all be Pareto-suboptimal at acquiring informative experience. A curiosity signal that keeps rising while the agent learns nothing is a symptom, not a success.

Don't ship FHE-based RL with regularization as your only defense. L2 weight decay and gradient clipping breached the safe approximation domain on 3 of 5 seeds in the HAO evaluation, and the unstabilized baseline breached in 83.8% of episodes. Bellman drift is structural, so it needs a structural fix like zero-mean centered TD targets.

Don't reduce video safety detection to binary classification. The multi-label framing is what makes a controllable precision-recall trade-off possible, and ATPO's jump from 40.66 to 75.44 Jaccard shows how much signal binary framing throws away.

Don't select checkpoints for competitive agents by imitation loss. Faynt's supervised 10M beat the pretrained 75M on win rate with worse held-out prediction loss. Evaluate on the game outcome.

Sources ​

One thing to remember: the strongest results in this batch come from rethinking what gets optimized. FERPO changes the policy update. F-CIP and InterEvolve change where rewards come from. ATPO changes what the reward penalizes. HAO changes what the TD target contains. And the one scale-heavy entry, Faynt, still shows a smaller model beating a larger one once the training objective is fixed.

The bottom line ​

If you're planning RL under fully homomorphic encryption, budget real time for approximation stability, because Bellman drift will breach your safety bounds in 83.8% of episodes without a structural fix like HAO's zero-mean centered TD targets.

If you have a frozen controller and new tasks keep arriving, adopt test-time reward program evolution instead of retraining, because InterEvolve shows human-designed rewards leave most of the controller's competence untapped.

If you're picking checkpoints for a competitive agent, optimize on game outcomes, not imitation loss, because Faynt's 10M model beat its 75M counterpart at 69.7% versus 45.4% despite worse held-out prediction loss.