Appearance
The bottleneck moved past the weights
For two years, making an agent smarter meant one thing: a bigger model underneath. The prompt, the tools, the loop around the model were plumbing. Nobody trained the plumbing.
That's over. The improvement surface has moved to the loop itself. This fall, six groups hit the same conclusion from different directions: the structure around the model, what the harness-learning paper calls "the executable program that organizes model calls, tool use, and information flow," is now the thing these systems optimize. Microsoft's SkillOpt trains a skill document like a neural network, with epochs, learning rates, and validation gates. A companion line of work trains a proposer model to rewrite a solver's harness from execution feedback. UMM-Reflection runs reinforcement learning over complete render-critique-revise trajectories inside one model. ARISE evolves its own rubrics as the agent's weaknesses shift. AI Night-Scientist teaches models when to stop being predictable. Ouroboros packages the same instinct as a spec-first workflow for coding agents.
Sometimes that means updating weights with RL over the whole loop. Sometimes the weights stay frozen and only the harness changes. Either way, the unit of optimization is no longer a single rollout. It's the loop.
The loop that keeps showing up
A language-model agent is two things: a model and a harness. The harness decides when to call the model, what context it gets, what tools it can reach, and how its outputs are checked. Prompts, skill documents, rubrics, reflection instructions, tool orchestration: all of it.
Every system in this cluster runs the same skeleton. Draw it once, because the papers use different words for the same parts.
SkillOpt gates every edit on a held-out validation score and accepts only strict improvements. ARISE re-derives its rubrics from rollout evidence instead of keeping them fixed. Ouroboros runs a three-stage evaluation gate before a new spec is accepted. UMM-Reflection replaces the gate with a group-relative advantage computed across sibling trajectories. The shapes differ. The skeleton doesn't.
What differs is what counts as the harness, and who does the editing. SkillOpt edits a text skill document using a separate optimizer model. Harness learning edits executable programs with a trained proposer. ARISE edits rubrics and skill activations. UMM-Reflection edits reflection tokens and the image generation jointly, inside one model. Ouroboros edits the seed spec and workflow of a coding task. Same loop, six different edit targets.
SkillOpt: a skill document as trainable state
Microsoft's SkillOpt is the most explicit about the analogy. Train agent skills like you train neural networks, with epochs, mini-batches, learning rates, and validation gates, but without touching model weights. The repo's complaint about current practice is sharp: agent skills are hand-crafted, generated one-shot by a strong LLM, or evolved through loosely controlled self-revision. None of those behaves like a deep-learning optimizer. None reliably improves over its starting point under feedback.
SkillOpt treats the skill document as the trainable state of a frozen agent. A separate optimizer model turns scored rollouts into bounded add, delete, or replace edits on a single skill document. In the default path, an edit is accepted only when it strictly improves a held-out validation score. A textual learning-rate budget, a rejected-edit buffer, and an epoch-wise slow/meta update keep training stable.
The deployed artifact is a compact best_skill.md, typically 300 to 2,000 tokens. That's small enough to drop straight into a system prompt or a skills folder, and it runs against the unchanged target model with zero extra inference calls. You pay for training once; the skill costs nothing at runtime.
Results: across six benchmarks, seven target models, and three execution harnesses, SkillOpt is best or tied-best on all 52 evaluated (model, benchmark, harness) cells. On GPT-5.5 it lifts average no-skill accuracy by 23.5 points in direct chat, 24.8 points inside the Codex agentic loop, and 19.1 points inside Claude Code. A double-digit average lift is the difference between an agent you demo and an agent you ship. Optimized skills transfer across model scales, between Codex and Claude Code harnesses, and to nearby benchmarks without further optimization.
Harness learning: a proposer that rewrites the playbook
Where SkillOpt edits text, harness learning edits programs. The starting point is the same: a language-model agent is jointly defined by its model and its harness, and different tasks call for different ways of organizing model calls, tool use, and information flow. Since the harness needs to adapt, the paper formulates the problem as meta-learning over executable programs. Harness revisions play the role that weight updates play in gradient-based adaptation.
The training setup is clean. A proposer model is trained with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, with no parameter-space update at all. Adaptation happens entirely in the program around the model.
Two findings stand out. Policies trained on individual revisions keep improving harnesses over multiple rounds, which is the property you actually want in a continuing agent. And the ability to adapt at test time transfers to unseen tasks, so the proposer learns a generalizable improvement strategy rather than a patch for specific benchmarks. Training on revision sequences helps in some settings and not others, so multi-round credit still needs work.
Quick Take: the shared diagnosis across this cluster is that agent performance is now limited less by the model and more by the executable structure around it, and that structure responds to disciplined optimization just as weights do.
UMM-Reflection: credit across a render-fix-render loop
UMM-Reflection is the technically hardest case in this cluster. Unified multimodal models can look at images and render them, so in principle a model can repair its own generations: diagnose what an image gets wrong, revise it, render the result, observe, and diagnose again. The obstacle is that whether a revision helped is only known after the revised image is rendered. Reflection text and image generation have to be learned jointly, over the whole loop.
SFT on reflection trajectories gives a cold start but doesn't find the high-success repair paths. Naive RL that optimizes only the renderer, or only one head, leaves most of the gain untapped. The fix is group-relative advantage over sibling trajectories: siblings share one initial image, so the relative advantage compares reflection strategies directly, while a trajectory-level advantage updates both the reflection tokens and the flow-based revisions in one pass. That avoids the combinatorial blow-up of per-round credit assignment. No verifier is needed at inference.
The numbers back the design. UMM-Reflection improves GenEval by 12.05 points over SFT, the jump from a model that can describe a fix to one that can render it. The gains transfer to benchmarks never seen in training: +10.97 on WISE, +3.48 on OneIG-Bench, and +4.63 on T2I-CompBench++.
The held-out transfer is the result that matters. A reflection skill that generalizes to entirely new benchmarks is a skill; one that only works on training distributions is a lookup table.
ARISE and AI Night-Scientist: the target moves too
Two papers attack a subtler failure: the evaluation criteria themselves go stale. ARISE starts from a sharp observation. As a long-horizon agent improves, previously observed weaknesses recede while new limitations emerge, so a static rubric and static training priorities become misaligned with what the agent still needs to learn. Even when capability gaps are identified, rollouts from the current policy keep reproducing the same failures instead of exploring better alternatives.
ARISE's answer is adaptive rubric-skill co-evolution. Rubrics evolve to reward partial behavioral progress. Paired skills are refined and selectively activated to guide exploration toward unresolved weaknesses. Capability-based adaptive sampling prioritizes tasks that target behaviors still needing work. On SkillsBench and Terminal-Bench, both long-horizon agent benchmarks, the framework improves task performance and training efficiency at the same time.
AI Night-Scientist is the same instinct pointed at scientific ideation, where LLMs have a low-entropy bias that produces homogeneous, predictable proposals. The system models creativity along three axes grounded in cognitive science: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea). GRPO training exposes the model to varying degrees and forms of creativity across those axes.
The results: 27.8% more diverse research directions, 14.9% more contribution types, predicted citation impact up by 32 percentage points, and originality up 66.2 points. A 27.8% wider spread of directions means the model proposes hypotheses that differ in scope, not reworded variants of one idea. The negative result is just as important: none of this comes from raising decoding temperature. Sampling noise without semantic guidance produces incoherent proposals, not creative ones. What works is RL training that teaches the model which kind of creativity to apply, and when.
Key numbers
SkillOpt lifts GPT-5.5's average accuracy by 23.5 points in direct chat and 24.8 points inside the Codex loop, from a single 300 to 2,000 token skill file. UMM-Reflection adds 12.05 points on GenEval over the SFT cold start and transfers to WISE at +10.97, a benchmark it never saw in training. AI Night-Scientist widens the spread of proposed research directions by 27.8% and lifts predicted citation impact by up to 32 points.
Ouroboros: the same loop, no training loop
The most complete engineering expression of this idea trains nothing at all. Ouroboros is an "Agent OS" for coding workflows: a local-first runtime that turns non-deterministic agent work into a replayable, observable, policy-bound execution contract. The workflow is interview, seed, execute, evaluate, evolve. A Socratic interview forces the operator to expose hidden assumptions before any code is written. Those answers crystallize into an immutable seed spec with acceptance criteria, an ontology, and an ambiguity gate. Execution runs through a Double Diamond decomposition. Evaluation runs a three-stage gate: mechanical checks, then semantic evaluation, then multi-model consensus. The evaluation output feeds the next iteration. The grading command and the expected result never appear in the spec the agent receives, which closes the standard reward-hacking hole where an agent optimizes the test instead of the task.
The GitHub threads around SkillOpt and Ouroboros tell the same story. Running a vague "build me a task CLI" through the interview surfaced assumptions I didn't know I was making; the repo's own log shows 12 hidden assumptions exposed and ambiguity scored down to 0.19 before any code was written. The "looks good" manual QA problem is the other recurring pain point. My team kept hitting review where the agent had built the wrong thing confidently, because a clarity failure looks identical to a capability failure in the execution logs. The three-stage gate exists to catch exactly that. One controversy here has nothing to do with the code: "ouroboros" is also a memecoin ticker, and the maintainers had to publish a disclaimer disavowing any token. That's open source in 2026.
How the approaches compare
| System | Trainable state | Editor | At deployment | Headline result |
|---|---|---|---|---|
| SkillOpt | Skill document, a 300 to 2,000 token markdown file | External optimizer model trained with RL | Zero extra model calls | Best or tied-best on all 52 evaluated cells; +23.5 points direct chat |
| Harness learning | Solver's executable harness | Proposer model trained with RL | No parameter updates | Revision quality improves; adaptation transfers to unseen tasks |
| UMM-Reflection | Reflection tokens and flow-based generation, jointly | Same model, group-relative RL | No verifier needed | +12.05 GenEval; +10.97 on held-out WISE |
| ARISE | Rubrics, skill activations, sampling priorities | Co-evolution from rollout evidence | Training-time only | Better task performance and training efficiency |
| AI Night-Scientist | Creativity policy along action, process, outcome | GRPO on the three axes | Same inference cost | +27.8% direction diversity; +32 points predicted citations |
| Ouroboros | Seed spec and workflow per task | Evolutionary loop over evaluations | Zero; workflow only | Ambiguity gated at 0.19, three-stage verified builds |
The split that matters: UMM-Reflection and AI Night-Scientist update weights with RL over the whole loop, while SkillOpt, harness learning, and Ouroboros keep the weights frozen and treat the harness as the trainable state. ARISE sits between, training both the policy and the rubrics that define progress. Both sides converge on the same skeleton, and the skeleton is the story. The validation gate shows up in nearly every variant, which makes it the load-bearing piece: the mechanism that keeps self-improvement from drifting into self-deception.
Common pitfalls
Don't optimize a skill against a frozen evaluator forever. ARISE exists because static rubrics go stale: the failure modes your loop fixed stop mattering, and new ones appear. If the evaluation mix never changes, your loop converges on a score instead of on competence.
Don't accept edits that look good on training rollouts. SkillOpt's accept-only-on-strict-validation-improvement rule is what keeps the loop stable. Remove the gate and the skill drifts until the agent gets worse in ways the training rollouts can't see. The rejected-edit buffer isn't bookkeeping. It's a memory of what didn't work.
Don't treat the skill artifact as a hand-written prompt. The failure mode SkillOpt names is one-shot generation by a strong LLM and loosely controlled self-revision. If your "training" is a human rewriting a prompt every few days, you're doing prompt engineering with extra steps, and you won't get the steady, verifiable improvement a real optimizer would.
Don't train one role of a self-correcting loop and freeze the other. UMM-Reflection's negative result is blunt: naive RL that optimizes only the renderer, or only the reflection head, leaves most of the gain untapped. In a look-revise-render loop, credit has to flow across the round to both roles.
Don't reach for higher temperature when you need diverse ideas. Night Science shows sampling noise without semantic guidance produces incoherent proposals, not creative ones. If you want structural departure, specify which creative axis you're rewarding.
One thing to remember
The trained state of an agent is shrinking. A 2,000 token skill file. A revised harness. An evolved rubric. A locked seed spec. In every system here, the deployable artifact is a text structure that orchestrates a model, and the loop that produces it looks the same: score, edit, gate, deploy. The next round of agent improvement will look less like model releases and more like disciplined loops over harnesses.
The bottom line
If you're building agents that run the same task family over and over, adopt a skill-training loop like SkillOpt. The 300 to 2,000 token artifact gives you double-digit point gains that transfer across models and harnesses, and it costs zero inference time at deployment.
If you're building self-correcting generative systems, image repair being the clearest case, use trajectory-level RL like UMM-Reflection. Splitting credit between reflection text and generation leaves most of the gain on the table, and an external verifier at inference is a cost you don't need.
If you can't run training loops at all, start with a spec-first workflow like Ouroboros. Most agent failures are input-clarity failures, and a Socratic interview plus a three-stage evaluation gate will beat prompt tweaks on a frozen model. One thing to watch: expect these loops to move from research repos into agent frameworks within two quarters. The validation gate, not the model call, is becoming the standard interface for agent improvement.