Appearance
The Control Problem in Diffusion Generation
The generation race cooled off
The latest round of diffusion papers reads like an admission that the base generation race is settled. Raw text-to-image and text-to-video quality, the thing everyone fought over a year ago, is good enough that the bragging moved elsewhere. The interesting work now happens in the control loop around the generator: how you condition it, how fast it runs, whether it obeys physics, and how you un-teach it things.
Nine arXiv papers and two Hugging Face releases landed in the same window, and they orbit a single question: now that we have the engine, how do we steer it?
Here is the cluster at a glance.
| Method | Steers | Mechanism | Headline result |
|---|---|---|---|
| NEPA | the conditioning signal itself | predicts next embeddings at every denoising step | FID 1.32 on ImageNet 256x256 at a third of REPA's training compute |
| DMAD | sampling speed | distribution matching recast as classification on shared-backbone discriminators | FID 1.04 one-step; four-step SDXL FID 14.47 beats the teachers it was compared against |
| 4Director | camera and object motion | canonical mesh per object, rigid transform per frame, depth-video conditioning | view-consistent control across viewpoint changes |
| HiPhy | physics compliance | RL with local and global physical rewards | largest gains on concurrent multi-principle scenes |
| TVRL | RL credit assignment | token-level credit maps from VLM gradient magnitudes | +3.60 VBench-2.0 Overall over base |
| CEASE | continual concept erasure | two subspace constraints on a closed-form solver | consistent erase-preserve tradeoff across sequential requests |
| RASteer | erasure specificity | retain-orthogonal steering with overlap-adaptive calibration | matches or beats activation steering and weight editing |
Better conditions beat bigger models
In a diffusion transformer, the condition is embedded once and reused at every denoising step. A class token or text prompt gets projected to an embedding, and the model leans on the same signal whether it is predicting pure noise or fine detail. The NEPA paper (Next-Embedding Predictive Autoregression) asks whether the condition can be predicted instead of fixed. The clean image follows the noisy image in generation, so the clean image's embeddings are the next embeddings after the condition and the noisy image. NEPA trains a Transformer to predict them all at once with Multi-Embedding Prediction, and a DiT generator conditions on those predictions, recomputed at every denoising step. The conditioning signal adapts to the current noisy state.
The result, combined with REPA, is NEPA-DiT-XL at FID 1.32 on class-conditional ImageNet 256x256, using about a third of REPA's training compute. FID 1.32 on this benchmark is below the range where artifacts dominate, but this is really a training-efficiency result. A lab with a fixed GPU budget gets a better generator, or spends the saved hours elsewhere. The trade-off is structural: NEPA adds a second network to every sampling step, so you pay extra inference compute to save training compute. The paper studies that exchange directly, which is the honest way to report it.
Distillation without the second model
Distillation is how diffusion becomes practical, and it carries a tax. Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, which means keeping an auxiliary diffusion model fitted to the student's evolving distribution. That eats memory and compute. DMAD recasts distribution matching as classification and learns the log-density ratios directly. Two discriminator heads on a shared backbone separate real data and teacher samples from the student's, and linear losses on their logits train the student. The authors prove that at the discriminator optimum these losses recover the distribution-matching gradient behind DMD, through the standard identity linking discriminator logits to log-density ratios. The proof is what makes this a shortcut rather than a heuristic.
The numbers hold up across scales. DMAD reaches FID 1.04 with one-step generation on ImageNet-64x64, which means a generator that samples essentially instantly. On text-to-image, four-step SDXL hits FID 14.47 on COCO-10K, better than the multi-step teachers it was compared against. On video, a four-step student of Wan2.1-T2V-14B scores 85.15 total on VBench, also the best among the compared few-step methods and teachers. 14B parameters means the base model needs an A100-class card, and four steps is what turns it into an interactive tool. On MiniMax-H3-33B, the four-step student wins overall human preference over DMD2 79.1% of the time and over rCM 84.6% of the time, ties excluded. The normal "distillation costs quality" caveat does not apply in these comparisons.
A gap-based reweighting scheme ties the two heads together: teacher supervision is adapted across noise levels from the real-data head's empirical logit gap between real and teacher samples. That is the part that keeps the student from chasing a static target.
Quick Take The base generator stopped being the bottleneck. Every paper in this cluster is about steering: adaptive conditioning, one-step distillation, rigid geometry, token-level rewards, and erasure that doesn't bleed.
Grab the geometry
Existing object control in video models is coarse. Image-plane cues are ambiguous in depth and rotation, and 3D tracks or blobs lack complete geometry, so consistency breaks when the viewpoint changes. 4Director conditions a video world model on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh, then moved by one prescribed rigid transformation per frame. That gives an intuitive 3D control interface and stops unobserved geometry from being regenerated independently in every frame, which is where consistency usually dies. The controlled scene is rendered as a depth video, and a Motion Adapter turns that geometric scaffold into the final video while synthesizing view-consistent appearance, illumination, and non-rigid dynamics.
For training, the authors built RealCOD-Rigid, 20,774 clips annotated with rigid 3D scenes by an automatic pipeline. That is small by video-model standards, so the annotation pipeline matters as much as the architecture; it is the part that scales. They also introduce Identity-Gated IoU, which jointly scores adherence to the prescribed motion and preservation of object identity, two objectives that usually fight each other. For production use, the key claim is the interface: you drag a mesh instead of writing a prompt that hopes the model interprets "pan left" correctly.
Physics by reward, credit at the token level
Video models fail physics most visibly when several principles have to hold at once. "A balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently in the same shot, and existing methods mostly handle one principle per video. HiPhy is a reinforcement learning framework with a dual-level objective: locally it enforces the temporal dynamics of individual physical principles, and globally it keeps the physical and semantic coherence of the scene. The authors built a 50K-prompt dataset covering co-occurring physical events and a benchmark, MultiPhyBench, to measure them. HiPhy's largest gains come exactly on the multi-principle scenes where competing methods degrade most sharply. If a method survives balloon-plus-steam, the single-principle cases are easy.
TVRL attacks a different localization problem. In video RL, one scalar reward for a whole video cannot say which tokens are wrong, so the optimizer perturbs already-correct regions while under-correcting the bad ones. TVRL derives token-level credit from the reward model itself: the answer likelihood of a frozen vision-language model feeds the video-level reward, and the magnitudes of its video-input gradients reveal which generated tokens drive that score. TVRL instantiates this in GRPO, averaging question rewards into a group-relative advantage and using detached token-credit maps to reweight dense denoising log-probabilities inside the clipped policy ratio. On VBench-2.0 it scores 57.69 Overall, 3.60 points over the base model, and it beats matched GRPO baselines across three SDE samplers and four reward models by margins from 1.33 to 3.15 points. The spread across reward models matters: token-level credit is not tied to one scoring function.
The shared move in both papers: reward the part that matters, locally, instead of hoping a global signal drags the sample in the right direction.
Erasure you can ship without retraining
Continual concept erasure is the compliance problem nobody wants to have. Erasure requests arrive over time, and if consecutive edits are not constrained, residual perturbations outside the retain set interact and accumulate. Unrelated generations degrade, and previously erased targets can collapse back toward noise. CEASE is training-free and imposes two subspace constraints on a closed-form solver: it adds the token representation of the shared replacement to the solver's invariance matrix, and when interference is detected, it projects the current update onto the orthogonal complement of dominant output directions from cumulative past updates. A closed-form decomposition attributes the accumulated interference to repeated activation of the shared replacement and overlap between successive update directions, so each constraint targets a measured source. Across continual erasure of celebrities, styles, and instances, CEASE keeps the most consistent erase-preserve trade-off, where existing methods either degrade general generation or fail to erase.
RASteer handles the single-shot version of the same problem. An erasure direction built from the target concept also contains shared components that retained concepts rely on, because target and retained concepts overlap in representation space. RASteer builds a retain subspace from the concepts to preserve, removes components aligned with it from the erasure direction via Retain-Orthogonal Steering, then uses Overlap-Adaptive Calibration to decide, per layer and per denoising step, how much of each shared component to strip. Removing all shared components weakens erasure, so the calibration is the part that balances erasure against preservation. On unsafe-content, instance, and style erasure across multiple backbones, RASteer matches or beats the activation-steering and weight-editing baselines.
If you host an image API, this is the difference between removing one style and quietly degrading every prompt that shares visual features with it.
Key numbers NEPA-DiT-XL hits FID 1.32 on ImageNet 256x256 at about a third of REPA's training compute. DMAD reaches FID 1.04 with one-step generation on ImageNet-64x64, four-step SDXL at 14.47 on COCO-10K, and VBench 85.15 with four-step Wan2.1-T2V-14B. TVRL gains 3.60 points on VBench-2.0 Overall over its base model, and the training-free planner routes 300+ agents around 100+ obstacles in under 6 seconds on a GPU.
Composition as a staged workflow
Positive and negative space is one of those composition principles humans internalize and models ignore. Generating it means coordinating two semantic concepts that share a boundary, and single-pass prompting does not get you there. FaV-A treats it as a staged agent workflow: generate a base object, analyze its shape and spatial structure to identify candidate negative-space semantics, then produce compositional instructions for the final generation stage. Staged workflow beats clever prompting here. FaV-A is a pipeline with a vision-language model in the analysis role, and ablations show it produces more visually coherent and semantically aligned positive-negative space than direct zero-shot MLLM baselines. Nothing about the architecture is exotic, which is the point: some controls are workflows, not weights.
Motion planning with zero demonstrations
The same denoising machinery shows up in robotics. Diffusion planners cast trajectory generation as iterative denoising, which handles multi-modal trajectory distributions and whole-trajectory refinement, but they train on large collections of feasible trajectories. That makes them map-specific and useless when high-quality demonstrations do not exist.
The new work replaces learned global trajectory scores with analytical local scores from obstacle, smoothness, velocity, and inter-agent feasibility terms. The structural observation is that a trajectory's score can be reconstructed from local interactions between neighboring waypoints and nearby constraints. That yields a decomposed denoising procedure which keeps the optimization structure of classical trajectory methods and inherits the iterative refinement of diffusion. It generates feasible paths for 300+ agents in environments with 100+ obstacles in under 6 seconds on a GPU, beating strong learned and optimization baselines. Under 6 seconds for 300 agents is the difference between mid-mission replanning and precomputed plans, and there is no training set to collect.
What the community is saying
The gap between the arXiv stack and the Hugging Face front page is instructive. Viggle's Qwen-Image-2.1-viggle-turbo is the distillation promise in product form: fewer steps on a strong base model, positioned for character animation. I ran it because the base sampling loop is slow for iterating on poses, and the turbo variant cut that noticeably. Character consistency still degrades when I push large rotations, which is exactly the failure 4Director's rigid geometry is designed to fix.
The other artifact sits at the opposite end of the control spectrum. An all-in-one "uncensored" LoRA space for Qwen-Image-2.1 (arudradey/qwen-image-2.1-uncensored-aio-loras) has 95 likes, runs on zero persistent GPUs, and I found it does what the name says. The failure mode that matters is that the bundle degrades general prompt following, which is the same bleed-over problem CEASE and RASteer subtract with subspace constraints. Same week: researchers publish closed-form math to remove copyrighted styles, and hobbyists publish aggregate adapter weights to keep everything in. Both are steering a frozen model toward a target, just with opposite signs.
Common pitfalls
Five mistakes show up again and again when people build with this stack.
Don't distill with an auxiliary score model when you are VRAM-bound. Keeping a diffusion model fitted to the student's evolving distribution roughly doubles the memory footprint of DMD-style training. DMAD's shared-backbone discriminators recover the same distribution-matching gradient without that second model.
Don't erase concepts by subtracting the target embedding at inference. Target and retained concepts share representation space, so plain subtraction suppresses unrelated generations. RASteer's retain-subspace projection exists because this fails in practice, and CEASE's invariance constraints exist because it fails repeatedly.
Don't reward a whole video with one scalar in RL. The optimizer cannot localize errors, so it perturbs already-correct tokens and under-corrects the bad ones. Use reward-model gradient magnitudes for token-level credit like TVRL, or budget for wasted training runs.
Don't assume a clever prompt handles multi-principle physics. Models trained on single-principle events degrade most on scenes that combine buoyancy and fluid dynamics. If you need concurrent physical laws, you need a physics reward, not prompt engineering.
Don't mistake depth maps for 3D control. A depth video steers the camera but not object identity across viewpoint changes. If your use case needs identity-consistent motion, you need an explicit geometry representation like 4Director's canonical meshes.
One thing to remember: the value of a diffusion model now comes from what you can do to it after training, not what it does out of the box. Adaptive conditioning, few-step distillation, token-localized rewards, and erasure that does not bleed are the four skills that separate a model you can build on from a demo you show once.
The Bottom Line
If you are training a class- or text-conditioned DiT and GPU hours are the constraint, adopt NEPA-style embedding prediction as your conditioning path. You pay a second network at sampling time, and you get roughly three times the model quality per training compute.
If you are deploying image or video generation at scale, plan for distilled few-step students as the default artifact. DMAD's four-step SDXL at FID 14.47 and four-step Wan2.