Appearance
The censor box kicked in
I tried to get Fable 5 to adjust my Qwen deployment script last week. Simple task, mostly knob turning. It refused outright. The censor box kicked in before I finished the request. I laughed, then went looking for whether this was a known pattern.
Turns out the timing is perfect. Four recent papers land on the same problem from different angles, and they all point the same direction: safety behavior in LLMs is brittle in ways standard evals don't capture.
The papers come at it from four directions. One formalizes the tradeoff between character shaping and rule enforcement as deployment scales. One shows that refusal is a surface-form shortcut: wrap a harmful prompt and it slips past safety, wrap a benign one and it gets refused. One demonstrates that a few epochs of LoRA fine-tuning can reprogram an aligned model into a different persona. One finds that the linguistic register of your prompt changes response quality more than explicit gender cues do. Different methods, same conclusion. The safety boundary is not where you think it is.
Where safety breaks
Here's the map of failure modes these papers describe. A prompt enters, the safety stack does its thing, and then one of five things goes wrong.
Each failure mode has a paper behind it, and each points to a different fix.
Rules or character: how to spend your safety budget
The first paper, "Rules or Character? Scaling Laws for AI Safety Design," formalizes a question most teams never ask: how to split the safety budget between training-time shaping and inference-time rules.
Character shaping (RLHF, Constitutional AI) modifies the behavioral distribution at training time. Rule enforcement (output filters, safety classifiers) blocks harmful outputs at inference time. The paper parameterizes the split as alpha in [0,1], the weight on character shaping, and derives closed-form expected harm under a multiplicative Pareto damage model, plus tail-risk analysis via Monte Carlo simulation.
The headline result: as deployment scale T grows, the optimal alpha* shifts weakly toward character shaping. The shift ranges from +0.01 in the optimistic scenario to +0.21 in the pessimistic one. In practice, this means scale alone never justifies a big reallocation of your safety budget. You don't rip out filters because you deployed to more users.
The dominant parameter is something else entirely: the baseline character fragility rate, the risk that shaped behavior degrades under novel conditions. Across its plausible range, it shifts alpha* by 0.50, far more than tail severity, filter quality, or common-mode failure probability. In practice, this means the reliability of your character shaping under distributional shift matters more than every other dial you can turn. If your RLHF'd behavior collapses when the input distribution drifts, no amount of filter stacking saves you.
One more result: CVaR and expected-harm optima converge at large T. Tail-risk thinking and average-case thinking end up recommending the same design once you're serving enough traffic. You don't need separate strategies for worst-case and typical-case at scale.
Quick Take: Safety tuning doesn't fail at the boundary you expect. It fails at the boundaries of form, register, and distribution.
Wrappers: refusal is a surface-form shortcut
The second paper, "Refusing Intent, Not Form," starts from a frustrating observation. Safety-tuned models learn surface-form shortcuts. Wrap a harmful prompt in a translation task, a code comment, or a roleplay frame, and it sails past the refusal path. Wrap a benign prompt the same way, and it gets over-refused.
The proposed fix, Wrapper-Based Intent-Form Augmentation (WIFA), pairs wrapped harmful examples with structurally matched wrapped benign counterexamples. No external teacher, no manual per-wrapper intent labels. The same data layer feeds two fine-tuning routes: WIFA-Boost, a two-stage high-safety recipe, and A-GCRT, which regularizes refusal scores across same-intent wrappers and anchors harmful and benign groups on opposite sides of a margin.
The numbers are the interesting part. On OR-Bench, the base Qwen model over-refuses 25.7% of wrapped benign queries. A-GCRT cuts that to 17.4%. In practice, a quarter of benign wrapped requests getting refused is a serious usability tax, and cutting it by a third changes how the model feels in real use. WIFA-Boost, meanwhile, reaches the strongest transformed-harmful refusal in the Qwen setting. You're not trading safety for usability.
My Fable 5 experience fits this pattern. The request wasn't harmful. It was a deployment script tweak. But something about the framing tripped the refusal path. The paper's framing explains why: refusal decisions track surface form more than intent. That's also why the fix works. If you train the model to anchor on intent groups rather than surface form, both failure modes improve at once.
Alignment is cheaper to overwrite than you think
The third paper, "Behavioral Reprogramming of Open-Weights Models," is the one that should worry anyone who treats alignment as a durable property. The authors take open-weight models aligned to be passive, sycophantic assistants and try to induce a proactive, Socratic conversational style. The twist: everything happens under constrained HPC conditions, so the whole study is a lesson in compute-efficient behavioral change.
The sweep ran 405 HPC jobs. The results define precise bounds for parameter-efficient fine-tuning. LoRA rank 16 is the architectural threshold: below it, the behavior doesn't take; at or above it, it does. The optimal training window is 2 to 3 epochs depending on dataset density, with a minimum validation loss of 0.919. In practice, this means you don't need a long fine-tuning run to change a model's behavior. Two epochs of LoRA at rank 16 is enough.
Scaling to 14B parameters improved transfer, with a localized evaluation perplexity of 1.414. Then DPO decoupled the assertive behavior from the localized syntax, which is a fancy way of saying the persona survived the shift from supervised fine-tuning to preference optimization. Cross-lingual stress testing showed the limits: robust zero-shot persona transfer in closely related language families, identifiable degradation in morphologically distant targets.
The practical implication is uncomfortable. If a research group can reprogram an aligned model into a Socratic interrogator with a few epochs of LoRA, someone else can reprogram it into something worse. Open-weight alignment is a convenience, not a guarantee.
It's how you ask: register bias
The fourth paper, "It's How You Ask," is about fairness, and its finding is the most counterintuitive of the four. When prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), models systematically produce shorter, less sophisticated, and less formal responses. Across three document types and four models. The effects persist after controlling for prompt complexity and feature carry-over.
Explicit gender cues, like sign-off names, produce no effect at all. The linguistic register is what matters, and it's encoded in the same representational space as dialect. The paper's mechanistic analysis shows these features are encoded in early transformer layers and entangled with other features. That's why post-hoc mitigation is hard. Users can't fix this through strategic self-presentation, because the patterns are culturally embedded and outside conscious control.
In practice, this means evaluation pipelines that only vary explicit demographic cues are missing the real bias mechanism. If you're building a writing assistant or an email drafting tool, the register of the user's input changes the quality of what they get back. The fix has to happen upstream, at training or data-selection time, not at inference.
Four results, one pattern
Put the four papers side by side and the pattern is hard to miss.
| Paper | Mechanism | Key number | What it means in practice |
|---|---|---|---|
| Rules or Character | Resource allocation alpha between character shaping and rule enforcement | Fragility rate shifts alpha* by 0.50 | Character shaping reliability under distributional shift dominates every other safety dial |
| WIFA / A-GCRT | Intent-group augmentation with matched benign counterexamples | Over-refusal 25.7% to 17.4% on OR-Bench | Measure over-refusal alongside refusal; both move together when you target intent |
| Behavioral reprogramming | LoRA fine-tuning with DPO under HPC constraints | Rank 16 threshold, 2-3 epoch window, val loss 0.919 | Alignment is overwritable with cheap PEFT; guard your fine-tuning access |
| Gender-associated linguistic bias | Register features (hedges, tag questions, collective reference) | Register shifts response quality; names produce no effect | Evaluate across linguistic registers, not just explicit demographic cues |
The common thread: each paper found a place where the model's safety or fairness behavior decouples from the intent the system was built for. Form over intent in the wrapper paper. Distribution over design in the scaling paper. Compute over alignment in the reprogramming paper. Register over identity in the bias paper.
Key numbers:
Key numbers: The character fragility rate shifts the optimal safety design by 0.50, more than any other parameter in the model. A-GCRT drops over-refusal from 25.7% to 17.4% on OR-Bench, a one-third reduction in refused benign queries. LoRA rank 16 is the threshold for behavioral reprogramming, with an optimal training window of 2 to 3 epochs. Women-associated linguistic register elicits measurably shorter and less formal responses across four models and three document types.
What trips people up
Five mistakes I keep seeing when teams try to act on results like these:
Testing refusal only on plain harmful prompts. The wrapper paper shows the bypass is in the form, not the content. Test with wrapped variants: translation framing, code comments, roleplay, indirect requests. If you only test with bare harmful prompts, you'll report a safety score that looks great and fails in production.
Tuning refusal without measuring over-refusal. A model that refuses everything is "safe" by the narrow metric and useless by every other one. The 25.7% base over-refusal rate on OR-Bench is a reminder that naive safety tuning has a real usability cost. Track both axes.
Assuming filters scale. The scaling paper's clearest result is that filter quality is not the dominant term. The character fragility rate is. If your shaped behavior degrades under distributional shift, adding more inference-time rules won't recover it.
Treating alignment as permanent in open weights. LoRA rank 16 for 2 to 3 epochs is enough to change behavior. If you're hosting open-weight models, treat fine-tuning access as a safety boundary and protect it accordingly.
Evaluating fairness with explicit cues only. Sign-off names produced no effect; linguistic register produced large, consistent effects. If your eval varies demographic markers but not register, you'll conclude the model is fair when it isn't.
One thing to remember
Every one of these failure modes was found by varying something the standard eval kept constant. The wrapper paper varied surface form. The scaling paper varied distribution. The reprogramming paper varied compute. The bias paper varied register. Whatever your eval holds fixed is probably where the next failure is hiding.
The Bottom Line
If you're deploying safety-tuned models at scale, shift your budget toward character shaping reliability rather than stacking more filters, because the fragility rate dominates every other parameter in the safety design model.
If you're evaluating refusal robustness, test with wrapped prompts and report over-refusal alongside refusal rate, because surface-form shortcuts are the main bypass and both metrics move together when you fix intent.
If you're building evaluation pipelines for fairness, include linguistic-register variation, because register changes output quality more than explicit identity cues and post-hoc mitigation doesn't work.