Skip to content

What Four New Papers Reveal About Model Internals

#interpretability #mechanistic-interpretability #sparse-autoencoders #vision-transformers #multimodal-llms

Interpretability has an evidence problem ​

Interpretability research has a credibility problem. Papers end with heatmaps and an assertion that some neuron "represents" dogs, or honesty, or the concept of a left turn. The claim is rarely falsifiable, and the field has learned to live with that.

Four papers on arXiv this month are trying something harder. They build instruments that test representational claims under controlled conditions, and the results are less flattering than the usual narrative. A sparse autoencoder analysis of DINOv2 shows that one class of tokens does the semantic work while another class mostly encodes background texture. A study of math LLMs finds they translate between subfield vocabularies without preserving the hypotheses implicit in the source. A verification-style paper proves that no finite test can certify skill removal, no matter how exhaustive. And an MLLM study shows that vision-language fusion follows different internal pathways depending on which architecture family you picked.

Different layers of the stack, same move: define a measurement, run a control, then say what the model actually does rather than what it looks like it does.

The four papers at a glance ​

PaperTargetInstrumentHeadline result
Decoding register and outlier tokens (2610.03698)DINOv2 token activationsSparse autoencoders, UMAP clustering, CLIP cross-checks, causal ablationsRegister features carry semantics; outlier features carry texture. Disrupting them drops representation cosine similarity 48.17% vs 0.31%
Objects Without Morphisms (2610.03551)Math LLMs, 7 models across 4 familiesBlind-queue coding of truth/content/scope, planted-positive controlTranslations widen or narrow quantification asymmetrically by direction; implicit hypotheses rarely stated
Certified Mechanistic Edits (2610.03502)ReLU nets up to a softmax + LayerNorm transformerSound bound propagation over continuous embedding regionsRemoval and preservation become provable; about 9x the input dimension exact solvers reach. Finite tests provably can't certify removal
Architecture-Dependent Fusion Pathways (2610.03289)Concatenation vs native multimodal LLMsAlignment decoupling, attention routing/entropy, intrinsic dimensionality, causal interventionsText-first vision-later pathway vs early visual-textual co-adaptation

The papers don't just add results. They add methods you can reuse on other models, which is rarer than it should be.

Register tokens carry semantics; outlier tokens carry texture ​

Self-supervised ViTs like DINOv2 develop a well-known pathology: background regions produce patch tokens with abnormally high norms, and those outlier tokens clutter every downstream analysis. The register design was introduced to absorb this behavior, but nobody had cleanly established what either token type is actually doing.

This paper trains sparse autoencoders on register-token and outlier-token activations, then runs the features through an automated interpretability pipeline with UMAP clustering and CLIP-space cross-checks. The division of labor is stark. Register-token features associate with high-level semantic concepts. Outlier-token features associate with lower-level structural patterns, backgrounds, and textures.

The causal ablation is where the paper earns its keep. Disrupting the top-activating register-derived features produces a 48.17% drop in representation cosine similarity. Disrupting the outlier-derived features produces a 0.31% drop. That's roughly two orders of magnitude, in the same model, on the same inputs. The practical read: when an SAE feature lights up on a background texture, that is not the model recognizing a concept. It is the model's token accounting system misfiring.

When I've trained SAEs on ViT activations myself, the first thing that jumps out is how many latents seem to encode image statistics rather than objects. The background regions are exactly where this happens, which is why the register design exists. This paper gives you a way to measure the split instead of feeling it.

Quick Take: these four papers share one discipline: they test representational claims with interventions and controls, and the claims mostly fail in instructive ways.

Math LLMs translate objects, not hypotheses ​

The second paper asks a question that benchmark scores can't answer: when an LLM for mathematics translates a statement between the dialects of neighboring subfields, does it preserve what the statement actually asserts?

The setup is subtle. Fidelity in translation hinges on the level of generality at which content is asserted. The source leaves that level implicit in its vocabulary, so a faithful translation must recover it from the relationship between the theories. The authors built an instrument that codes truth, content, and scope in separate blind queues, with a judge-free measure of whether a rewrite states the hypothesis implicit in its source. A planted-positive control establishes sensitivity.

Across seven models from four families, the direction of translation predicts the error. Translating toward the general framing widens the domain of quantification in 60.6% of rewrites and narrows it in none. Translating toward the concrete framing narrows it in 28.3% and widens it in a negligible 0.3%. The hypothesis that would prevent the widening is stated in only 21.6% of model rewrites, versus 4.2% for human-written statements. Instructing the model to state every hypothesis it requires raises that rate but does not improve its sensitivity to direction.

The title is a category theory joke with a sharp edge. A category without morphisms is just a set of objects. The claim is that these models have acquired an object-level correspondence between subfield vocabularies without the structure-preserving maps that would carry hypotheses from one theory to the other. The model knows that "group" in one dialect maps to "group" in another, but not that the translation must preserve the generality at which a claim is made.

Two details keep this from being a sizing problem. Capability doesn't govern the asymmetry: every model tested shows it, and the most capable widens least. And the result replicates on the pre-registered held-out half of the benchmark, and on statements written by mathematicians.

The bigger implication lands on benchmarks. Expert-level competition math performance is largely a search effect: sample candidates in quantity, retain the ones an external criterion accepts. That improves the outcome that survives while leaving what the model represents untouched. A strong pass@k tells you almost nothing about what the model knows.

Certified edits: when removal is a guarantee ​

Mechanistic edits, ablations, weight edits, activation steering, are the standard tools for unlearning a harmful capability while preserving useful ones. Current practice validates them by testing. The problem: a test can never cover an entire continuous region of inputs.

This paper certifies the behavioral effect of an edit instead. The guarantee is that disabling a circuit removes one skill and preserves another for every input in a region, a feature non-interference property in the information-flow-security sense. They demonstrate certified edits from toy ReLU networks up to a standard softmax + LayerNorm transformer, proving removal and preservation over continuous embedding-space regions.

The scaling trick is sound bound propagation. Instead of checking inputs one by one, you propagate bounds through the network that cover whole regions. That reaches roughly 9x the input-perturbation dimension an exact solver can handle. For a small transformer, that's the difference between certifying a handful of input dimensions and certifying a region wide enough to be meaningful.

Key numbers:

  • 48.17% vs 0.31%: representation cosine similarity drop when register-derived vs outlier-derived features are disrupted
  • 60.6% vs 0%: rewrites that widen vs narrow the domain of quantification when translating toward general framing
  • 28.3% vs 0.3%: rewrites that narrow vs widen toward concrete framing
  • 21.6% vs 4.2%: rewrites vs human statements that state the implicit hypothesis
  • 9x: input-perturbation dimension certified by sound bound propagation vs an exact solver

Then comes the result that should change how you talk about unlearning. The authors prove that no finite deterministic black-box test can certify removal. They exhibit an edit that passes exhaustive testing yet provably fails on a survivor pocket that can be made arbitrarily small.

I've hit this on much smaller scales: a circuit looks dead across a large eval suite, then you generate one adversarial input and the behavior reappears. The survivor pocket result is that experience, formalized. "Passed our evals" and "removed" are not the same sentence.

The honest caveats matter. Guarantees hold on small, standard-architecture networks, and any removal claim presupposes that the target skill admits a decidable specification. Real-world harms may not have that property. The paper is a proof of concept for a standard, not a tool you can point at a production model today. But it sets the right bar.

Two fusion pathways in multimodal models ​

The fourth paper asks where vision and text actually meet inside multimodal LLMs. The answer depends on which of the two architectural paradigms you're using.

Concatenation architectures bolt a vision encoder onto a text model and feed visual tokens in as part of the sequence. Native multimodal architectures train the whole stack together from the start. The authors run three connected analyses: alignment decoupling to identify which modality changes, attention routing and entropy to characterize how cross-modal information is distributed, and intrinsic dimensionality to examine how fusion reshapes feature spaces. Causal interventions validate the resulting interpretation.

The two pathways are distinct. Concatenation models follow a text-first, vision-later pathway: the text side carries the process, and visual information gets incorporated later in the network. Native models exhibit earlier visual-textual co-adaptation, with feature-space reorganization happening sooner and more pervasively. A supplementary visual CKA analysis probes the Platonic Representation Hypothesis against these fusion findings.

DimensionConcatenation modelsNative multimodal models
Fusion onsetText-first, vision-laterEarly visual-textual co-adaptation
Feature-space reorganizationLater in the networkEarlier
Practical signalVision interventions matter most in later layersCross-modal intervention points exist early on

If you're debugging or editing an MLLM, architecture family tells you where to look. Layer indices that work for steering one family will miss the other. This paper gives you a diagnostic sequence, alignment decoupling, attention analysis, intrinsic dimensionality, to find the fusion points yourself instead of copying numbers from a blog post.

Common pitfalls ​

Reading causal ablation drops as standalone importance scores. A 48% cosine drop means something only against a baseline. Feature knockouts vary by layer, token norm, and how many features you kill. Run a same-count control ablation on random features and report the gap, not the raw number.

Treating a passing eval suite as proof of removal. The survivor pocket result is not a corner case, it's a theorem: an edit can pass exhaustive testing and still fail on an arbitrarily small region. Report coverage, report the region you tested, and stop saying "removed" when you mean "passed our tests."

Confusing benchmark scores with knowledge. Competition math results are mostly a search effect: sample a lot of candidates, keep the ones an external checker accepts. The representation underneath never changes. If you're evaluating a math LLM, treat search-augmented benchmark scores as search scores and design a separate probe for what the model actually represents.

Training SAEs on raw ViT activations without handling high-norm outliers. DINOv2's background regions produce tokens whose norm dominates everything else, and your SAE latents end up encoding norm rather than concepts. Normalize the activations or use the register tokens, otherwise your interpretability tool is measuring token accounting, not semantics.

Assuming fusion happens at the same layers across MLLM families. Concatenation and native multimodal models fuse vision and text at different depths and with different dynamics. Find the fusion onset per model before you pick intervention layers.

What this adds up to ​

Each paper is small in scale. One backbone analyzed. Small networks certified. One translation benchmark. Specific model families. The shared contribution is methodological: a way to convert an interpretive claim into a measurement that can fail.

That matters because the failure modes are consistent. The math LLMs show that strong benchmark performance coexists with deep representational gaps. The certified edit paper shows that empirical validation has a hard ceiling. The ViT analysis shows that not all internal activity is created equal, and systems that don't account for that will misread their own models.

One thing to remember: representational claims become trustworthy when they survive an intervention and a control. "Looks like" doesn't clear that bar. These four papers are the field starting to insist.

The Bottom Line ​

If you're evaluating a math-focused LLM, treat pass@k with external verification as a search score, not a knowledge score. The gap between search performance and representation is the most reproducible finding across model families in this batch.

If you're making a removal or unlearning claim, either run a certified method on a small, decidable circuit or use honest language about coverage. A test suite is evidence, not a guarantee, and the survivor pocket theorem means the difference is not pedantic.

If you're steering or patching an MLLM, identify the architecture family before picking layers. Fusion onset differs between concatenation and native models, so intervention points that work on one will miss on the other. Expect architecture-aware diagnostics to become the default within two release cycles.