Skip to content

Reading and Writing Transformer Weights: SAE Explanations, Task Vectors, and a Compiled Doom

#mechanistic-interpretability #sparse-autoencoders #in-context-learning #task-vectors #interpretability

The explanation bottleneck ​

Sparse autoencoders give us features by the thousands. Each one is a direction in activation space that fires on some recognizable pattern: a concept, a style, a token context. But a direction is just a vector. Figuring out what it means still requires the slow, manual loop of collecting inputs where the feature activates, staring at them, and writing a description. That loop is shallow too. You end up explaining what the feature correlates with, not what it computes.

Three things landed this week that attack that problem from different angles. One fine-tunes a model to read SAE features out as text. One gives you a criterion for when a simple intervention captures everything a set of demonstrations does. And one skips interpretation entirely by compiling a known algorithm into transformer weights, which gives interpretability researchers something they almost never have: ground truth.

SAEVerbalizer: making features talk ​

The SAEVerbalizer paper starts from a simple observation. SAE decoder directions are the vocabulary the feature dictionary uses. If you inject a direction into an LLM's representations and fine-tune the downstream layers, the model can learn to describe what that direction means. No behavioral experiments, no dataset mining. The explanation comes from the direction itself.

The training setup is straightforward. You take SAE feature directions, inject them into the model's activations, and supervise the model to produce natural-language explanations. Once trained, the verbalizer reads new features directly from their decoder directions.

The interesting results are the generalization ones. The verbalizer explains features it never saw during fine-tuning. It transfers across SAE dictionaries trained separately, which suggests decoder directions share a common geometry even when the dictionaries don't. And a lightweight adapter extends it to features extracted from different LLMs entirely.

The intervention experiments are the cleanest evidence that this is real reading, not memorization. Inject two directions and the explanation combines their meanings. Reverse a direction and the meaning shifts accordingly. That's the behavior you'd expect from a system that has learned how direction relates to meaning, not one that has memorized a lookup table.

Key numbers:

  • Unseen features: explanations generalize to features held out from fine-tuning.
  • Cross-dictionary transfer: the verbalizer explains features from separately trained SAE dictionaries.
  • Cross-LLM transfer: a lightweight adapter extends it to features from other models.
  • Composition: injecting two directions yields an explanation that combines their meanings.

Quick Take: SAEVerbalizer turns feature explanation from a manual audit into a learned readout, and the cross-dictionary transfer suggests decoder directions are more portable than the dictionaries that produce them.

Task vectors: when a single shift is enough ​

The second paper asks a related but different question. When you show a model demonstrations, the model changes its internal computation. Sometimes that change can be compressed into a static task vector: one additive shift applied to the activations. Sometimes it can't, and you need query-conditioned transformations or attention routing. When is the cheap option enough?

The paper's Selection-Realization Hypothesis says demonstrations induce a compact family of internal changes, the query selects from that family, and the model's computation constrains how the selected change can be implemented. In plain terms: the demonstrations set up a space of possible interventions, the query picks one, and the architecture decides what that pick looks like internally.

The empirical test is clever. The authors built controlled multimodal tasks where query dependence varies while the underlying task primitives and prompt format stay fixed, then compared correct demonstrations against matched counterfactuals. The finding: a static task vector works when most of the demonstration-induced change is shared across queries. When explicit in-context learning contains query-specific or distributed structure, a local additive shift can't recover it, and you need the more expressive interventions. The paper shows this relationship extends to natural VQA benchmarks and supports cost-aware method selection without test performance.

That's practically useful. If you're building a system that uses in-context learning, you can measure how much of the demonstration-induced change is shared across queries and decide upfront whether a task vector will do. No need to train and evaluate both.

Three ways to relate weights to meaning ​

DimensionSAEVerbalizerTask-vector analysisCompiled Doom transformer
Direction of analysisReads features out as textMeasures how demonstrations compress into interventionsWrites a known algorithm into weights
What you getNatural-language explanations of SAE featuresA criterion for when static task vectors sufficeA checkpoint with known circuits
Training requiredFine-tune downstream layers onceNone, analysis onlyNone, compilation only
Ground truthPost-hoc, needs validationBehavioral, via counterfactualsExact, by construction
Compute costOne fine-tune, then cheap inferenceForward passes only40 minutes per frame on a B200

The table shows how the three approaches fit together. SAEVerbalizer and task-vector analysis are read paths: they extract meaning from weights. The Doom compiler is a write path: it puts meaning into weights. The write path is what makes the read paths testable.

The other direction: compiling Doom into weights ​

The Doom project is the payoff of a compiler that converts computation graphs into transformer weights. The author ported Doom's rendering algorithm into a compatible graph, compiled it, and produced a standard Hugging Face checkpoint. No training anywhere. You load it without trust_remote_code, feed it a prompt representing the scene data, and generate until the model stops. The output is a token stream of pixel-drawing commands: move the cursor, draw a pixel. Apply them mechanically and you get the E1M1 frame.

The numbers are absurd in the best way. One frame is a 3,614-token prompt plus 53,747 generated tokens. Just over 40 minutes on a B200. The original Doom ran at 35 FPS on a 486. This runs at 35 FPD, frames per day, on hardware that costs more than the 486 did new.

The host program that loads the checkpoint, generates the render, and parses the output is 43 lines of Python. The computation graph definition is much longer, but that's what gets compiled into the weights. The model is Doom's algorithm, expressed in parameters. No approximation, no distillation, no training.

Why compiled transformers are the ground-truth testbed ​

This matters for interpretability beyond the spectacle. When you compile a known algorithm into weights, you know exactly what every circuit does. That's ground truth, something SAE research almost never has. SAE explanations are post-hoc inferences from behavior. Compiled weights are verified by construction.

The Doom checkpoint is a stress test for interpretability tools. Run an SAE on it and check whether the features you find correspond to the actual algorithm components: the ray casting, the texture mapping, the wall distance calculations. If your tools can't recover known structure in a model where the structure is guaranteed to exist, what does that say about their outputs on models where it isn't?

There's a deeper point here. The interpretability field has spent years building read tools for weights we don't understand. The compiler work gives us write tools for weights we do understand. The combination is the closest thing to a controlled experiment the field has.

What the community is saying ​

The Reddit thread on the Doom transformer had the usual mix of awe and pointed questions. When I loaded the checkpoint, the first thing that struck me was the host code. 43 lines. The whole pipeline, from prompt to rendered frame, fits on one screen. That's the part that makes the project feel less like a stunt and more like a tool.

The debate in the comments was about what this actually proves. Some argued the model is just a very elaborate interpreter, and the real work happened in the compiler. That's fair, but it misses the point. The compiler is the contribution. If you can compile arbitrary computation graphs into transformer weights, you can build interpretability testbeds on demand, with whatever structure you want to study.

The practical complaints were predictable. 40 minutes per frame means you're not iterating quickly. The 21B parameter checkpoint doesn't run on consumer hardware. And the generated token stream is verbose: 53,747 tokens to draw one frame that the original engine rasterized in milliseconds. Nobody is shipping this as a renderer. It's a proof that the write path exists.

Common pitfalls ​

The biggest mistake I see in SAE work is explaining features from external behavior alone. You collect the activating examples, find the common thread, and call it an explanation. That describes correlation, not computation. SAEVerbalizer's premise is that the decoder direction itself carries the meaning, and its transfer results back that up. If your explanation procedure never touches the direction, you're explaining the dataset, not the feature.

A second trap is assuming a static task vector is always sufficient. The Selection-Realization results are explicit: a task vector works only when the demonstration-induced change is shared across queries. If your task has query-specific structure, the additive shift silently fails. The paper's contribution is that you can measure this upfront instead of discovering it after deployment.

A third one: treating SAE features as independent atoms. The intervention experiments show directions compose and reverse. Inject two and the explanation merges. Flip one and the meaning flips. Features are relational, and explanations that treat them in isolation will miss how they interact.

A fourth, specific to the compiled transformer line of work: don't assume compiled checkpoints are cheap to run. The Doom model is 21B parameters and takes 40 minutes per frame on a B200. Plan your experiments around that constraint or pick a smaller compiled model. The compiler approach scales down, but you have to actually do the scaling.

One thing to remember ​

The read path and the write path are converging. SAEVerbalizer reads features out as text, task-vector theory tells you when a simple intervention captures a behavior, and the Doom compiler writes known algorithms into weights. Each one alone is a useful tool. Together they give interpretability something it has been missing: a way to check its own answers.

The bottom line ​

If you're building SAE-based interpretability pipelines, adopt the verbalization approach from SAEVerbalizer, because replacing manual feature audits with a learned readout that transfers across dictionaries removes the main bottleneck in scaling explanations.

If you're choosing between task vectors and more expressive interventions for in-context learning, measure how much of the demonstration-induced change is shared across queries first, because that single number tells you whether the cheap option works.

One thing to watch: compiled transformers are going to become the standard validation ground truth for interpretability tools within the next year, because they're the only setting where you can check explanations against circuits you know are there.