Topic
Interpretability and model internals
Reading models rather than only testing them. The ROME editing result and the limits we found in it, model diffing as the instrument that is still missing, and the question underneath all of it: how do you tell what a model knows from what it has merely seen?
7 posts.
Better model diffing is needed
The alignment technique I wish existed. Cheap sensors that show what changed inside a network during training, so training can be steered while it runs rather than audited afterwards.
Using data attribution for AI alignment
In-Run Data Shapley is cheap enough to attribute behaviour during pre-training. What I would do with that: order the curriculum so a model holds human values before situational awareness.
But is it really in Rome? Limitations of the ROME model editing technique
The ROME edit does not generalise the way you would expect. It runs in one direction only, and cheese and fromage have to be edited separately.
Differential Training Process: Delaying capabilities until inner aligned
Keep the model from working out that it is in a training loop until its values are where we want them. Plus the two objections I have not answered.
Is the "Valley of Confused Abstractions" real?
Chris Olah's curve says models get harder to read before they get easier. Neel told me that came out of vision models, so I am posting my confusion.
Notes on Cicero
Cicero never lies. It states the plan it holds and then changes plan, and that is what players experienced as betrayal. Interpretability on the model alone would have missed it.
Detail about factual knowledge in Transformers
The model writes a spread of facts about the subject into the residual stream before it knows what will be asked. An appendix to the ROME post.