Topic

Interpretability and model internals

Reading models rather than only testing them. The ROME editing result and the limits we found in it, model diffing as the instrument that is still missing, and the question underneath all of it: how do you tell what a model knows from what it has merely seen?

7 posts.

All topics or everything, by year