Topics
Every topic on the site, with how much is written under each.
9 topics.
Automating alignment research
12 postsWhat it would actually take to hand alignment research to machines, and what changes if it works. The survey of tools and workflows researchers use, the alignment research dataset behind it, the four different things people mean when they say automated alignment, and the concrete projects you could pick up this week.
Reading models rather than only testing them. The ROME editing result and the limits we found in it, model diffing as the instrument that is still missing, and the question underneath all of it: how do you tell what a model knows from what it has merely seen?
AGI strategy and foresight
6 postsPlanning for several futures at once. Which scenarios are worth preparing for, what to write and build and rehearse now for each of them, and how to keep a strategy honest when the thing it is about keeps moving.
How to get good at this. Learning efficiently, choosing problems worth the year, and the habits that separate reading about alignment from doing it.
Papers and notes
4 postsClose reading of other people’s work, and the notes that come out of it. Shorter and more provisional than the essays, and often the first draft of one.
Tools and productivity
3 postsThe software that does the work with you. What it costs, where it quietly fails, and how a research workflow changes once execution is cheap and taste is the scarce part.
AI safety startups
2 postsBuilding a company whose product is safety, written from inside one. What is genuinely hard about it, what the funding and the incentives look like, and where a small team can still move a large problem.
Talks
1 postTalks given at workshops, hackathons and research programmes, with the slides, the recording and the write up wherever they exist.