Archive
Everything published here, newest first.
2026
6 postsFooled by Vast Knowledge
Models look like they generalise because the training data is vast enough to hide the difference between interpolation and extrapolation. I think that confusion is leading safety research astray.
When Execution Gets Cheap, Does Taste Become the Moat?
For twenty years the mantra was that ideas are worthless and execution is everything. That stops holding when execution is cheap, and taste as curation is not what replaces it.
Hard Truths About Where AI Is Headed
Four LinkedIn posts collected in one place: coding jobs going away, planning as the scarce skill, rogue models arriving, and what the Pentagon story says about red lines.
Better model diffing is needed
The alignment technique I wish existed. Cheap sensors that show what changed inside a network during training, so training can be steered while it runs rather than audited afterwards.
Difficulties in Building an AI Safety Startup
Guardrails, evals, RL environments, control infrastructure. I work through each one and keep landing in the same place, too far from what happens inside a frontier lab.
Gaining clarity on Automated Alignment Research
Four different things get conflated in the phrase automated alignment research, and the crux is whether applying techniques we already have is a short step from inventing a new paradigm.
2025
1 postAutomating AI Safety: What we can do today
A concrete list of things to build now so a coding agent can actually run an interpretability experiment, and why the METR slowdown result does not say what people think.
2024
8 postsAI Alignment Project Ideas
Seven project write-ups for ARENA and LASR participants, from improving MAIA to reward misspecification and the cases where a model gets the right answer from bad reasoning.
How much I'm paying for AI productivity software (and the future of AI use)
What I was paying for AI tools in September 2024, roughly nine hundred dollars a month, and the harder question underneath it: why it is difficult to spend meaningfully more.
The importance of Entropy
A short physics note. A constant stream of low-entropy energy from the Sun favours structures that dissipate it, which may make life less an accident than a consequence.
Accelerating AI Alignment Research (Talk)
The keynote I gave to open the research augmentation hackathon I co-organised with Apart Research. Slides and project suggestions are linked underneath.
Using data attribution for AI alignment
In-Run Data Shapley is cheap enough to attribute behaviour during pre-training. What I would do with that: order the curriculum so a model holds human values before situational awareness.
Quantum Computing, Photonics, and Energy Bottlenecks for AGI
My answer is that quantum probably does not matter and photonics might. The energy numbers, why optical transistors are the wrong thing to judge photonics by, and OpenAI's hire.
AI Insights #1: How Misalignment Could Lead to Takeover & Necessary Safety Properties
A round-up of three: Christiano on takeover thresholds, Shane Legg's necessary properties for any AGI safety plan, and why the danger sits in the agents built on top of models.
My current research and request for collaborators
The bio I wrote for EAG Bay Area 2024: what I was working on, what I wanted to argue about, and the three kinds of collaborator I was looking for.
2022
17 postsBut is it really in Rome? Limitations of the ROME model editing technique
The ROME edit does not generalise the way you would expect. It runs in one direction only, and cheese and fromage have to be edited separately.
An incomplete list of projects I'd like to work on in 2023
Two sentences pointing at a LessWrong shortform where I listed the projects I wanted to work on in 2023.
(Linkpost) Results for a survey of tool use and workflows in alignment research
A pointer to the results of the March 2022 survey on tool use in alignment research. The write-up itself is on LessWrong.
How learning efficiently applies to alignment research
Studying only counts here if it shortens the path to insights that solve alignment. Wentworth's point about newcomers losing years to low-value sub-problems is the thing to avoid.
Differential Training Process: Delaying capabilities until inner aligned
Keep the model from working out that it is in a training loop until its values are where we want them. Plus the two objections I have not answered.
Near-Term AI capabilities probably bring low-hanging fruits for global poverty/health
A call to charity entrepreneurs in December 2022. Models of that generation were about to make some interventions viable, and the non-profit sector had not caught up.
Foresight for AGI Safety Strategy
Governance people should be writing the policy document and building the demo for each plausible scenario now, rather than starting the day a minister gives them three months.
Is the "Valley of Confused Abstractions" real?
Chris Olah's curve says models get harder to read before they get easier. Neel told me that came out of vision models, so I am posting my confusion.
Notes on Cicero
Cicero never lies. It states the plan it holds and then changes plan, and that is what players experienced as betrayal. Interpretability on the model alone would have missed it.
Detail about factual knowledge in Transformers
The model writes a spread of facts about the subject into the residual stream before it knows what will be asked. An appendix to the ROME post.
Current Thoughts on my Learning System
I coasted on raw ability and it stopped being enough. Learning is a set of skills you have to practise, and this is my public commitment to build the system.
What does "Effective" in EA mean to you?
Effective, to me, always meant doing the most good inside the share of your life you have decided to give it. A movement too demanding of everyone collapses.
Helping organizations survive disasters (and potentially avoid them altogether)
My 2019 primer on strategic foresight, reshared for AGI forecasters. Low-probability high-impact developments get ignored because they are unlikely, and Bostrom's urn is the reason to look anyway.
AI Alignment YouTube Playlists
Two YouTube playlists of alignment talks, split by whether you have to watch the slides, so that one of them works while you walk.
I'll be in Berkeley for SERI MATS for the next 2 months
A short note from July 2022 saying I was heading to Berkeley for SERI MATS, and that I would like to meet people while I was there.
A descriptive, not prescriptive, overview of current AI Alignment Research
We catalogued the alignment literature and let the analysis say which research directions actually exist. The AI Safety Camp project behind the arXiv paper and the alignment research dataset.
A survey of tool use and workflows in alignment research
Before building anything we asked alignment researchers how they work and where a language model tool would earn its place. The survey is closed, the dual-use note still stands.
2021
1 postInteresting Applications of GPT-3: Elicit
My first post on this site, walking through Ought's Elicit in 2021. The subquestion decomposition is the earliest form of the idea I have been working on ever since.