Topic
Automating alignment research
What it would actually take to hand alignment research to machines, and what changes if it works. The survey of tools and workflows researchers use, the alignment research dataset behind it, the four different things people mean when they say automated alignment, and the concrete projects you could pick up this week.
This tag is now called Automating alignment research. The old address still works and still lists the same writing.
12 posts.
Fooled by Vast Knowledge
Models look like they generalise because the training data is vast enough to hide the difference between interpolation and extrapolation. I think that confusion is leading safety research astray.
Gaining clarity on Automated Alignment Research
Four different things get conflated in the phrase automated alignment research, and the crux is whether applying techniques we already have is a short step from inventing a new paradigm.
Automating AI Safety: What we can do today
A concrete list of things to build now so a coding agent can actually run an interpretability experiment, and why the METR slowdown result does not say what people think.
AI Alignment Project Ideas
Seven project write-ups for ARENA and LASR participants, from improving MAIA to reward misspecification and the cases where a model gets the right answer from bad reasoning.
Accelerating AI Alignment Research (Talk)
The keynote I gave to open the research augmentation hackathon I co-organised with Apart Research. Slides and project suggestions are linked underneath.
Using data attribution for AI alignment
In-Run Data Shapley is cheap enough to attribute behaviour during pre-training. What I would do with that: order the curriculum so a model holds human values before situational awareness.
My current research and request for collaborators
The bio I wrote for EAG Bay Area 2024: what I was working on, what I wanted to argue about, and the three kinds of collaborator I was looking for.
An incomplete list of projects I'd like to work on in 2023
Two sentences pointing at a LessWrong shortform where I listed the projects I wanted to work on in 2023.
(Linkpost) Results for a survey of tool use and workflows in alignment research
A pointer to the results of the March 2022 survey on tool use in alignment research. The write-up itself is on LessWrong.
A descriptive, not prescriptive, overview of current AI Alignment Research
We catalogued the alignment literature and let the analysis say which research directions actually exist. The AI Safety Camp project behind the arXiv paper and the alignment research dataset.
A survey of tool use and workflows in alignment research
Before building anything we asked alignment researchers how they work and where a language model tool would earn its place. The survey is closed, the dual-use note still stands.
Interesting Applications of GPT-3: Elicit
My first post on this site, walking through Ought's Elicit in 2021. The subquestion decomposition is the earliest form of the idea I have been working on ever since.