About
What I work on: making sense of current AI safety work, the cruxes that decide what labs and funders do next, and the alignment of superintelligent AI.
This website is under reconstruction, so some of the text (like on this page) has been AI-generated from conversations and references.
I’m Jacques Thibodeau. I work on reducing risks from advanced AI, and more broadly on whether the shift to powerful AI goes well for people.
Most of my time now goes to three things. The first is finding the disagreements that decide what labs, funders and governments do next, and working out what evidence would settle them. The second is making sense of current AI safety work, which is tangled enough that people often argue past each other. The third is the alignment of superintelligent AI, including what models actually know rather than what they only appear to know.
The work
The alignment research dataset and paper. In 2022 I collected and cataloged the AI alignment research literature and analyzed it to find the field’s real subfields, rather than the ones people assert exist. We released a paper, “Researching Alignment Research: Unsupervised Analysis”, and an open dataset, and showed that a classifier trained on the corpus finds relevant work nobody had put in it. The project started at AI Safety Camp. Jan Leike, then at OpenAI, called this direction his “favored approach to solving the alignment problem” and mentioned the work in a post at the time.
SERI MATS. In July 2022 I went to Berkeley for two months on the SERI MATS program, to work on aligning language models and on making language models useful for accelerating alignment research.
The ROME result. ROME was one of the most influential model editing papers in prosaic alignment. I tested it and found the edit does not generalize the way people assumed. It is not bidirectional, it mostly edits the token association rather than the concept, and it over or under optimizes depending on the new fact. Read it. Later work connected this finding to the “reversal curse”: Chughtai, Cooney and Nanda (2024) cite my 2022 investigation as an earlier observation, and Kim and colleagues cite it in discussing why model edits fail to update the reverse relationship. Changing a model’s answer to one question does not establish that its underlying knowledge has changed consistently.
Finding the side-effects of model interventions. I co-authored Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models with Quintin Pope, Ajay Hayagreeve Balaji and Xiaoli Fern (2026). We built an automated pipeline that compares models before and after an intervention and tests plain-language hypotheses about how their behaviour changed. Across reasoning distillation, knowledge editing and unlearning, it found both intended changes and unexpected side-effects.
Automating alignment research. This was my main focus for a couple of years. It is now one of the areas I know well rather than the thing I work on. Nearly every AI safety plan includes “automate alignment research”, and almost nobody says which of four different things they mean. I wrote Gaining clarity on automated alignment research to separate them, because the disagreements people think are empirical are usually about which sense they have in mind. Automating AI safety: what we can do today lists concrete projects that would make current coding agents better at running safety experiments. That post came out of my mentorship during the SPAR program. The main work that came out of my PIBBSS fellowship was Automated alignment research needs a better plan, which I presented at the PIBBSS Symposium 2025. That page has the video, the deck and the transcript of the roughly hour-long talk.
Now. I’m starting a non-profit research organisation. It works on the disagreements that actually decide what labs, funders and governments do next, and on the evidence that would settle them or at least sharpen them. The output is reports, research and scenario planning, written for the people making those calls. It’s early, and quiet for now. If you work on this, get in touch on X or LessWrong.
How I think about this
I want a world with less unintentional suffering and no existential catastrophes. I work on AI because I think superintelligent AI is the most consequential thing humans will build, and it can produce either outcome. Above that, I want people to be able to meet their needs and to have room for the kind of experience Scott Barry Kaufman calls transcendence. That is the whole reason the technical work matters to me.
Disclosures
- I received a grant from the Long-Term Future Fund in 2022 to continue the work on accelerating alignment research.
- The main work that came out of my PIBBSS fellowship was Automated alignment research needs a better plan.
- In 2025 and 2026 I worked on an AI safety startup and decided not to keep pursuing it. Why.