How I Came to Take AI Risk Seriously
My path from dismissing AI risk to working on AI safety, and the effort to think clearly amid fear, social pressure, and uncertainty.
The transition to advanced AI, and how to make it go well.
My path from dismissing AI risk to working on AI safety, and the effort to think clearly amid fear, social pressure, and uncertainty.
My PIBBSS Symposium talk from September 2025, with the deck and the full transcript, on why the dual-use worry is overblown.
Models look like they generalise because the training data is vast enough to hide the difference between interpolation and extrapolation. I think that confusion is leading safety research astray.
Guardrails, evals, RL environments, control infrastructure. I work through each one and keep landing in the same place, too far from what happens inside a frontier lab.
Four different things get conflated in the phrase automated alignment research, and the crux is whether applying techniques we already have is a short step from inventing a new paradigm.
A concrete list of things to build now so a coding agent can actually run an interpretability experiment, and why the METR slowdown result does not say what people think.
The ROME edit does not generalise the way you would expect. It runs in one direction only, and cheese and fromage have to be edited separately.
Governance people should be writing the policy document and building the demo for each plausible scenario now, rather than starting the day a minister gives them three months.
Email when a post is finished. You get the post itself, not a summary, and I send nothing else. Every post is also narrated, and the narration has a feed of its own.
The RSS feed keeps the address and the item GUIDs it has today, so nothing you already follow breaks.