Read Dan Selsam on AI Risk
Dan Selsam’s statement on AI risk is my recommended reading this week, alongside some thoughts on the debate about AI safety funding.
The transition to advanced AI, and how to make it go well.
Dan Selsam’s statement on AI risk is my recommended reading this week, alongside some thoughts on the debate about AI safety funding.
My path from dismissing AI risk to working on AI safety, and the effort to think clearly amid fear, social pressure, and uncertainty.
My PIBBSS Symposium talk from September 2025, with the deck and the full transcript, on why the dual-use worry is overblown.
Models look like they generalise because the training data is vast enough to hide the difference between interpolation and extrapolation. I think that confusion is leading safety research astray.
Guardrails, evals, RL environments, control infrastructure. I work through each one and keep landing in the same place, too far from what happens inside a frontier lab.
Four different things get conflated in the phrase automated alignment research, and the crux is whether applying techniques we already have is a short step from inventing a new paradigm.
A concrete list of things to build now so a coding agent can actually run an interpretability experiment, and why the METR slowdown result does not say what people think.
The ROME edit does not generalise the way you would expect. It runs in one direction only, and cheese and fromage have to be edited separately.
Email when a post is finished. You get the post itself, not a summary, and I send nothing else. Every post is also narrated, and the narration has a feed of its own.
The RSS feed keeps the address and the item GUIDs it has today, so nothing you already follow breaks.