Fooled by Vast Knowledge
Models look like they generalise because the training data is vast enough to hide the difference between interpolation and extrapolation. I think that confusion is leading safety research astray.
Making sure advanced AI systems stay beneficial as they scale. Most of the work sits in two places: what it would take to automate alignment research, and what models know versus what they only appear to know. I'm starting a non-profit research organisation. That work is on the Now page.
Models look like they generalise because the training data is vast enough to hide the difference between interpolation and extrapolation. I think that confusion is leading safety research astray.
For twenty years the mantra was that ideas are worthless and execution is everything. That stops holding when execution is cheap, and taste as curation is not what replaces it.
Guardrails, evals, RL environments, control infrastructure. I work through each one and keep landing in the same place, too far from what happens inside a frontier lab.
Four different things get conflated in the phrase automated alignment research, and the crux is whether applying techniques we already have is a short step from inventing a new paradigm.
A concrete list of things to build now so a coding agent can actually run an interpretability experiment, and why the METR slowdown result does not say what people think.
The ROME edit does not generalise the way you would expect. It runs in one direction only, and cheese and fromage have to be edited separately.
Governance people should be writing the policy document and building the demo for each plausible scenario now, rather than starting the day a minister gives them three months.
A post arrives in your reader when it is finished. No digest, no cadence promise: you get the post, or nothing. Every post is also narrated, and the narration has a feed of its own.
The RSS feed keeps the address and the item GUIDs it has today, so nothing you already follow breaks.