Topic
Papers and notes
Close reading of other people’s work, and the notes that come out of it. Shorter and more provisional than the essays, and often the first draft of one.
4 posts.
The importance of Entropy
A short physics note. A constant stream of low-entropy energy from the Sun favours structures that dissipate it, which may make life less an accident than a consequence.
AI Insights #1: How Misalignment Could Lead to Takeover & Necessary Safety Properties
A round-up of three: Christiano on takeover thresholds, Shane Legg's necessary properties for any AGI safety plan, and why the danger sits in the agents built on top of models.
Notes on Cicero
Cicero never lies. It states the plan it holds and then changes plan, and that is what players experienced as betrayal. Interpretability on the model alone would have missed it.
A descriptive, not prescriptive, overview of current AI Alignment Research
We catalogued the alignment literature and let the analysis say which research directions actually exist. The AI Safety Camp project behind the arXiv paper and the alignment research dataset.