Epistemic Status: I'm trying to keep up a pace of a post per week on average as I've found it a good habit to get more into writing. Inspired by this post by Eukaryote I've tried to do the thing where I create something that is the easier to digest version of the way I think about AI Safety problems in terms of an intuition pump. I'll also add my usual disclaimer that claude was part of the writing process of this including the intial generation of tikzpictures.
The AI Society Lens
Philosophy has given us several powerful tools for escaping the limits of individual perspective. Kant's categorical imperative asks "what if everyone did this?"—transforming local decisions into universal policies to reveal hidden contradictions. Rawls' veil of ignorance asks "would you design society this way if you didn't know your position in it?"—forcing impartiality by stripping away self-interest. Virtue ethics asks "what kind of person does this action make me?"—shifting focus from isolated acts to the character they cultivate.
These are what Daniel Dennett would have called intuition pumps—thought experiments that teleport you to a vantage point where consequences become visible that were hidden from your original position.
I think there’s a lens that is quite interesting for AI Safety.
The AI Society Lens: Take a property of a single AI system and ask "what would a society run by such agents look like?"
The first figure illustrates this move. On the left, a single agent with some pr...
Make this part of your paper trail
Save this paper to a shelf, write a review, and keep your own notes.