Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Immanuel Kant and the Decision Theory App Store, published by Daniel Kokotajlo on July 10, 2022 on LessWrong.
[Epistemic status: About as silly as it sounds.]
Prepare to be astounded by this rationalist reconstruction of Kant, drawn out of an unbelievably tiny parcel of Kant literature!
Kant argues that all rational agents will:
“Act only according to that maxim whereby you can at the same time will that it should become a universal law.” (421)
“Act in such a way that you treat humanity, whether in your own person or in the person of another, always at the same time as an end and never simply as a means.” (429)
Kant clarifies that treating someone as an end means striving to further their ends, i.e. goals/values. (430)
Kant clarifies that strictly speaking it’s not just humans that should be treated this way, but all rational beings. He specifically says that this does not extend to non-rational beings. (428)
“Act in accordance with the maxims of a member legislating universal laws for a merely possible kingdom of ends.” (439)
Not only are all of these claims allegedly derivable from the concept of instrumental rationality, they are supposedly equivalent!
Bold claims, lol. What is he smoking?
Well, listen up.
Taboo “morality.” We are interested in functions that map [epistemic state, preferences, set of available actions] to [action].
Suppose there is an "optimal" function. Call this "instrumental rationality," a.k.a. “Systematized Winning.”
Kant asks: Obviously what the optimal function tells you to do depends heavily on your goals and credences; the best way to systematically win depends on what the victory conditions are. Is there anything interesting we can say about what the optimal function recommends that isn’t like this? Any non-trivial things that it tells everyone to do regardless of what their goals are?
Kant answers: Yes! Consider the twin Prisoner's Dilemma--a version of the PD in which it is common knowledge that both players implement the same algorithm and thus will make the same choice. Suppose (for contradiction) that the optimal function defects. We can now construct a new function, Optimal+, that seems superior to the optimal function:
IF in twin PD against someone who you know runs Optimal+: Cooperate
ELSE: Do whatever the optimal function will do.
Optimal+ is superior to the optimal function because it is exactly the same except that it gets better results in the twin PD (because the opponent will cooperate too, because they are running the same algorithm as you).
Contradiction! Looks like our "optimal function" wasn't optimal after all. Therefore the real optimal function must cooperate in the twin PD.
Generalizing this reasoning, Kant says, the optimal function will choose as if it is choosing for all instances of the optimal function in similar situations. Thus we can conclude the following interesting fact: Regardless of what your goals are, the optimal function will tell you to avoid doing things that you wouldn’t want other rational agents in similar situations to do. (rational agents := agents obeying the optimal function.)
To understand this, and see how it generalizes still further, I hereby introduce the following analogy:
The Decision Theory App Store
Imagine an ideal competitive market for advice-giving AI assistants. Tech companies code them up and then you download them for free from the app store. There is AlphaBot, MetaBot, OpenBot, DeepBot.
When installed, the apps give advice. Specifically they scan your brain to extract your credences and values/utility function, and then they tell you what to do. You can follow the advice or not.
Sometimes users end up in Twin Prisoner’s Dilemmas. That is, situations where they are in some sort of prisoner’s dilemma with someone else where there is common knowledge that they both are likely t...