Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Why The Focus on Expected Utility Maximisers?, published by DragonGod on December 27, 2022 on LessWrong.
Epistemic Status
Unsure, partially noticing my own confusion. Hoping Cunningham's Law can help resolve it.
Confusions About Arguments From Expected Utility Maximisation
Some MIRI people (e.g. Rob Bensinger) still highlight EU maximisers as the paradigm case for existentially dangerous AI systems. I'm confused by this for a few reasons:
Not all consequentialist/goal directed systems are expected utility maximisers
E.g. humans
Some recent developments make me sceptical that VNM expected utility are a natural form of generally intelligent systems
Wentworth's subagents provide a model for inexploitable agents that don't maximise a simple unitary utility function
The main requirement for subagents to be a better model than unitary agents is path dependent preferences or hidden state variables
Alternatively, subagents natively admit partial orders over preferences
If I'm not mistaken, utility functions seem to require a (static) total order over preferences
This might be a very unreasonable ask; it does not seem to describe humans, animals, or even existing sophisticated AI systems
I think the strongest implication of Wentworth's subagents is that expected utility maximisation is not the limit or idealised form of agency
Shard Theory suggests that trained agents (via reinforcement learning) form value "shards"
Values are inherently "contextual influences on decision making"
Hence agents do not have a static total order over preferences (what a utility function implies) as what preferences are active depends on the context
Preferences are dynamic (change over time), and the ordering of them is not necessarily total
This explains many of the observed inconsistencies in human decision making
A multitude of value shards do not admit analysis as a simple unitary utility function
Reward is not the optimisation target
Reinforcement learning does not select for reward maximising agents in general
Reward "upweight certain kinds of actions in certain kinds of situations, and therefore reward chisels cognitive grooves into agents"
I'm thus very sceptical that systems optimised via reinforcement learning to be capable in a wide variety of domains/tasks converge towards maximising a simple expected utility function
I am not aware that humanity actually knows training paradigms that select for expected utility maximisers
Our most capable/economically transformative AI systems are not agents and are definitely not expected utility maximisers
Such systems might converge towards general intelligence under sufficiently strong selection pressure but do not become expected utility maximisers in the limit
The do not become agents in the limit and expected utility maximisation is a particular kind of agency
I am seriously entertaining the hypothesis that expected utility maximisation is anti-natural to selection for general intelligence
I'm not under the impression that systems optimised by stochastic gradient descent to be generally capable optimisers converge towards expected utility maximisers
The generally capable optimisers produced by evolution aren't expected utility maximisers
I'm starting to suspect that "search like" optimisation processes for general intelligence do not in general converge towards expected utility maximisers
I.e. it may end up being the case that the only way to create a generally capable expected utility maximiser is to explicitly design one
And we do not know how to design capable optimisers for rich environments
We can't even design an image classifier
I currently disbelieve the strong orthogonality thesis translated to practice
While it may be in theory feasible to design systems at any intelligence level with any final goal
In practice, we cannot design capab...