Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: P₂B: Plan to P₂B Better, published by Ramana Kumar on October 24, 2021 on The AI Alignment Forum.
tl;dr: Most good plans involve taking steps to make better plans. Making better plans is the convergent instrumental goal, of which all familiar convergent instrumental goals are an instance. This is key to understanding what agency is and why it is powerful.
Planning means using a world model to predict the consequences of various courses of actions one could take, and taking actions that have good predicted consequences. (We think of this with the handle “doing things for reasons,” though we acknowledge this may be an idiosyncratic use of “reasons.”)
We take “planning” to include things that are relevantly similar to this procedure, such as following a bag of heuristics that approximates it. We’re also including actually following the plans, in what might more clunkily be called “planning-acting.”
Planning, in this broad sense, seems essential to the kind of goal-directed, consequential, agent-like intelligence that we expect to be highly impactful. This sequence explains why.
One Convergent Instrumental Goal to Rule them All
Consider the maxim
“make there be more and/or better planning towards your goal.”
This section argues that all the classic convergent instrumental goals are special cases of this maxim.
To flesh this out a little, here are some categories of ways to follow the maxim. Remember that a planner is typically close (in terms of what it might affect via action) to at least one planner – itself – so these directions can typically be applied in the first case to the planner itself.
Make the planners with your goal better at planning. For example, get them new relevant data to work with, get them to run faster or more effective algorithms, build protections against value drift, etc.
Make the planners with your goal have better options. For example, move them to better locations, get them more resources, get them more power or a greater number of options to select from, have them take steps in an object-level plan towards the goal.
Make there be more planners with your goal. For example, keep yourself running and aligned with your goal, acquire delegates and subordinates, convince followers and converts, build successors.
Reviewing Omohundro's “The Basic AI Drives” and Bostrom’s “The Superintelligent Will,” we extract a list of convergent instrumental goals, and find that they are all instances of the maxim “make there be more/better planners for the current goal:”
Self-preservation / self-protection:
Make there be more planners that have your goals, focusing on reusing the existing planner, that is, preventing its destruction.
Self-improvement:
Make there be better planners with your goals, focusing on making the existing one better.
Resource Acquisition:
Make there be better planners with your goals, focusing on making the existing one able to take more effective actions.
Goal-content Integrity:
Make there be more planners that have your goals, focusing on ensuring the existing planners that have your goals keep those goals and avoid them being changed.
Resource-use Efficiency:
Cognitive Enhancement:
Creativity:
Same as self-improvement
Technological Perfection:
Same as resource acquisition and/or self-improvement
Rationality:
Omohundro takes this to be something like “make the utility function explicit” along with “maximize expected utility.”
Thus it is similar to goal-content integrity and self-improvement.
Utility-function preservation:
Similar to goal-content integrity.
Prevent counterfeit utility:
Essentially this is avoiding wireheading. Omohundro: “An important class of vulnerabilities arises when the subsystems for measuring utility become corrupted.”
Thus it is similar to goal-content integrity.
Seeking a concise, memorable-yet-accurate name for t...