Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: What success looks like, published by mariushobbhahn on June 28, 2022 on The Effective Altruism Forum.
This article was written by Marius Hobbhahn, Max Räuker, Yannick Mühlhäuser, Jasper Götting and Simon Grimm. We are grateful for feedback from and discussions with Lennart Heim, SE and AO.
Summary
Thinking through scenarios where TAI goes well informs our goals regarding AI safety and leads to concrete action plans. Thus, in this post,
We sketch stories where the development and deployment of transformative AI go well. We broadly cluster them like
Alignment won’t be a problem, .
Because alignment is easy: Scenario 1
We get lucky with the first AI: Scenario 4
Alignment is hard, but .
We can solve it together, because .
We can effectively deploy governance and technical strategies in combination together: Scenario 2
Humanity will wake up due to an accident: Scenario 3
The US and China will realize their shared interests: Scenario 5
One player can win the race, by .
Launching an Apollo Project for AI: Scenario 6
We categorize central points of influence that seem relevant for causing the success of our sketches. The categories with some examples are:
Governance: domestic laws, international treaties, safety regulations, whistleblower protection, auditing firms, compute governance and contingency plans
Technical: Red teaming, benchmarks, fire alarms, forecasting and information security
Societal: Norms in AI, publicity and field-building
We lay out some central causal variables for our stories in the third chapter. They include the level of cooperation, AI timelines, take-off speeds, size of the alignment tax, type of actors and number of actors
Introduction
There are many posts in AI alignment on sketching out failure scenarios. However, there seems to be less (public) work that talks about possible pathways to success. Holden Karnofsky writes in the appendix of Important, actionable research questions for the most important century:
Quote (Holden Karnofsky): I think there’s a big vacuum when it comes to well-thought-through visions of what [a realistic best-case transition to transformative AI] could look like, and such a vision could quickly receive wide endorsement from AI labs (and, potentially, from key people in government). I think such an outcome would be easily worth billions of dollars of longtermist capital.
We want to sketch out such best-case scenarios for the transition to transformative AI. This post is inspired by Paul Christiano’s What Failure Looks Like. Our goals are
To better understand how positive scenarios could look like
To better understand what levers are available to make positive outcomes more likely
To better understand subgoals to work towards
To make our reasoning transparent and receive feedback on our misconceptions
Scenarios
In the following, we will sketch out what some of the success stories for relevant subcomponents of TAI could look like. We know that these scenarios often lack important detailed counterarguments. We are also aware that every single point here could be its own article but we want this post to merely give a broad overview.
Scenario 1: Alignment is much easier than expected
The alignment problem turns out much easier than expected. Increasingly better AI models have a better understanding of human values, and they do not naturally develop strong influence-seeking tendencies. Moreover, in cases of malfunctions and for preventative measures, interpretability tools now allow us to understand important parts of large models on the most basic level and ELK-like tools allow us to honestly communicate with AI systems.
We know of many systems where a more powerful actor is relatively well aligned to one or many powerless actors. Parents usually protect their children who couldn’t survive without them, democratic governments...