Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: AI Safety and Neighboring Communities: A Quick-Start Guide, as of Summer 2022, published by Sam Bowman on September 1, 2022 on LessWrong.
Getting into AI safety involves working with a mix of communities, subcultures, goals, and ideologies that you may not have encountered in the context of mainstream AI technical research. This document attempts to briefly map these out for newcomers.
This is inevitably going to be biased by what sides of these communities I (Sam) have encountered, and it will quickly become dated. I expect it will still be a useful resource for some people anyhow, at least in the short term.
AI Safety/AI Alignment/AGI Safety/AI Existential Safety/AI X-Risk
The research project of ensuring that future AI progress doesn’t yield civilization-endingly catastrophic results.
Good intros:
Carlsmith Report
What misalignment looks like as capabilities scale
Vox piece
Why are people concerned about this?
My rough summary:
It’s plausible that future AI systems could be much faster or more effective than us at real-world reasoning and planning.
Probably not plain generative models, but possibly models derived from generative models in cheap ways
Once you have a system with superhuman reasoning and planning abilities, it’s easy to make it dangerous by accident.
Most simple objective functions or goals become dangerous in the limit, usually because of secondary or instrumental subgoals that emerge along the way.
Pursuing typical goals arbitrarily well requires a system to prevent itself from being turned off, by deception or force if needed.
Pursuing typical goals arbitrarily well requires acquiring any power or resources that could increase the chances of success, by deception or force if needed.
Toy example: Computing pi to an arbitrarily high precision eventually requires that you spend all the sun’s energy output on computing.
Knowledge and values are likely to be orthogonal: A model could know human values and norms well, but not have any reason to act on them. For agents built around generative models, this is the default outcome.
Sufficiently powerful AI systems could look benign in pre-deployment training/research environments, because they would be capable of understanding that they’re not yet in a position to accomplish their goals.
Simple attempts to work around this (like the more abstract goal ‘do what your operators want’) don’t tend to have straightforward robust implementations.
If such a system were single-mindedly pursuing a dangerous goal, we probably wouldn’t be able to stop it.
Superhuman reasoning and planning would give models with a sufficiently good understanding of the world many ways to effectively gain power with nothing more than an internet connection. (ex: Cyberattacks on banks.)
Consensus within the field is that these risks could become concrete within ~4–25 years, and have a >10% chance of being leading to a global catastrophe (i.e., extinction or something comparably bad). If true, it’s bad news.
Given the above, we either need to stop all development toward AGI worldwide (plausibly undesirable or impossible), or else do three possible-but-very-difficult things:
(i) build robust techniques to align AGI systems with the values and goals of their operators,
(ii) ensure that those techniques are understood and used by any group that could plausibly build AGI, and
(iii) ensure that we're able to govern the operators of AGI systems in a way that makes their actions broadly positive for humanity as a whole.
Does this have anything to do with sentience or consciousness?
No.
Influential people and institutions:
Present core community as I see it: Paul Christiano, Jacob Steinhardt, Ajeya Cotra, Jared Kaplan, Jan Leike, Beth Barnes, Geoffrey Irving, Buck Shlegeris, David Krueger, Chris Olah, Evan Hubinger, Richard Ngo, Rohin Shah; ARC, R...