Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Some abstract, non-technical reasons to be non-maximally-pessimistic about AI alignment, published by Rob Bensinger on December 12, 2021 on LessWrong.
I basically agree with Eliezer’s picture of things in the AGI interventions post.
But I’ve seen some readers rounding off Eliezer’s ‘the situation looks very dire’-ish statements to ‘the situation is hopeless’, and ‘solving alignment still looks to me like our best shot at a good future, but so far we’ve made very little progress, we aren’t anywhere near on track to solve the problem, and it isn’t clear what the best path forward is’-ish statements to ‘let’s give up on alignment’.
It’s hard to give a technical argument for ‘alignment isn’t doomed’, because I don’t know how to do alignment (and, to the best of my knowledge in December 2021, no one else does either). But I can give some of the more abstract reasons I think that.
I feel sort of wary of sharing a ‘reasons to be less pessimistic’ list, because it’s blatantly filtered evidence, it makes it easy to overcorrect, etc. In my experience, people tend to be way too eager to classify problems as either ‘easy’ or ‘impossible’; just adding more evidence may cause people to bounce back and forth between the two rather than planting flag in the middle ground.
I did write a version of 'reasons not to be maximally pesimistic' for a few friends in 2018. I’m warily fine with sharing that below, with the caveats ‘holy shit is this ever filtered evidence!’ and ‘these my own not-MIRI-vetted personal thoughts’. And 'this is a casual thing I jotted down for friends in 2018'.
Today, I would add some points (e.g., 'AGI may be surprisingly far off; timelines are hard to predict'), and I'd remove others (e.g., 'Nate and Eliezer feel pretty good about MIRI's current research'). Also, since the list is both qualitative and one-sided, it doesn’t reflect the fact that I’m quantitatively a bit more pessimistic now than I was in 2018.
Lo:
[...S]ome of the main reasons I'm not extremely pessimistic about artificial general intelligence outcomes.
(Warning: one-sided lists of considerations can obviously be epistemically bad. I mostly mean to correct for the fact that I see a lot of rationalists who strike me as overly pessimistic about AGI outcomes. Also, I don't try to argue for most of these points in any detail; I'm just trying to share my own views for others' stack.)
1. AGI alignment is just a technical problem, and humanity actually has a remarkably good record when it comes to solving technical problems. It's historically common for crazy-seeming goals to fall to engineering ingenuity, even in the face of seemingly insuperable obstacles.
Some of the underlying causes for this are 'it's hard to predict what clever ideas are hiding in the parts of your map that aren't filled in yet', and 'it's hard to prove a universal negation'. A universal negation is what you need in order to say that there's no clever engineering solution; whereas even if you've had ten thousand failed attempts, a single existence proof — a single solution to the problem — renders those failures totally irrelevant.
2. We don't know very much yet about the alignment problem. This isn't a reason for optimism, but it's a reason not to have confident pessimism, because no confident view can be justified by a state of uncertainty. We just have to learn more and do it the hard way and see how things go.
A blank map can feel like 'it's hopeless' for various reasons, even when you don't actually have enough Bayesian evidence to assert a technical problem is hopeless. For example: you think really hard about the problem and can't come up with a solution, which to some extent feels like there just isn't a solution. And: people aren't very good at knowing which parts of their map are blank, so it may feel like there aren't mo...