Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: AXRP Episode 18 - Concept Extrapolation with Stuart Armstrong, published by DanielFilan on September 3, 2022 on The AI Alignment Forum.
Google Podcasts link
Concept extrapolation is the idea of taking concepts an AI has about the world - say, “mass” or “does this picture contain a hot dog” - and extending them sensibly to situations where things are different - like learning that the world works via special relativity, or seeing a picture of a novel sausage-bread combination. For a while, Stuart Armstrong has been thinking about concept extrapolation and how it relates to AI alignment. In this episode, we discuss where his thoughts are at on this topic, what the relationship to AI alignment is, and what the open questions are.
Topics we discuss:
What is concept extrapolation
When is concept extrapolation possible
A toy formalism
Uniqueness of extrapolations
Unity of concept extrapolation methods
Concept extrapolation and corrigibility
Is concept extrapolation possible?
Misunderstandings of Stuart’s approach
Following Stuart’s work
Daniel Filan: Hello, everybody. In this episode, I’ll be speaking with Stuart Armstrong. Stuart was previously a senior researcher at the Future of Humanity Institute at Oxford, where he worked on AI safety and x-risk, as well as how to spread between galaxies by disassembling the planet Mercury. He’s currently the head boffin at Aligned AI, where he works on concept extrapolation, the subject of our discussion. For links to what we’re discussing, you can check the description of this episode, and you can read the transcript at axrp.net. Well, Stuart, welcome to the show.
Stuart Armstrong: Thank you.
Daniel Filan: Cool.
Stuart Armstrong: Good to be on.
What is concept extrapolation
Daniel Filan: Yeah, it’s nice to have you. So I guess the thing I want to be talking about today is your work and your thoughts on concept extrapolation and model splintering, which I guess you’ve called it. Can you just tell us: what is concept extrapolation?
Stuart Armstrong: Model splintering is when the features or the concepts on which you built your goals or your reward functions break down. Traditional examples are in physics when (for instance) the ether disappeared. It didn’t mean when the ether disappeared that all the previous physics that had been based on ether suddenly became completely wrong. You had to extend the old results into a new framework. You had to find a new framework and you had to extend it. So model splintering is when the model falls apart or the features or concepts fall apart and concept extrapolation is what you do to extend the concept across that divide.
Daniel Filan: Okay.
Stuart Armstrong: Like there was a concept of energy before relativity, and there’s a concept of energy after relativity. They’re not exactly the same thing, but there’s a definite continuity to it.
Daniel Filan: Cool. So you mentioned that at some point we used to think there was ether, and now we think there isn’t. What’s an example of a concept or something that splintered when we realized there wasn’t an ether anymore, just to get a really concrete example.
Stuart Armstrong: Maxwell’s equations - Maxwell’s non-relativistic equations are based on a non-constant speed of light. Maxwell’s equations are not relativistic, though they have a relativistic formulation.
Daniel Filan: Hang on. I thought they were. Isn’t that why you get the constant speed of light out of them?
Stuart Armstrong: Okay. If that’s uncertain, then let’s try another example.
Daniel Filan: We could do energy, once you discovered general relativity, or.
Stuart Armstrong: Energy, inertial mass, for example: those concepts needed a Newtonian universe to make sense, when it wasn’t so much the absence of ether, but it was the surprisingly constant speed of light that broke those. So when you m...