LessWrong (30+ Karma)

“What Is The Alignment Problem?” by johnswentworth


Listen Later

So we want to align future AGIs. Ultimately we’d like to align them to human values, but in the shorter term we might start with other targets, like e.g. corrigibility.

That problem description all makes sense on a hand-wavy intuitive level, but once we get concrete and dig into technical details… wait, what exactly is the goal again? When we say we want to “align AGI”, what does that mean? And what about these “human values” - it's easy to list things which are importantly not human values (like stated preferences, revealed preferences, etc), but what are we talking about? And don’t even get me started on corrigibility!

Turns out, it's surprisingly tricky to explain what exactly “the alignment problem” refers to. And there's good reasons for that! In this post, I’ll give my current best explanation of what the alignment problem is (including a few variants and the [...]

---

Outline:

(01:27) The Difficulty of Specifying Problems

(01:50) Toy Problem 1: Old MacDonald's New Hen

(04:08) Toy Problem 2: Sorting Bleggs and Rubes

(06:55) Generalization to Alignment

(08:54) But What If The Patterns Don't Hold?

(13:06) Alignment of What?

(14:01) Alignment of a Goal or Purpose

(19:47) Alignment of Basic Agents

(23:51) Alignment of General Intelligence

(27:40) How Does All That Relate To Todays AI?

(31:03) Alignment to What?

(32:01) What are a Humans Values?

(36:14) Other targets

(36:43) Paul!Corrigibility

(39:11) Eliezer!Corrigibility

(40:52) Subproblem!Corrigibility

(42:55) Exercise: Do What I Mean (DWIM)

(43:26) Putting It All Together, and Takeaways

The original text contained 10 footnotes which were omitted from this narration.

---

First published:

January 16th, 2025

Source:

https://www.lesswrong.com/posts/dHNKtQ3vTBxTfTPxu/what-is-the-alignment-problem

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more
View all episodesView all episodes
Download on the App Store

LessWrong (30+ Karma)By LessWrong


More shows like LessWrong (30+ Karma)

View all
Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,324 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,403 Listeners

The Peter Attia Drive by Peter Attia, MD

The Peter Attia Drive

7,916 Listeners

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas by Sean Carroll | Wondery

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

4,114 Listeners

ManifoldOne by Steve Hsu

ManifoldOne

87 Listeners

Your Undivided Attention by Tristan Harris and Aza Raskin, The Center for Humane Technology

Your Undivided Attention

1,446 Listeners

All-In with Chamath, Jason, Sacks & Friedberg by All-In Podcast, LLC

All-In with Chamath, Jason, Sacks & Friedberg

8,775 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

90 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

355 Listeners

Hard Fork by The New York Times

Hard Fork

5,375 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,295 Listeners

Moonshots with Peter Diamandis by PHD Ventures

Moonshots with Peter Diamandis

472 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

126 Listeners

Latent Space: The AI Engineer Podcast by swyx + Alessio

Latent Space: The AI Engineer Podcast

74 Listeners

BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

BG2Pod with Brad Gerstner and Bill Gurley

443 Listeners