December 02, 2024

“39 - Evan Hubinger on Model Organisms of Misalignment” by DanielFilan

1 hour 53 minutes

YouTube link

The ‘model organisms of misalignment’ line of research creates AI models that exhibit various types of misalignment, and studies them to try to understand how the misalignment occurs and whether it can be somehow removed. In this episode, Evan Hubinger talks about two papers he's worked on at Anthropic under this agenda: “Sleeper Agents” and “Sycophancy to Subterfuge”.

Topics we discuss:

Model organisms and stress-testing

Sleeper Agents

Do ‘sleeper agents’ properly model deceptive alignment?

Surprising results in “Sleeper Agents”

Sycophancy to Subterfuge

How models generalize from sycophancy to subterfuge

Is the reward editing task valid?

Training away sycophancy and subterfuge

Model organisms, AI control, and evaluations

Other model organisms research

Alignment stress-testing at Anthropic

Following Evan's work

Daniel Filan:

Hello, everybody. In this episode, I’ll be speaking with Evan Hubinger. [...]

---

Outline:

(01:46) Model organisms and stress-testing

(09:02) Sleeper Agents

(25:18) Do ‘sleeper agents’ properly model deceptive alignment?

(42:08) Surprising results in “Sleeper Agents”

(01:02:51) Sycophancy to Subterfuge

(01:15:27) How models generalize from sycophancy to subterfuge

(01:23:27) Is the reward editing task valid?

(01:28:53) Training away sycophancy and subterfuge

(01:36:42) Model organisms, AI control, and evaluations

(01:41:12) Other model organisms research

(01:43:11) Alignment stress-testing at Anthropic

(01:51:07) Following Evan's work

---

First published:

December 1st, 2024

Source:

https://www.lesswrong.com/posts/sookiqxkzzLmPYB3r/39-evan-hubinger-on-model-organisms-of-misalignment

---

Narrated by TYPE III AUDIO.

...more

View all episodes

By LessWrong

December 02, 2024

“39 - Evan Hubinger on Model Organisms of Misalignment” by DanielFilan

1 hour 53 minutes

YouTube link

Topics we discuss:

Model organisms and stress-testing

Sleeper Agents

Do ‘sleeper agents’ properly model deceptive alignment?

Surprising results in “Sleeper Agents”

Sycophancy to Subterfuge

How models generalize from sycophancy to subterfuge

Is the reward editing task valid?

Training away sycophancy and subterfuge

Model organisms, AI control, and evaluations

Other model organisms research

Alignment stress-testing at Anthropic

Following Evan's work

Daniel Filan:

Hello, everybody. In this episode, I’ll be speaking with Evan Hubinger. [...]

---

Outline:

(01:46) Model organisms and stress-testing

(09:02) Sleeper Agents

(25:18) Do ‘sleeper agents’ properly model deceptive alignment?

(42:08) Surprising results in “Sleeper Agents”

(01:02:51) Sycophancy to Subterfuge

(01:15:27) How models generalize from sycophancy to subterfuge

(01:23:27) Is the reward editing task valid?

(01:28:53) Training away sycophancy and subterfuge

(01:36:42) Model organisms, AI control, and evaluations

(01:41:12) Other model organisms research

(01:43:11) Alignment stress-testing at Anthropic

(01:51:07) Following Evan's work

---

First published:

December 1st, 2024

Source:

https://www.lesswrong.com/posts/sookiqxkzzLmPYB3r/39-evan-hubinger-on-model-organisms-of-misalignment

---

Narrated by TYPE III AUDIO.

...more

More shows like LessWrong (30+ Karma)

View all

The Daily

112,193 Listeners

Astral Codex Ten Podcast

131 Listeners

Interesting Times with Ross Douthat

7,227 Listeners

Dwarkesh Podcast

564 Listeners

The Ezra Klein Show

16,216 Listeners

AI Article Readings

4 Listeners

Doom Debates!

14 Listeners

LessWrong posts by zvi

2 Listeners

Share “39 - Evan Hubinger on Model Organisms of Misalignment” by DanielFilan

Sign up to save your podcasts

“39 - Evan Hubinger on Model Organisms of Misalignment” by DanielFilan

“39 - Evan Hubinger on Model Organisms of Misalignment” by DanielFilan

More shows like LessWrong (30+ Karma)

The Daily

Astral Codex Ten Podcast

Interesting Times with Ross Douthat

Dwarkesh Podcast

The Ezra Klein Show

AI Article Readings

Doom Debates!

LessWrong posts by zvi