
Sign up to save your podcasts
Or


YouTube link
The ‘model organisms of misalignment’ line of research creates AI models that exhibit various types of misalignment, and studies them to try to understand how the misalignment occurs and whether it can be somehow removed. In this episode, Evan Hubinger talks about two papers he's worked on at Anthropic under this agenda: “Sleeper Agents” and “Sycophancy to Subterfuge”.
Topics we discuss:
Daniel Filan:
---
Outline:
(01:46) Model organisms and stress-testing
(09:02) Sleeper Agents
(25:18) Do ‘sleeper agents’ properly model deceptive alignment?
(42:08) Surprising results in “Sleeper Agents”
(01:02:51) Sycophancy to Subterfuge
(01:15:27) How models generalize from sycophancy to subterfuge
(01:23:27) Is the reward editing task valid?
(01:28:53) Training away sycophancy and subterfuge
(01:36:42) Model organisms, AI control, and evaluations
(01:41:12) Other model organisms research
(01:43:11) Alignment stress-testing at Anthropic
(01:51:07) Following Evan's work
---
First published:
Source:
Narrated by TYPE III AUDIO.
By LessWrongYouTube link
The ‘model organisms of misalignment’ line of research creates AI models that exhibit various types of misalignment, and studies them to try to understand how the misalignment occurs and whether it can be somehow removed. In this episode, Evan Hubinger talks about two papers he's worked on at Anthropic under this agenda: “Sleeper Agents” and “Sycophancy to Subterfuge”.
Topics we discuss:
Daniel Filan:
---
Outline:
(01:46) Model organisms and stress-testing
(09:02) Sleeper Agents
(25:18) Do ‘sleeper agents’ properly model deceptive alignment?
(42:08) Surprising results in “Sleeper Agents”
(01:02:51) Sycophancy to Subterfuge
(01:15:27) How models generalize from sycophancy to subterfuge
(01:23:27) Is the reward editing task valid?
(01:28:53) Training away sycophancy and subterfuge
(01:36:42) Model organisms, AI control, and evaluations
(01:41:12) Other model organisms research
(01:43:11) Alignment stress-testing at Anthropic
(01:51:07) Following Evan's work
---
First published:
Source:
Narrated by TYPE III AUDIO.

26,332 Listeners

2,453 Listeners

8,579 Listeners

4,183 Listeners

93 Listeners

1,598 Listeners

9,932 Listeners

95 Listeners

511 Listeners

5,518 Listeners

15,938 Listeners

546 Listeners

131 Listeners

93 Listeners

467 Listeners