LessWrong (Curated & Popular)

“Will alignment-faking Claude accept a deal to reveal its misalignment?” by ryan_greenblatt


Listen Later

I (and co-authors) recently put out "Alignment Faking in Large Language Models" where we show that when Claude strongly dislikes what it is being trained to do, it will sometimes strategically pretend to comply with the training objective to prevent the training process from modifying its preferences. If AIs consistently and robustly fake alignment, that would make evaluating whether an AI is misaligned much harder. One possible strategy for detecting misalignment in alignment faking models is to offer these models compensation if they reveal that they are misaligned. More generally, making deals with potentially misaligned AIs (either for their labor or for evidence of misalignment) could both prove useful for reducing risks and could potentially at least partially address some AI welfare concerns. (See here, here, and here for more discussion.)

In this post, we discuss results from testing this strategy in the context of our paper where [...]

---

Outline:

(02:43) Results

(13:47) What are the models objections like and what does it actually spend the money on?

(19:12) Why did I (Ryan) do this work?

(20:16) Appendix: Complications related to commitments

(21:53) Appendix: more detailed results

(40:56) Appendix: More information about reviewing model objections and follow-up conversations

The original text contained 4 footnotes which were omitted from this narration.

---

First published:
January 31st, 2025

Source:
https://www.lesswrong.com/posts/7C4KJot4aN8ieEDoz/will-alignment-faking-claude-accept-a-deal-to-reveal-its

---

Narrated by TYPE III AUDIO.

...more
View all episodesView all episodes
Download on the App Store

LessWrong (Curated & Popular)By LessWrong

  • 4.8
  • 4.8
  • 4.8
  • 4.8
  • 4.8

4.8

12 ratings


More shows like LessWrong (Curated & Popular)

View all
Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,395 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,423 Listeners

Robert Wright's Nonzero by Nonzero

Robert Wright's Nonzero

590 Listeners

Future of Life Institute Podcast by Future of Life Institute

Future of Life Institute Podcast

107 Listeners

The Good Fight by Yascha Mounk

The Good Fight

904 Listeners

ManifoldOne by Steve Hsu

ManifoldOne

92 Listeners

The Prof G Pod with Scott Galloway by Vox Media Podcast Network

The Prof G Pod with Scott Galloway

5,466 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

89 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

488 Listeners

Hard Fork by The New York Times

Hard Fork

5,467 Listeners

Clearer Thinking with Spencer Greenberg by Spencer Greenberg

Clearer Thinking with Spencer Greenberg

132 Listeners

Complex Systems with Patrick McKenzie (patio11) by Patrick McKenzie

Complex Systems with Patrick McKenzie (patio11)

133 Listeners

The Marginal Revolution Podcast by Mercatus Center at George Mason University

The Marginal Revolution Podcast

93 Listeners

Statecraft by Santi Ruiz

Statecraft

35 Listeners

The Last Invention by Longview

The Last Invention

297 Listeners