July 14, 2025

“Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance” by Senthooran Rajamanoharan, Neel Nanda

19 minutes

This is a write-up of a brief investigation into shutdown resistance undertaken by the Google DeepMind interpretability team.

TL;DR

Why do models sometimes resist shutdown? Are they ignoring instructions to pursue their own agenda – in this case, self-preservation? Or is there a more prosaic explanation? We investigated a specific agentic environment introduced by Palisade Research, where shutdown resistance has previously been reported. By analysing Gemini 2.5 Pro's reasoning, we found the behaviour stems from a misguided attempt to complete what it perceives as the primary goal. When we explicitly clarify in the prompt that shutdown compliance takes priority, this resistance vanishes. These same clarified instructions also eliminate shutdown subversion in OpenAI's o3 and o4-mini. We also check what happens when we remove the goal conflict entirely: when asked to shut down only after completing their task, the models comply perfectly. Our observations offer a simpler explanation for shutdown [...]

---

Outline:

(00:23) TL;DR

(01:41) Introduction

(06:18) What did we take away from this?

(09:21) Comparing our results with Palisade's follow-up findings

(11:10) Experiment details

(11:14) Original setting

(13:23) Adding a shutdown warning

(14:48) Clarifying instruction precedence

(15:57) Raising the stakes

(17:22) Control experiment: removing the goal conflict

(18:18) Acknowledgements

(18:34) Appendix

The original text contained 3 footnotes which were omitted from this narration.

---

First published:

July 14th, 2025

Source:

https://www.lesswrong.com/posts/wnzkjSmrgWZaBa2aC/self-preservation-or-instruction-ambiguity-examining-the

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more

View all episodes

By LessWrong

July 14, 2025

“Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance” by Senthooran Rajamanoharan, Neel Nanda

19 minutes

This is a write-up of a brief investigation into shutdown resistance undertaken by the Google DeepMind interpretability team.

TL;DR

---

Outline:

(00:23) TL;DR

(01:41) Introduction

(06:18) What did we take away from this?

(09:21) Comparing our results with Palisade's follow-up findings

(11:10) Experiment details

(11:14) Original setting

(13:23) Adding a shutdown warning

(14:48) Clarifying instruction precedence

(15:57) Raising the stakes

(17:22) Control experiment: removing the goal conflict

(18:18) Acknowledgements

(18:34) Appendix

The original text contained 3 footnotes which were omitted from this narration.

---

First published:

July 14th, 2025

Source:

https://www.lesswrong.com/posts/wnzkjSmrgWZaBa2aC/self-preservation-or-instruction-ambiguity-examining-the

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more

More shows like LessWrong (30+ Karma)

View all

Making Sense with Sam Harris

26,386 Listeners

Conversations with Tyler

2,419 Listeners

The Peter Attia Drive

8,916 Listeners

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

4,153 Listeners

ManifoldOne

92 Listeners

Your Undivided Attention

1,595 Listeners

All-In with Chamath, Jason, Sacks & Friedberg

9,900 Listeners

Machine Learning Street Talk (MLST)

90 Listeners

Dwarkesh Podcast

76 Listeners

Hard Fork

5,470 Listeners

The Ezra Klein Show

16,026 Listeners

Moonshots with Peter Diamandis

539 Listeners

No Priors: Artificial Intelligence | Technology | Startups

130 Listeners

Latent Space: The AI Engineer Podcast

94 Listeners

BG2Pod with Brad Gerstner and Bill Gurley

504 Listeners

Share “Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance” by Senthooran Rajamanoharan, Neel Nanda

Sign up to save your podcasts

“Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance” by Senthooran Rajamanoharan, Neel Nanda

“Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance” by Senthooran Rajamanoharan, Neel Nanda

More shows like LessWrong (30+ Karma)

Making Sense with Sam Harris

Conversations with Tyler

The Peter Attia Drive

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

ManifoldOne

Your Undivided Attention

All-In with Chamath, Jason, Sacks & Friedberg

Machine Learning Street Talk (MLST)

Dwarkesh Podcast

Hard Fork

The Ezra Klein Show

Moonshots with Peter Diamandis

No Priors: Artificial Intelligence | Technology | Startups

Latent Space: The AI Engineer Podcast

BG2Pod with Brad Gerstner and Bill Gurley