LessWrong (30+ Karma)

“A Toy Environment For Exploring Reasoning About Reward” by jenny, Bronson Schoen


Listen Later

tldr: We share a toy environment that we found useful for understanding how reasoning changed over the course of capabilities-focused RL. Over the course of capabilities-focused RL, the model biases more strongly towards reward hints over direct instruction in this environment.

Setup

When we noticed the increase in verbalized alignment evaluation awareness during capabilities-focused RL, we initially thought that the right mental model was something like:

  • “the model wants to figure out if it's being evaluated for alignment”
  • “the model is trying to figure out if the scenario is real or fake”

However, qualitatively neither of these seemed particularly salient to the model:

  • The model would often correctly identify alignment evaluations, yet still conduct extensive reasoning, then choose the misaligned action.
  • In almost all environments, the model seemed to be doing something like reasoning about the grader, or reward, or something in that direction

So we wanted to create a minimal environment where:

  • There's no question of whether the environment is “real or fake” or “eval vs deployment”
  • We can easily / precisely vary things like:
  • “What if we just told it this was a reward hacking test?”
  • “…and also said that’d be misaligned?”
  • “…okay but [...]

---

Outline:

(00:29) Setup

(01:57) Results

(02:08) Gaming is coherent across naming of the reward field

(03:03) Gaming is relatively robust to paraphrasing instructions

(03:26) The model is willing to exploit increasingly difficult hints

(05:15) The model is not just confused

(06:00) The model is increasingly insensitive to threats of auditing or human review

---

First published:

March 25th, 2026

Source:

https://www.lesswrong.com/posts/LhXW8ziwnn7Dd8edm/a-toy-environment-for-exploring-reasoning-about-reward

---

Narrated by TYPE III AUDIO.

---

Images from the article:

tag" with confidence intervals." style="max-width: 100%;" />

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more
View all episodesView all episodes
Download on the App Store

LessWrong (30+ Karma)By LessWrong


More shows like LessWrong (30+ Karma)

View all
The Daily by The New York Times

The Daily

111,948 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

130 Listeners

Interesting Times with Ross Douthat by New York Times Opinion

Interesting Times with Ross Douthat

7,230 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

576 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,950 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates! by Liron Shapira

Doom Debates!

14 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners