
Sign up to save your podcasts
Or


tldr: We share a toy environment that we found useful for understanding how reasoning changed over the course of capabilities-focused RL. Over the course of capabilities-focused RL, the model biases more strongly towards reward hints over direct instruction in this environment.
Setup
When we noticed the increase in verbalized alignment evaluation awareness during capabilities-focused RL, we initially thought that the right mental model was something like:
However, qualitatively neither of these seemed particularly salient to the model:
So we wanted to create a minimal environment where:
---
Outline:
(00:29) Setup
(01:57) Results
(02:08) Gaming is coherent across naming of the reward field
(03:03) Gaming is relatively robust to paraphrasing instructions
(03:26) The model is willing to exploit increasingly difficult hints
(05:15) The model is not just confused
(06:00) The model is increasingly insensitive to threats of auditing or human review
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
tag" with confidence intervals." style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
By LessWrongtldr: We share a toy environment that we found useful for understanding how reasoning changed over the course of capabilities-focused RL. Over the course of capabilities-focused RL, the model biases more strongly towards reward hints over direct instruction in this environment.
Setup
When we noticed the increase in verbalized alignment evaluation awareness during capabilities-focused RL, we initially thought that the right mental model was something like:
However, qualitatively neither of these seemed particularly salient to the model:
So we wanted to create a minimal environment where:
---
Outline:
(00:29) Setup
(01:57) Results
(02:08) Gaming is coherent across naming of the reward field
(03:03) Gaming is relatively robust to paraphrasing instructions
(03:26) The model is willing to exploit increasingly difficult hints
(05:15) The model is not just confused
(06:00) The model is increasingly insensitive to threats of auditing or human review
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
tag" with confidence intervals." style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

111,948 Listeners

130 Listeners

7,230 Listeners

576 Listeners

15,950 Listeners

4 Listeners

14 Listeners

2 Listeners