July 09, 2024

“What and Why: Developmental Interpretability of Reinforcement Learning” by Garrett Baker

Listen Later

12 minutes

Introduction

I happen to be in that happy stage in the research cycle where I ask for money so I can continue to work on things I think are important. Part of that means justifying what I want to work on to the satisfaction of the people who provide that money.

This presents a good opportunity to say what I plan to work on in a more layman-friendly way, for the benefit of LessWrong, potential collaborators, interested researchers, and funders who want to read the fun version of my project proposal

It also provides the opportunity for people who are very pessimistic about the chances I end up doing anything useful by pursuing this to have their say. So if you read this (or skim it), and have critiques (or just recommendations), I'd love to hear them! Publicly or privately.

So without further ado, in this post I will [...]

---

Outline:

(00:06) Introduction

(02:40) Reinforcement learning

(02:44) Agentic AIs vs Tool AIs

(04:54) Data walls

(07:05) Values

(07:49) Ok, but concretely what will you actually do?

(11:10) Call to action

The original text contained 1 footnote which was omitted from this narration.

---

First published:

July 9th, 2024

Source:

https://www.lesswrong.com/posts/Bczmi8vjiugDRec7C/what-and-why-developmental-interpretability-of-reinforcement

---

Narrated by TYPE III AUDIO.

...more

View all episodes

View all episodes

Download on the App Store

Download on the App Store

Get it on Google Play

LessWrong (30+ Karma)

By LessWrong

July 09, 2024

“What and Why: Developmental Interpretability of Reinforcement Learning” by Garrett Baker

Listen Later

12 minutes

Introduction

I happen to be in that happy stage in the research cycle where I ask for money so I can continue to work on things I think are important. Part of that means justifying what I want to work on to the satisfaction of the people who provide that money.

This presents a good opportunity to say what I plan to work on in a more layman-friendly way, for the benefit of LessWrong, potential collaborators, interested researchers, and funders who want to read the fun version of my project proposal

It also provides the opportunity for people who are very pessimistic about the chances I end up doing anything useful by pursuing this to have their say. So if you read this (or skim it), and have critiques (or just recommendations), I'd love to hear them! Publicly or privately.

So without further ado, in this post I will [...]

---

Outline:

(00:06) Introduction

(02:40) Reinforcement learning

(02:44) Agentic AIs vs Tool AIs

(04:54) Data walls

(07:05) Values

(07:49) Ok, but concretely what will you actually do?

(11:10) Call to action

The original text contained 1 footnote which was omitted from this narration.

---

First published:

July 9th, 2024

Source:

https://www.lesswrong.com/posts/Bczmi8vjiugDRec7C/what-and-why-developmental-interpretability-of-reinforcement

---

Narrated by TYPE III AUDIO.

...more

More shows like LessWrong (30+ Karma)

The Daily by The New York Times

The Daily

112,144 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

131 Listeners

Interesting Times with Ross Douthat by New York Times Opinion

Interesting Times with Ross Douthat

7,238 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

577 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

16,139 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates! by Liron Shapira

Doom Debates!

14 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners