LessWrong (30+ Karma)

“~80 Interesting Questions about Foundation Model Agent Safety” by RohanS, Govind Pimpale


Listen Later

Many people helped us a great deal in developing the questions and ideas in this post, including people at CHAI, MATS, various other places in Berkeley, and Aether. To all of them: Thank you very much! Any mistakes are our own.

Foundation model agents - systems like AutoGPT and Devin that equip foundations models with planning, memory, tool use, and other affordances to perform autonomous tasks - seem to have immense implications for AI capabilities and safety. As such, I (Rohan) am planning to do foundation model agent safety research.

Following the spirit of an earlier post I wrote, I thought it would be fun and valuable write as many interesting questions as I could about foundation model agent safety. I shared these questions with my collaborators, and Govind wrote a bunch more questions that he is interested in. This post includes questions from both of us. [...]

---

Outline:

(01:14) Rohan

(01:28) Basics and Current Status

(03:16) Chain-of-Thought (CoT) Interpretability

(08:02) Goals

(10:18) Forecasting (Technical and Sociological)

(16:43) Broad Conceptual Safety Questions

(21:50) Miscellaneous

(25:21) Govind

(25:24) OpenAI o1 and other RL CoT Agents

(26:30) Linguistic Drift, Neuralese, and Steganography

(27:32) Agentic Performance

(28:57) Forecasting

---

First published:

October 28th, 2024

Source:

https://www.lesswrong.com/posts/ZJzyDdKsDhFvqQhdQ/80-interesting-questions-about-foundation-model-agent-safety

---

Narrated by TYPE III AUDIO.

...more
View all episodesView all episodes
Download on the App Store

LessWrong (30+ Karma)By LessWrong


More shows like LessWrong (30+ Karma)

View all
The Daily by The New York Times

The Daily

113,069 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

132 Listeners

Interesting Times with Ross Douthat by New York Times Opinion

Interesting Times with Ross Douthat

7,281 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

558 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

16,464 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates by Liron Shapira

Doom Debates

14 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners