LessWrong (30+ Karma)

“Protocol evaluations: good analogies vs control” by Fabien Roger


Listen Later

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.

Let's say you want to use AIs that are capable of causing x-risk. You have a plan that describes how to train, validate and deploy AIs. For example, you could use pretraining+RLHF, check that it doesn’t fall for honeypots and never inserts a single backdoor in your coding validation set, and have humans look at the deployment outputs GPT-4 finds the most suspicious. How do you know if this plan allows you to use these potentially dangerous AIs safely?

Evaluate safety using very good analogies for the real situation?

If you want to evaluate a protocol (a set of techniques you use to train, validate, and deploy a model), then you can see how well it works in a domain where you have held-out validation data that you can use to check if your protocol works. That [...]

---

Outline:

(00:40) Evaluate safety using very good analogies for the real situation?

(02:40) The string theory analogy

(04:12) Noticing rare failures

(04:55) Scheming AIs can game these sandwiching experiments

(05:44) Evaluating how robust protocols are against scheming AIs

(06:19) Malign initializations

(09:29) Control evaluations (Black-box)

(12:07) Aggregating non-adversarial and adversarial evaluations

(12:51) Theoretically-sound arguments about protocols and AIs

(13:50) Other non-adversarial evaluations

(16:44) Overview of evaluation methods

(20:02) Appendix: Analogies vs high-stakes reward hacking

The original text contained 8 footnotes which were omitted from this narration.

---

First published:

February 19th, 2024

Source:

https://www.lesswrong.com/posts/qhaSoR6vGmKnqGYLE/protocol-evaluations-good-analogies-vs-control

---

Narrated by TYPE III AUDIO.

...more
View all episodesView all episodes
Download on the App Store

LessWrong (30+ Karma)By LessWrong


More shows like LessWrong (30+ Karma)

View all
The Daily by The New York Times

The Daily

113,164 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

130 Listeners

Interesting Times with Ross Douthat by New York Times Opinion

Interesting Times with Ross Douthat

7,255 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

535 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

16,266 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates by Liron Shapira

Doom Debates

14 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners