November 03, 2024

“Current safety training techniques do not fully transfer to the agent setting” by Simon Lermen, Govind Pimpale

Listen Later

10 minutes

TL;DR: I'm presenting three recent papers which all share a similar finding, i.e. the safety training techniques for chat models don’t transfer well from chat models to the agents built from them. In other words, models won’t tell you how to do something harmful, but they are often willing to directly execute harmful actions. However, all papers find that different attack methods like jailbreaks, prompt-engineering, or refusal-vector ablation do transfer.

Here are the three papers:

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
Applying Refusal-Vector Ablation to Llama 3.1 70B Agents

What are language model agents

Language model agents are a combination of a language model and a scaffolding software. Regular language models are typically limited to being chat bots, i.e. they receive messages and reply to them. However, scaffolding gives these models access to tools which they can [...]

---

Outline:

(00:55) What are language model agents

(01:36) Overview

(03:31) AgentHarm Benchmark

(05:27) Refusal-Trained LLMs Are Easily Jailbroken as Browser Agents

(06:47) Applying Refusal-Vector Ablation to Llama 3.1 70B Agents

(08:23) Discussion

---

First published:

November 3rd, 2024

Source:

https://www.lesswrong.com/posts/ZoFxTqWRBkyanonyb/current-safety-training-techniques-do-not-fully-transfer-to

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more

View all episodes

View all episodes

Download on the App Store

Download on the App Store

Get it on Google Play

LessWrong (30+ Karma)

By LessWrong

November 03, 2024

“Current safety training techniques do not fully transfer to the agent setting” by Simon Lermen, Govind Pimpale

Listen Later

10 minutes

TL;DR: I'm presenting three recent papers which all share a similar finding, i.e. the safety training techniques for chat models don’t transfer well from chat models to the agents built from them. In other words, models won’t tell you how to do something harmful, but they are often willing to directly execute harmful actions. However, all papers find that different attack methods like jailbreaks, prompt-engineering, or refusal-vector ablation do transfer.

Here are the three papers:

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
Applying Refusal-Vector Ablation to Llama 3.1 70B Agents

What are language model agents

Language model agents are a combination of a language model and a scaffolding software. Regular language models are typically limited to being chat bots, i.e. they receive messages and reply to them. However, scaffolding gives these models access to tools which they can [...]

---

Outline:

(00:55) What are language model agents

(01:36) Overview

(03:31) AgentHarm Benchmark

(05:27) Refusal-Trained LLMs Are Easily Jailbroken as Browser Agents

(06:47) Applying Refusal-Vector Ablation to Llama 3.1 70B Agents

(08:23) Discussion

---

First published:

November 3rd, 2024

Source:

https://www.lesswrong.com/posts/ZoFxTqWRBkyanonyb/current-safety-training-techniques-do-not-fully-transfer-to

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more

More shows like LessWrong (30+ Karma)

The Daily by The New York Times

The Daily

112,216 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

131 Listeners

Interesting Times with Ross Douthat by New York Times Opinion

Interesting Times with Ross Douthat

7,229 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

568 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

16,205 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates! by Liron Shapira

Doom Debates!

14 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners