October 14, 2024

“The case for unlearning that removes information from LLM weights” by Fabien Roger

11 minutes

What if you could remove some information from the weights of an AI? Would that be helpful?

It is clearly useful against some misuse concerns: if you are concerned that LLMs will make it easier to build bioweapons because they have memorized such information, removing the memorized facts would remove this misuse concern.

In a paper Aghyad Deeb and I just released, we show it is tractable to evaluate the presence of certain undesirable facts in an LLM: take independent facts that should have all been removed, fine-tune on some of them, and see if accuracy increases on the other ones. The fine-tuning process should make the model “try” to answer, but if the information was removed from the weights (and if the facts are actually independent), then accuracy on the held-out facts should remain low.

Removing information from the weights is stronger than the usual notion of [...]

---

Outline:

(01:50) Do current unlearning techniques remove facts from model weights?

(04:24) Hopes for successful information removal

(06:51) Using information removal to reduce x-risk

(06:56) Information you should probably remove from the weights

(08:20) How removing information helps you

(09:20) Information you probably can’t remove - and why this won’t work for superintelligent AIs

The original text contained 5 footnotes which were omitted from this narration.

The original text contained 2 images which were described by AI.

---

First published:

October 14th, 2024

Source:

https://www.lesswrong.com/posts/9AbYkAy8s9LvB7dT5/the-case-for-unlearning-that-removes-information-from-llm

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more

View all episodes

By LessWrong

October 14, 2024

“The case for unlearning that removes information from LLM weights” by Fabien Roger

11 minutes

What if you could remove some information from the weights of an AI? Would that be helpful?

Removing information from the weights is stronger than the usual notion of [...]

---

Outline:

(01:50) Do current unlearning techniques remove facts from model weights?

(04:24) Hopes for successful information removal

(06:51) Using information removal to reduce x-risk

(06:56) Information you should probably remove from the weights

(08:20) How removing information helps you

(09:20) Information you probably can’t remove - and why this won’t work for superintelligent AIs

The original text contained 5 footnotes which were omitted from this narration.

The original text contained 2 images which were described by AI.

---

First published:

October 14th, 2024

Source:

https://www.lesswrong.com/posts/9AbYkAy8s9LvB7dT5/the-case-for-unlearning-that-removes-information-from-llm

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more

More shows like LessWrong (30+ Karma)

View all

Making Sense with Sam Harris

26,366 Listeners

Conversations with Tyler

2,384 Listeners

The Peter Attia Drive

7,944 Listeners

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

4,137 Listeners

ManifoldOne

87 Listeners

Your Undivided Attention

1,459 Listeners

All-In with Chamath, Jason, Sacks & Friedberg

9,050 Listeners

Machine Learning Street Talk (MLST)

88 Listeners

Dwarkesh Podcast

386 Listeners

Hard Fork

5,422 Listeners

The Ezra Klein Show

15,228 Listeners

Moonshots with Peter Diamandis

473 Listeners

No Priors: Artificial Intelligence | Technology | Startups

120 Listeners

Latent Space: The AI Engineer Podcast

76 Listeners

BG2Pod with Brad Gerstner and Bill Gurley

456 Listeners

Share “The case for unlearning that removes information from LLM weights” by Fabien Roger

Sign up to save your podcasts

“The case for unlearning that removes information from LLM weights” by Fabien Roger

“The case for unlearning that removes information from LLM weights” by Fabien Roger

More shows like LessWrong (30+ Karma)

Making Sense with Sam Harris

Conversations with Tyler

The Peter Attia Drive

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

ManifoldOne

Your Undivided Attention

All-In with Chamath, Jason, Sacks & Friedberg

Machine Learning Street Talk (MLST)

Dwarkesh Podcast

Hard Fork

The Ezra Klein Show

Moonshots with Peter Diamandis

No Priors: Artificial Intelligence | Technology | Startups

Latent Space: The AI Engineer Podcast

BG2Pod with Brad Gerstner and Bill Gurley