LessWrong (Curated & Popular)

“Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research (GDM Mech Interp Team Progress Update #2)” by Neel Nanda, lewis smith, Senthooran Rajamanoharan, Arthur Conmy, Callum McDougall, Tom Lieberum, János Kramár, Rohin Shah


Listen Later

Audio note: this article contains 31 uses of latex notation, so the narration may be difficult to follow. There's a link to the original text in the episode description.

Lewis Smith*, Sen Rajamanoharan*, Arthur Conmy, Callum McDougall, Janos Kramar, Tom Lieberum, Rohin Shah, Neel Nanda

* = equal contribution

The following piece is a list of snippets about research from the GDM mechanistic interpretability team, which we didn’t consider a good fit for turning into a paper, but which we thought the community might benefit from seeing in this less formal form. These are largely things that we found in the process of a project investigating whether sparse autoencoders were useful for downstream tasks, notably out-of-distribution probing.

TL;DR
  • To validate whether SAEs were a worthwhile technique, we explored whether they were useful on the downstream task of OOD generalisation when detecting harmful intent in user prompts
  • [...]
---

Outline:

(01:08) TL;DR

(02:38) Introduction

(02:41) Motivation

(06:09) Our Task

(08:35) Conclusions and Strategic Updates

(13:59) Comparing different ways to train Chat SAEs

(18:30) Using SAEs for OOD Probing

(20:21) Technical Setup

(20:24) Datasets

(24:16) Probing

(26:48) Results

(30:36) Related Work and Discussion

(34:01) Is it surprising that SAEs didn't work?

(39:54) Dataset debugging with SAEs

(42:02) Autointerp and high frequency latents

(44:16) Removing High Frequency Latents from JumpReLU SAEs

(45:04) Method

(45:07) Motivation

(47:29) Modifying the sparsity penalty

(48:48) How we evaluated interpretability

(50:36) Results

(51:18) Reconstruction loss at fixed sparsity

(52:10) Frequency histograms

(52:52) Latent interpretability

(54:23) Conclusions

(56:43) Appendix

The original text contained 7 footnotes which were omitted from this narration.

---

First published:
March 26th, 2025

Source:
https://www.lesswrong.com/posts/4uXCAJNuPKtKBsi28/sae-progress-update-2-draft

---

Narrated by TYPE III AUDIO.

---

Images from the article:

...more
View all episodesView all episodes
Download on the App Store

LessWrong (Curated & Popular)By LessWrong

  • 4.8
  • 4.8
  • 4.8
  • 4.8
  • 4.8

4.8

11 ratings


More shows like LessWrong (Curated & Popular)

View all
Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,388 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

123 Listeners

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas by Sean Carroll | Wondery

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

4,133 Listeners

ManifoldOne by Steve Hsu

ManifoldOne

87 Listeners

The Jim Rutt Show by The Jim Rutt Show

The Jim Rutt Show

251 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

87 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

389 Listeners

Hard Fork by The New York Times

Hard Fork

5,432 Listeners

Clearer Thinking with Spencer Greenberg by Spencer Greenberg

Clearer Thinking with Spencer Greenberg

128 Listeners

Razib Khan's Unsupervised Learning by Razib Khan

Razib Khan's Unsupervised Learning

198 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

121 Listeners

Latent Space: The AI Engineer Podcast by swyx + Alessio

Latent Space: The AI Engineer Podcast

75 Listeners

"Econ 102" with Noah Smith and Erik Torenberg by Turpentine

"Econ 102" with Noah Smith and Erik Torenberg

145 Listeners

Complex Systems with Patrick McKenzie (patio11) by Patrick McKenzie

Complex Systems with Patrick McKenzie (patio11)

121 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

1 Listeners