May 07, 2025

“Negative Results on Group SAEs” by Josh Engels

Listen Later

18 minutes

Introduction

Soon after we released Not All Language Model Features Are One-Dimensionally Linear, I started working with @Logan Riggs and @Jannik Brinkmann on a natural followup to the paper: could we build a variant of SAEs that could find multi-dimensional features directly, instead of needing to cluster SAE latents post-hoc like we did in the paper.

We worked on this for a few months last summer and tried a bunch of things. Unfortunately, none of our results were that compelling, and eventually our interest in the project died down and we didn’t publish our (mostly negative) results. Recently, multiple people (@Noa Nabeshima , @chanind, Goncalo Paulo) said they were interested in working on SAEs that could find multi-dimensional features, so I decided I would write up what we tried.

At this point the results are almost a year old, but I think the overall narrative should still [...]

---

Outline:

(00:10) Introduction

(02:32) Group SAEs

(03:23) Synthetic Circles Experiments

(07:15) Training Group SAEs on GPT-2

(07:27) High level metrics

(09:28) Do the Group SAEs Capture Known Circular Subspaces

(11:46) Other Things We Tried

(12:03) Experimenting with learned groups

(12:08) Motivation and Ideas

(15:43) Learned Group Space

(18:13) Conclusion

---

First published:

May 6th, 2025

Source:

https://www.lesswrong.com/posts/jKKbRKuXNaLujnojw/untitled-draft-okbt

---

Narrated by TYPE III AUDIO.

---

Images from the article:

" showing colored points from 1-12" style="max-width: 100%;" />

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more

View all episodes

View all episodes

Download on the App Store

Download on the App Store

Get it on Google Play

LessWrong (30+ Karma)

By LessWrong

May 07, 2025

“Negative Results on Group SAEs” by Josh Engels

Listen Later

18 minutes

Introduction

Soon after we released Not All Language Model Features Are One-Dimensionally Linear, I started working with @Logan Riggs and @Jannik Brinkmann on a natural followup to the paper: could we build a variant of SAEs that could find multi-dimensional features directly, instead of needing to cluster SAE latents post-hoc like we did in the paper.

We worked on this for a few months last summer and tried a bunch of things. Unfortunately, none of our results were that compelling, and eventually our interest in the project died down and we didn’t publish our (mostly negative) results. Recently, multiple people (@Noa Nabeshima , @chanind, Goncalo Paulo) said they were interested in working on SAEs that could find multi-dimensional features, so I decided I would write up what we tried.

At this point the results are almost a year old, but I think the overall narrative should still [...]

---

Outline:

(00:10) Introduction

(02:32) Group SAEs

(03:23) Synthetic Circles Experiments

(07:15) Training Group SAEs on GPT-2

(07:27) High level metrics

(09:28) Do the Group SAEs Capture Known Circular Subspaces

(11:46) Other Things We Tried

(12:03) Experimenting with learned groups

(12:08) Motivation and Ideas

(15:43) Learned Group Space

(18:13) Conclusion

---

First published:

May 6th, 2025

Source:

https://www.lesswrong.com/posts/jKKbRKuXNaLujnojw/untitled-draft-okbt

---

Narrated by TYPE III AUDIO.

---

Images from the article:

" showing colored points from 1-12" style="max-width: 100%;" />

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more

More shows like LessWrong (30+ Karma)

Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,469 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,395 Listeners

The Peter Attia Drive by Peter Attia, MD

The Peter Attia Drive

7,967 Listeners

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas by Sean Carroll | Wondery

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

4,145 Listeners

ManifoldOne by Steve Hsu

ManifoldOne

90 Listeners

Your Undivided Attention by Tristan Harris and Aza Raskin, The Center for Humane Technology

Your Undivided Attention

1,480 Listeners

All-In with Chamath, Jason, Sacks & Friedberg by All-In Podcast, LLC

All-In with Chamath, Jason, Sacks & Friedberg

9,236 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

88 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

428 Listeners

Hard Fork by The New York Times

Hard Fork

5,462 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,335 Listeners

Moonshots with Peter Diamandis by PHD Ventures

Moonshots with Peter Diamandis

483 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

121 Listeners

Latent Space: The AI Engineer Podcast by swyx + Alessio

Latent Space: The AI Engineer Podcast

75 Listeners

BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

BG2Pod with Brad Gerstner and Bill Gurley

469 Listeners