December 14, 2024

“Matryoshka Sparse Autoencoders” by Noa Nabeshima

Listen Later

21 minutes

View trees here
Search through latents with a token-regex language
View individual latents here
See code here (github.com/noanabeshima/matryoshka-saes)
Continually updated version of this document

Abstract

Sparse autoencoders (SAEs)[1][2] break down neural network internals into components called latents. Smaller SAE latents seem to correspond to more abstract concepts while larger SAE latents seem to represent finer, more specific concepts.

While increasing SAE size allows for finer-grained representations, it also introduces two key problems: feature absorption introduced in Chanin et al. [3], where latents develop unintuitive "holes" as other latents in the SAE take over specific cases, and what I term fragmentation, where meaningful abstract concepts in the small SAE (e.g. 'female names' or 'words in quotes') shatter (via feature splitting[1:1]) into many specific latents, hiding real structure in the model.

This paper introduces Matryoshka SAEs, a training approach that addresses these challenges. Inspired by prior work[4][5], Matryoshka SAEs are trained [...]

---

Outline:

(00:18) Abstract

(01:40) Introduction

(04:08) Problem

(04:11) Terminology

(04:34) Reference SAEs

(05:27) Feature Absorption Example

(08:26) Method

(11:07) Results

(11:10) Toy Model

(15:58) Reconstruction Quality

(17:14) Limitations and Future Work

(20:30) Acknowledgements

The original text contained 30 footnotes which were omitted from this narration.

The original text contained 8 images which were described by AI.

---

First published:

December 14th, 2024

Source:

https://www.lesswrong.com/posts/zbebxYCqsryPALh8C/matryoshka-sparse-autoencoders

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more

View all episodes

View all episodes

Download on the App Store

Download on the App Store

Get it on Google Play

LessWrong (30+ Karma)

By LessWrong

December 14, 2024

“Matryoshka Sparse Autoencoders” by Noa Nabeshima

Listen Later

21 minutes

View trees here
Search through latents with a token-regex language
View individual latents here
See code here (github.com/noanabeshima/matryoshka-saes)
Continually updated version of this document

Abstract

Sparse autoencoders (SAEs)[1][2] break down neural network internals into components called latents. Smaller SAE latents seem to correspond to more abstract concepts while larger SAE latents seem to represent finer, more specific concepts.

While increasing SAE size allows for finer-grained representations, it also introduces two key problems: feature absorption introduced in Chanin et al. [3], where latents develop unintuitive "holes" as other latents in the SAE take over specific cases, and what I term fragmentation, where meaningful abstract concepts in the small SAE (e.g. 'female names' or 'words in quotes') shatter (via feature splitting[1:1]) into many specific latents, hiding real structure in the model.

This paper introduces Matryoshka SAEs, a training approach that addresses these challenges. Inspired by prior work[4][5], Matryoshka SAEs are trained [...]

---

Outline:

(00:18) Abstract

(01:40) Introduction

(04:08) Problem

(04:11) Terminology

(04:34) Reference SAEs

(05:27) Feature Absorption Example

(08:26) Method

(11:07) Results

(11:10) Toy Model

(15:58) Reconstruction Quality

(17:14) Limitations and Future Work

(20:30) Acknowledgements

The original text contained 30 footnotes which were omitted from this narration.

The original text contained 8 images which were described by AI.

---

First published:

December 14th, 2024

Source:

https://www.lesswrong.com/posts/zbebxYCqsryPALh8C/matryoshka-sparse-autoencoders

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more

More shows like LessWrong (30+ Karma)

Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,346 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,384 Listeners

The Peter Attia Drive by Peter Attia, MD

The Peter Attia Drive

7,976 Listeners

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas by Sean Carroll | Wondery

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

4,133 Listeners

ManifoldOne by Steve Hsu

ManifoldOne

87 Listeners

Your Undivided Attention by Tristan Harris and Aza Raskin, The Center for Humane Technology

Your Undivided Attention

1,444 Listeners

All-In with Chamath, Jason, Sacks & Friedberg by All-In Podcast, LLC

All-In with Chamath, Jason, Sacks & Friedberg

8,905 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

87 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

373 Listeners

Hard Fork by The New York Times

Hard Fork

5,414 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,293 Listeners

Moonshots with Peter Diamandis by PHD Ventures

Moonshots with Peter Diamandis

468 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

122 Listeners

Latent Space: The AI Engineer Podcast by swyx + Alessio

Latent Space: The AI Engineer Podcast

77 Listeners

BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

BG2Pod with Brad Gerstner and Bill Gurley

452 Listeners