June 07, 2024

“Scaling and evaluating sparse autoencoders” by leogao

1 minute

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.

[Blog] [Paper] [Visualizer]

Abstract:

Sparse autoencoders provide a promising unsupervised approach for extracting interpretable features from a language model by reconstructing activations from a sparse bottleneck layer. Since language models learn many concepts, autoencoders need to be very large to recover all relevant features. However, studying the properties of autoencoder scaling is difficult due to the need to balance reconstruction and sparsity objectives and the presence of dead latents. We propose using k-sparse autoencoders [Makhzani and Frey, 2013] to directly control sparsity, simplifying tuning and improving the reconstruction-sparsity frontier. Additionally, we find modifications that result in few dead latents, even at the largest scales we tried. Using these techniques, we find clean scaling laws with respect to autoencoder size and sparsity. We also introduce several new metrics for evaluating feature quality based on the recovery of hypothesized [...]

---

First published:

June 6th, 2024

Source:

https://www.lesswrong.com/posts/Fg2gAgxN6hHSaTjkf/scaling-and-evaluating-sparse-autoencoders

---

Narrated by TYPE III AUDIO.

...more