June 03, 2024

“Comments on Anthropic’s Scaling Monosemanticity” by Robert_AIZI

Listen Later

11 minutes

These are some of my notes from reading Anthropic's latest research report, Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet.

TL;DR

In roughly descending order of importance:

Its great that Anthropic trained an SAE on a production-scale language model, and that the approach works to find interpretable features. Its great those features allow interventions like the recently-departed Golden Gate Claude. I especially like the code bug feature.
I worry that naming features after high-activating examples (e.g. "the Golden Gate Bridge feature") gives a false sense of security. Most of the time that feature activates, it is irrelevant to the golden gate bridge. That feature is only well-described as "related to the golden gate bridge" if you condition on a very high activation, and that's <10% of its activations (from an eyeballing of the graph).
This work does not address my major concern about dictionary learning: it is [...]

---

Outline:

(00:16) TL;DR

(02:23) A Feature Isnt Its Highest Activating Examples

(04:38) Finding Specific Features

(06:02) Architecture - The Classics, but Wider

(07:26) Correlations - Strangely Large?

(09:48) Future Tests

The original text contained 2 footnotes which were omitted from this narration.

---

First published:

June 3rd, 2024

Source:

https://www.lesswrong.com/posts/zzmhsKx5dBpChKhry/comments-on-anthropic-s-scaling-monosemanticity

---

Narrated by TYPE III AUDIO.

...more

View all episodes

View all episodes

Download on the App Store

Download on the App Store

Get it on Google Play

LessWrong (30+ Karma)

By LessWrong

June 03, 2024

“Comments on Anthropic’s Scaling Monosemanticity” by Robert_AIZI

Listen Later

11 minutes

These are some of my notes from reading Anthropic's latest research report, Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet.

TL;DR

In roughly descending order of importance:

Its great that Anthropic trained an SAE on a production-scale language model, and that the approach works to find interpretable features. Its great those features allow interventions like the recently-departed Golden Gate Claude. I especially like the code bug feature.
I worry that naming features after high-activating examples (e.g. "the Golden Gate Bridge feature") gives a false sense of security. Most of the time that feature activates, it is irrelevant to the golden gate bridge. That feature is only well-described as "related to the golden gate bridge" if you condition on a very high activation, and that's <10% of its activations (from an eyeballing of the graph).
This work does not address my major concern about dictionary learning: it is [...]

---

Outline:

(00:16) TL;DR

(02:23) A Feature Isnt Its Highest Activating Examples

(04:38) Finding Specific Features

(06:02) Architecture - The Classics, but Wider

(07:26) Correlations - Strangely Large?

(09:48) Future Tests

The original text contained 2 footnotes which were omitted from this narration.

---

First published:

June 3rd, 2024

Source:

https://www.lesswrong.com/posts/zzmhsKx5dBpChKhry/comments-on-anthropic-s-scaling-monosemanticity

---

Narrated by TYPE III AUDIO.

...more

More shows like LessWrong (30+ Karma)

The Daily by The New York Times

The Daily

112,144 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

131 Listeners

Interesting Times with Ross Douthat by New York Times Opinion

Interesting Times with Ross Douthat

7,238 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

577 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

16,139 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates! by Liron Shapira

Doom Debates!

14 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners