LessWrong (30+ Karma)

“The Most Forbidden Technique” by Zvi


Listen Later

The Most Forbidden Technique is training an AI using interpretability techniques.

An AI produces a final output [X] via some method [M]. You can analyze [M] using technique [T], to learn what the AI is up to. You could train on that. Never do that.

You train on [X]. Only [X]. Never [M], never [T].

Why? Because [T] is how you figure out when the model is misbehaving.

If you train on [T], you are training the AI to obfuscate its thinking, and defeat [T]. You will rapidly lose your ability to know what is going on, in exactly the ways you most need to know what is going on.

Those bits of optimization pressure from [T] are precious. Use them wisely.

Table of Contents

  1. New Paper Warns Against the Most Forbidden Technique.
  2. Reward Hacking Is The Default.
  3. Using [...]
  4. ---

    Outline:

    (00:57) New Paper Warns Against the Most Forbidden Technique

    (06:52) Reward Hacking Is The Default

    (09:25) Using CoT to Detect Reward Hacking Is Most Forbidden Technique

    (11:49) Not Using the Most Forbidden Technique Is Harder Than It Looks

    (14:10) It's You, It's Also the Incentives

    (17:41) The Most Forbidden Technique Quickly Backfires

    (18:58) Focus Only On What Matters

    (19:33) Is There a Better Way?

    (21:34) What Might We Do Next?

    ---

    First published:

    March 12th, 2025

    Source:

    https://www.lesswrong.com/posts/mpmsK8KKysgSKDm2T/the-most-forbidden-technique

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    ...more
    View all episodesView all episodes
    Download on the App Store

    LessWrong (30+ Karma)By LessWrong


    More shows like LessWrong (30+ Karma)

    View all
    The Daily by The New York Times

    The Daily

    112,664 Listeners

    Astral Codex Ten Podcast by Jeremiah

    Astral Codex Ten Podcast

    130 Listeners

    Interesting Times with Ross Douthat by New York Times Opinion

    Interesting Times with Ross Douthat

    7,216 Listeners

    Dwarkesh Podcast by Dwarkesh Patel

    Dwarkesh Podcast

    530 Listeners

    The Ezra Klein Show by New York Times Opinion

    The Ezra Klein Show

    16,132 Listeners

    AI Article Readings by Readings of great articles in AI voices

    AI Article Readings

    4 Listeners

    Doom Debates by Liron Shapira

    Doom Debates

    14 Listeners

    LessWrong posts by zvi by zvi

    LessWrong posts by zvi

    2 Listeners