LessWrong posts by zvi

“AI #23: Fundamental Problems with RLHF” by Zvi


Listen Later

After several jam-packed weeks, things slowed down to allow everyone to focus on the potential room temperature superconductor, check Polymarket to see how likely it is we are so back and bet real money, or Manifold for chats and better graphs and easier but much smaller trading.

The main thing I would highlight this week are an excellent paper laying out many of the fundamental difficulties with RLHF, and a systematic new exploit of current LLMs that seems to reliably defeat RLHF.

I’d also note that GPT-4 fine tuning is confirmed to be coming. That should be fun.

Table of Contents

  1. Introduction.
  2. Table of Contents.
  3. Language Models Offer Mundane Utility. Here’s what you’re going to do.
  4. Language Models Don’t Offer Mundane Utility. Universal attacks on LLMs.
  5. Fun With Image Generation. Videos might be a while.
  6. Deepfaketown and Botpocalypse Soon. An [...]

    ---

    Outline:

    (00:38) Table of Contents

    (02:32) Language Models Offer Mundane Utility

    (05:20) Language Models Don’t Offer Mundane Utility

    (16:57) Fun with Image Generation

    (17:40) Deepfaketown and Botpocalypse Soon

    (17:49) They Took Our Jobs

    (22:31) Get Involved

    (23:35) Introducing

    (31:33) In Other AI News

    (34:46) Quiet Speculations

    (36:13) China

    (38:31) The Quest for Sane Regulations

    (45:04) The Week in Audio

    (45:37) Rhetorical Innovation

    (47:51) No One Would Be So Stupid As To

    (49:44) Aligning a Smarter Than Human Intelligence is Difficult

    (01:10:30) Other People Are Not As Worried About AI Killing Everyone

    (01:11:34) The Wit and Wisdom of Sam Altman

    (01:11:56) The Lighter Side

    ---

  7. First published:

    August 3rd, 2023

    Source:

    https://www.lesswrong.com/posts/aKzwwKT2cy72awSyz/ai-23-fundamental-problems-with-rlhf

    ---

    Narrated by TYPE III AUDIO.

    ...more
    View all episodesView all episodes
    Download on the App Store

    LessWrong posts by zviBy zvi

    • 5
    • 5
    • 5
    • 5
    • 5

    5

    2 ratings


    More shows like LessWrong posts by zvi

    View all
    Making Sense with Sam Harris by Sam Harris

    Making Sense with Sam Harris

    26,391 Listeners

    Conversations with Tyler by Mercatus Center at George Mason University

    Conversations with Tyler

    2,463 Listeners

    The a16z Show by Andreessen Horowitz

    The a16z Show

    1,101 Listeners

    Future of Life Institute Podcast by Future of Life Institute

    Future of Life Institute Podcast

    109 Listeners

    ChinaTalk by Jordan Schneider

    ChinaTalk

    295 Listeners

    Politix by Politix

    Politix

    89 Listeners

    Dwarkesh Podcast by Dwarkesh Patel

    Dwarkesh Podcast

    555 Listeners

    Hard Fork by The New York Times

    Hard Fork

    5,553 Listeners

    Clearer Thinking with Spencer Greenberg by Spencer Greenberg

    Clearer Thinking with Spencer Greenberg

    139 Listeners

    LessWrong (Curated & Popular) by LessWrong

    LessWrong (Curated & Popular)

    14 Listeners

    No Priors: Artificial Intelligence | Technology | Startups by Conviction

    No Priors: Artificial Intelligence | Technology | Startups

    141 Listeners

    "Econ 102" with Noah Smith and Erik Torenberg by Turpentine

    "Econ 102" with Noah Smith and Erik Torenberg

    155 Listeners

    BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

    BG2Pod with Brad Gerstner and Bill Gurley

    458 Listeners

    LessWrong (30+ Karma) by LessWrong

    LessWrong (30+ Karma)

    0 Listeners

    Complex Systems with Patrick McKenzie (patio11) by Patrick McKenzie

    Complex Systems with Patrick McKenzie (patio11)

    143 Listeners