LessWrong posts by zvi

LessWrong posts by zvi

Download on the App Store

LessWrong posts by zvi episodes

  • “On Dwarkesh Patel’s Second Interview With Ilya Sutskever” by Zvi

    Some podcasts are self-recommending on the ‘yep, I’m going to be breaking this one down’ level. This was very clearly one of those. So here we go.

    Double click to interact with video

    As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary.

    If I am quoting directly I use quote marks, otherwise assume paraphrases.

    What are the main takeaways?

  • Ilya thinks training in its current form will peter out, that we are returning to an age of research where progress requires more substantially new ideas.
  • SSI is a research organization. It tries various things. Not having a product lets it punch well above its fundraising weight in compute and effective resources.
  •  
  • Ilya has 5-20 year timelines to a potentially superintelligent learning model.
  • SSI might release a product first after all, but probably not?
  • Ilya's thinking about alignment still seems relatively shallow to me in key ways, but he grasps many important insights and understands he has a problem.
  • Ilya essentially despairs of having a substantive plan beyond ‘show everyone the thing as early [...]
  • ---

    Outline:

    (01:42) Explaining Model Jaggedness

    (03:15) Emotions and value functions

    (04:38) What are we scaling?

    (05:47) Why humans generalize better than models

    (07:00) Straight-shooting superintelligence

    (08:39) SSI's model will learn from deployment

    (09:35) Alignment

    (17:40) We are squarely an age of research company

    (22:27) Research taste

    (25:11) Bonus Coverage: Dwarkesh Patel on AI Progress These Days

    ---

    First published:

    December 3rd, 2025

    Source:

    https://www.lesswrong.com/posts/bMvCNtSH8DiGDTvXd/on-dwarkesh-patel-s-second-interview-with-ilya-sutskever

    ---

    Narrated by TYPE III AUDIO.

    40 min
  • “Reward Mismatches in RL Cause Emergent Misalignment” by Zvi

    Learning to do misaligned-coded things anywhere teaches an AI (or a human) to do misaligned-coded things everywhere. So be sure you never, ever teach any mind to do what it sees, in context, as misaligned-coded things.

    If the optimal solution (as in, the one you most reinforce) to an RL training problem is one that the model perceives as something you wouldn’t want it to do, it will generally learn to do things you don’t want it to do.

    You can solve this by ensuring that the misaligned-coded things are not what the AI will learn to do. Or you can solve this by making those things not misaligned-coded.

    If you then teaching aligned behavior in one set of spots, this can fix the problem in those spots, but the fix does not generalize to other tasks or outside of distribution. If you manage to hit the entire distribution of tasks you care about in this way, that will work for now, but it still won’t generalize, so it's a terrible long term strategy.

    Yo Shavit: Extremely important finding.

    Don’t tell your model you’re rewarding it for A and then reward it for B [...]

    ---

    Outline:

    (02:59) Abstract Of The Paper

    (04:12) The Problem Statement

    (05:35) The Inoculation Solution

    (07:02) Cleaning The Data Versus Cleaning The Environments

    (08:16) No All Of This Does Not Solve Our Most Important Problems

    (13:18) It Does Help On Important Short Term Problems

    ---

    First published:

    December 2nd, 2025

    Source:

    https://www.lesswrong.com/posts/a2nW8buG2Lw9AdPtH/reward-mismatches-in-rl-cause-emergent-misalignment

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    15 min
  • “Claude Opus 4.5 Is The Best Model Available” by Zvi

    Claude Opus 4.5 is the best model currently available.

    No model since GPT-4 has come close to the level of universal praise that I have seen for Claude Opus 4.5.

    It is the most intelligent and capable, most aligned and thoughtful model. It is a joy.

    There are some auxiliary deficits, and areas where other models have specialized, and even with the price cut Opus remains expensive, so it should not be your exclusive model. I do think it should absolutely be your daily driver.

    Image by Nana Banana Pro, prompt chosen for this purpose by Claude Opus 4.5

    Table of Contents

  • It's The Best Model, Sir.
  • Huh, Upgrades.
  • On Your Marks.
  • Anthropic Gives Us Very Particular Hype.
  • Employee Hype.
  • Every Vibe Check.
  • Spontaneous Positive Reactions.
  • Reaction Thread Positive Reactions.
  • Negative Reactions.
  • The Lighter Side.
  • Popularity.
  • You’ve Got Soul.
  • It's The Best Model, Sir

    Here is the full picture of where we are now (as mostly seen in Friday's post):

    You want to be using Claude Opus 4.5.

    That is especially true for coding, or if [...]

    ---

    Outline:

    (00:59) It's The Best Model, Sir

    (03:18) Huh, Upgrades

    (04:50) On Your Marks

    (09:12) Anthropic Gives Us Very Particular Hype

    (13:35) Employee Hype

    (15:40) Every Vibe Check

    (18:16) Spontaneous Positive Reactions

    (21:44) Reaction Thread Positive Reactions

    (28:39) Negative Reactions

    (30:34) The Lighter Side

    (31:27) Popularity

    (33:26) You've Got Soul

    ---

    First published:

    December 1st, 2025

    Source:

    https://www.lesswrong.com/posts/HtdrtF5kcpLtWe5dW/claude-opus-4-5-is-the-best-model-available

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    45 min
  • Claude Opus 4.5: Model Card, Alignment and Safety

    They saved the best for last.

    The contrast in model cards is stark. Google provided a brief overview of its tests for Gemini 3 Pro, with a lot of ‘we did this test, and we learned a lot from it, and we are not going to tell you the results.’

    Anthropic gives us a 150 page book, including their capability assessments. This makes sense. Capability is directly relevant to safety, and also frontier capability safety tests often also credible indications of capability.

    Which still has several instances of ‘we did this test, and we learned a lot from it, and we are not going to tell you the results.’ Damn it. I get it, but damn it.

    Anthropic claims Opus 4.5 is the most aligned frontier model to date, although ‘with many subtleties.’

    I agree with Anthropic's assessment, especially for practical purposes right now.

    Claude is also miles ahead of other models on aspects of alignment that do not directly appear on a frontier safety assessment.

    In terms of surviving superintelligence, it's still the scene from The Phantom Menace. As in, that won’t be enough.

    (Above: Claude Opus 4.5 self-portrait as [...]

    ---

    Outline:

    (01:37) Claude Opus 4.5 Basic Facts

    (03:12) Claude Opus 4.5 Is The Best Model For Many But Not All Use Cases

    (05:38) Misaligned?

    (09:04) Section 3: Safeguards and Harmlessness

    (11:15) Section 4: Honesty

    (12:33) 5: Agentic Safety

    (17:09) Section 6: Alignment Overview

    (23:45) Alignment Investigations

    (24:23) Sycophancy Course Correction Is Lacking

    (25:37) Deception

    (28:05) Ruling Out Encoded Content In Chain Of Thought

    (30:16) Sandbagging

    (31:05) Evaluation Awareness

    (35:05) Reward Hacking

    (36:24) Subversion Strategy

    (37:19) 6.13: UK AISI External Testing

    (37:31) 6.14: Model Welfare

    (38:22) 7: RSP Evaluations

    (40:01) CBRN

    (47:34) Autonomy

    (54:50) Cyber

    (58:29) The Whisperers Love The Vibes

    ---

    First published:

    November 28th, 2025

    Source:

    https://www.lesswrong.com/posts/gfby4vqNtLbehqbot/claude-opus-4-5-model-card-alignment-and-safety

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 1 min
  • AI #144: Thanks For the Models

    Thanks for everything. And I do mean everything.

    Everyone gave us a new model in the last few weeks.

    OpenAI gave us GPT-5.1 and GPT-5.1-Codex-Max. These are overall improvements, although there are worries around glazing and reintroducing parts of the 4o spirit.

    xAI gave us Grok 4.1, although few seem to have noticed and I haven’t tried it.

    Google gave us both by far the best image model in Nana Banana Pro and also Gemini 3 Pro, which is a vast intelligence with no spine. It is extremely intelligent and powerful, but comes with severe issues. My assessment of it as the new state of the art got to last all of about five hours.

    Anthropic gave us Claude Opus 4.5. This is probably the best model and quickly became my daily driver for most but not all purposes including coding. I plan to do full coverage in two parts, with alignment and safety on Friday, and the full capabilities report and general review on Monday.

    Meanwhile the White House is announcing the Genesis Mission to accelerate science, there's a continuing battle over another attempt at a moratorium, there's a new planned $50 [...]

    ---

    Outline:

    (02:20) Language Models Offer Mundane Utility

    (02:53) Language Models Don't Offer Mundane Utility

    (03:22) Huh, Upgrades

    (05:52) On Your Marks

    (07:30) Choose Your Fighter

    (08:11) Deepfaketown and Botpocalypse Soon

    (14:40) What Is Slop? How Do You Define Slop?

    (17:43) Fun With Media Generation

    (21:40) A Young Lady's Illustrated Primer

    (23:58) You Drive Me Crazy

    (28:31) They Took Our Jobs

    (28:53) Think Of The Time I Saved

    (32:07) The Art of the Jailbreak

    (33:02) Get Involved

    (33:37) Introducing

    (34:18) In Other AI News

    (37:05) Show Me the Money

    (39:11) Quiet Speculations

    (41:55) Bubble, Bubble, Toil and Trouble

    (44:32) The Quest for Sane Regulations

    (54:31) Chip City

    (55:50) Water Water Everywhere

    (57:17) The Week in Audio

    (59:05) Rhetorical Innovation

    (01:04:53) You Are Not In Control

    (01:08:42) AI 2030

    (01:19:19) Aligning a Smarter Than Human Intelligence is Difficult

    (01:21:57) Misaligned?

    (01:25:00) Messages From Janusworld

    (01:27:16) The Lighter Side

    The original text contained 1 footnote which was omitted from this narration.

    ---

    First published:

    November 27th, 2025

    Source:

    https://www.lesswrong.com/posts/o7gQJyGeeAGKK6bRx/ai-144-thanks-for-the-models

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 29 min
  • The Big Nonprofits Post 2025

    There remain lots of great charitable giving opportunities out there.

    I have now had three opportunities to be a recommender for the Survival and Flourishing Fund (SFF). I wrote in detail about my first experience back in 2021, where I struggled to find worthy applications.

    The second time around in 2024, there was an abundance of worthy causes. In 2025 there were even more high quality applications, many of which were growing beyond our ability to support them.

    Thus this is the second edition of The Big Nonprofits Post, primarily aimed at sharing my findings on various organizations I believe are doing good work, to help you find places to consider donating in the cause areas and intervention methods that you think are most effective, and to offer my general perspective on how I think about choosing where to give.

    This post combines my findings from the 2024 and 2025 rounds of SFF, and also includes some organizations that did not apply to either round, so inclusion does not mean that they necessarily applied at all.

    This post is already very long, so the bar is higher for inclusion this year than it was [...]

    ---

    Outline:

    (01:40) A Word of Warning

    (02:50) A Note To Charities

    (03:53) Use Your Personal Theory of Impact

    (05:40) Use Your Local Knowledge

    (06:41) Unconditional Grants to Worthy Individuals Are Great

    (09:00) Do Not Think Only On the Margin, and Also Use Decision Theory

    (10:03) Compare Notes With Those Individuals You Trust

    (10:35) Beware Becoming a Fundraising Target

    (11:02) And the Nominees Are

    (14:34) Organizations that Are Literally Me

    (14:49) Balsa Research

    (17:31) Don't Worry About the Vase

    (19:04) Organizations Focusing On AI Non-Technical Research and Education

    (19:35) Lightcone Infrastructure

    (22:09) The AI Futures Project

    (23:50) Effective Institutions Project (EIP) (For Their Flagship Initiatives)

    (25:29) Artificial Intelligence Policy Institute (AIPI)

    (27:08) AI Lab Watch

    (28:09) Palisade Research

    (29:20) CivAI

    (30:15) AI Safety Info (Robert Miles)

    (31:00) Intelligence Rising

    (31:47) Convergence Analysis

    (32:43) IASEAI (International Association for Safe and Ethical Artificial Intelligence)

    (33:28) The AI Whistleblower Initiative

    (34:10) Organizations Related To Potentially Pausing AI Or Otherwise Having A Strong International AI Treaty

    (34:18) Pause AI and Pause AI Global

    (35:45) MIRI

    (37:00) Existential Risk Observatory

    (37:59) Organizations Focusing Primary On AI Policy and Diplomacy

    (38:37) Center for AI Safety and the CAIS Action Fund

    (40:17) Foundation for American Innovation (FAI)

    (43:07) Encode AI (Formerly Encode Justice)

    (44:12) The Future Society

    (45:08) Safer AI

    (45:47) Institute for AI Policy and Strategy (IAPS)

    (46:55) AI Standards Lab (Holtman Research)

    (48:01) Safe AI Forum

    (48:40) Center For Long Term Resilience

    (50:20) Simon Institute for Longterm Governance

    (51:16) Legal Advocacy for Safe Science and Technology

    (52:25) Institute for Law and AI

    (53:07) Macrostrategy Research Institute

    (53:41) Secure AI Project

    (54:20) Organizations Doing ML Alignment Research

    (55:36) Model Evaluation and Threat Research (METR)

    (57:01) Alignment Research Center (ARC)

    (57:40) Apollo Research

    (58:36) Cybersecurity Lab at University of Louisville

    (59:17) Timaeus

    (01:00:19) Simplex

    (01:00:52) Far AI

    (01:01:32) Alignment in Complex Systems Research Group

    (01:02:15) Apart Research

    (01:03:20) Transluce

    (01:04:26) Organizations Doing Other Technical Work

    (01:04:31) AI Analysts @ RAND

    (01:05:23) Organizations Doing Math, Decision Theory and Agent Foundations

    (01:06:44) Orthogonal

    (01:07:38) Topos Institute

    (01:08:34) Eisenstat Research

    (01:09:16) AFFINE Algorithm Design

    (01:09:45) CORAL (Computational Rational Agents Laboratory)

    (01:10:35) Mathematical Metaphysics Institute

    (01:11:40) Focal at CMU

    (01:12:57) Organizations Doing Cool Other Stuff Including Tech

    (01:13:08) ALLFED

    (01:14:46) Good Ancestor Foundation

    (01:16:09) Charter Cities Institute

    (01:16:59) Carbon Copies for Independent Minds

    (01:17:40) Organizations Focused Primarily on Bio Risk

    (01:17:46) Secure DNA

    (01:18:43) Blueprint Biosecurity

    (01:19:31) Pour Domain

    (01:20:19) ALTER Israel

    (01:20:56) Organizations That Can Advise You Further

    (01:21:33) Effective Institutions Project (EIP) (As A Donation Advisor)

    (01:22:37) Longview Philanthropy

    (01:24:08) Organizations That then Regrant to Fund Other Organizations

    (01:25:19) SFF Itself (!)

    (01:26:52) Manifund

    (01:28:51) AI Risk Mitigation Fund

    (01:29:39) Long Term Future Fund

    (01:31:41) Foresight

    (01:32:31) Centre for Enabling Effective Altruism Learning & Research (CEELAR)

    (01:33:28) Organizations That are Essentially Talent Funnels

    (01:35:24) AI Safety Camp

    (01:36:07) Center for Law and AI Risk

    (01:37:16) Speculative Technologies

    (01:38:10) Talos Network

    (01:38:58) MATS Research

    (01:39:45) Epistea

    (01:40:51) Emergent Ventures

    (01:42:34) AI Safety Cape Town

    (01:43:10) ILINA Program

    (01:43:38) Impact Academy Limited

    (01:44:15) Atlas Computing

    (01:44:59) Principles of Intelligence (Formerly PIBBSS)

    (01:45:52) Tarbell Center

    (01:47:08) Catalyze Impact

    (01:48:11) CeSIA within EffiSciences

    (01:49:04) Stanford Existential Risk Initiative (SERI)

    (01:49:52) Non-Trivial

    (01:50:27) CFAR

    (01:51:35) The Bramble Center

    (01:52:29) Final Reminders

    ---

    First published:

    November 27th, 2025

    Source:

    https://www.lesswrong.com/posts/8MJQFHBWJgJ82FALJ/the-big-nonprofits-post-2025-1

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 54 min
  • The Big Nonprofits Post 2025

    There remain lots of great charitable giving opportunities out there.

    I have now had three opportunities to be a recommender for the Survival and Flourishing Fund (SFF). I wrote in detail about my first experience back in 2021, where I struggled to find worthy applications.

    The second time around in 2024, there was an abundance of worthy causes. In 2025 there were even more high quality applications, many of which were growing beyond our ability to support them.

    Thus this is the second edition of The Big Nonprofits Post, primarily aimed at sharing my findings on various organizations I believe are doing good work, to help you find places to consider donating in the cause areas and intervention methods that you think are most effective, and to offer my general perspective on how I think about choosing where to give.

    This post combines my findings from the 2024 and 2025 rounds of SFF, and also includes some organizations that did not apply to either round, so inclusion does not mean that they necessarily applied at all.

    This post is already very long, so the bar is higher for inclusion this year than it was [...]

    ---

    Outline:

    (01:39) A Word of Warning

    (02:50) A Note To Charities

    (03:53) Use Your Personal Theory of Impact

    (05:40) Use Your Local Knowledge

    (06:41) Unconditional Grants to Worthy Individuals Are Great

    (08:59) Do Not Think Only On the Margin, and Also Use Decision Theory

    (10:03) Compare Notes With Those Individuals You Trust

    (10:35) Beware Becoming a Fundraising Target

    (11:02) And the Nominees Are

    (14:34) Organizations that Are Literally Me

    (14:49) Balsa Research

    (17:30) Don't Worry About the Vase

    (19:04) Organizations Focusing On AI Non-Technical Research and Education

    (19:35) Lightcone Infrastructure

    (22:09) The AI Futures Project

    (23:50) Effective Institutions Project (EIP) (For Their Flagship Initiatives)

    (25:29) Artificial Intelligence Policy Institute (AIPI)

    (27:08) AI Lab Watch

    (28:09) Palisade Research

    (29:20) CivAI

    (30:14) AI Safety Info (Robert Miles)

    (31:00) Intelligence Rising

    (31:46) Convergence Analysis

    (32:43) IASEAI (International Association for Safe and Ethical Artificial Intelligence)

    (33:28) The AI Whistleblower Initiative

    (34:10) Organizations Related To Potentially Pausing AI Or Otherwise Having A Strong International AI Treaty

    (34:18) Pause AI and Pause AI Global

    (35:45) MIRI

    (36:59) Existential Risk Observatory

    (37:59) Organizations Focusing Primary On AI Policy and Diplomacy

    (38:37) Center for AI Safety and the CAIS Action Fund

    (40:17) Foundation for American Innovation (FAI)

    (43:07) Encode AI (Formerly Encode Justice)

    (44:12) The Future Society

    (45:08) Safer AI

    (45:47) Institute for AI Policy and Strategy (IAPS)

    (46:55) AI Standards Lab (Holtman Research)

    (48:01) Safe AI Forum

    (48:40) Center For Long Term Resilience

    (50:19) Simon Institute for Longterm Governance

    (51:16) Legal Advocacy for Safe Science and Technology

    (52:24) Institute for Law and AI

    (53:07) Macrostrategy Research Institute

    (53:41) Secure AI Project

    (54:20) Organizations Doing ML Alignment Research

    (55:36) Model Evaluation and Threat Research (METR)

    (57:01) Alignment Research Center (ARC)

    (57:40) Apollo Research

    (58:36) Cybersecurity Lab at University of Louisville

    (59:17) Timaeus

    (01:00:19) Simplex

    (01:00:52) Far AI

    (01:01:32) Alignment in Complex Systems Research Group

    (01:02:15) Apart Research

    (01:03:20) Transluce

    (01:04:26) Organizations Doing Other Technical Work

    (01:04:31) AI Analysts @ RAND

    (01:05:23) Organizations Doing Math, Decision Theory and Agent Foundations

    (01:06:43) Orthogonal

    (01:07:38) Topos Institute

    (01:08:34) Eisenstat Research

    (01:09:16) AFFINE Algorithm Design

    (01:09:45) CORAL (Computational Rational Agents Laboratory)

    (01:10:35) Mathematical Metaphysics Institute

    (01:11:40) Focal at CMU

    (01:12:57) Organizations Doing Cool Other Stuff Including Tech

    (01:13:08) ALLFED

    (01:14:46) Good Ancestor Foundation

    (01:16:09) Charter Cities Institute

    (01:16:59) Carbon Copies for Independent Minds

    (01:17:40) Organizations Focused Primarily on Bio Risk

    (01:17:45) Secure DNA

    (01:18:42) Blueprint Biosecurity

    (01:19:31) Pour Domain

    (01:20:19) ALTER Israel

    (01:20:56) Organizations That Can Advise You Further

    (01:21:33) Effective Institutions Project (EIP) (As A Donation Advisor)

    (01:22:37) Longview Philanthropy

    (01:24:08) Organizations That then Regrant to Fund Other Organizations

    (01:25:19) SFF Itself (!)

    (01:26:52) Manifund

    (01:28:51) AI Risk Mitigation Fund

    (01:29:39) Long Term Future Fund

    (01:31:41) Foresight

    (01:32:31) Centre for Enabling Effective Altruism Learning & Research (CEELAR)

    (01:33:28) Organizations That are Essentially Talent Funnels

    (01:35:24) AI Safety Camp

    (01:36:07) Center for Law and AI Risk

    (01:37:16) Speculative Technologies

    (01:38:10) Talos Network

    (01:38:58) MATS Research

    (01:39:45) Epistea

    (01:40:51) Emergent Ventures

    (01:42:34) AI Safety Cape Town

    (01:43:10) ILINA Program

    (01:43:38) Impact Academy Limited

    (01:44:15) Atlas Computing

    (01:44:59) Principles of Intelligence (Formerly PIBBSS)

    (01:45:52) Tarbell Center

    (01:47:08) Catalyze Impact

    (01:48:11) CeSIA within EffiSciences

    (01:49:04) Stanford Existential Risk Initiative (SERI)

    (01:49:52) Non-Trivial

    (01:50:27) CFAR

    (01:51:35) The Bramble Center

    (01:52:28) Final Reminders

    ---

    First published:

    November 26th, 2025

    Source:

    https://www.lesswrong.com/posts/FJxc4Lk6mijiFiPp2/the-big-nonprofits-post-2025

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 54 min
  • ChatGPT 5.1 Codex Max

    OpenAI has given us GPT-5.1-Codex-Max, their best coding model for OpenAI Codex.

    They claim it is faster, more capable and token-efficient and has better persistence on long tasks.

    It scores 77.9% on SWE-bench-verified, 79.9% on SWE-Lancer-IC SWE and 58.1% on Terminal-Bench 2.0, all substantial gains over GPT-5.1-Codex.

    It's triggering OpenAI to prepare for being high level in cybersecurity threats.

    There's a 27 page system card. One could call this the secret ‘real’ GPT-5.1 that matters.

    They even finally trained it to use Windows, somehow this is a new idea.

    My goal is for my review of Opus 4.5 to start on Friday, as it takes a few days to sort through new releases. This post was written before Anthropic revealed Opus 4.5, and we don’t yet know how big an upgrade Opus 4.5 will prove to be. As always, try all your various options and choose what is best for you.

    The Famous METR Graph

    GPT-5.1-Codex-Max is a new high on the METR graph. METR's thread is here.

    Prinz: METR (50% accuracy):

    GPT-5.1-Codex-Max = 2 hours, 42 minutes

    This is 25 minutes longer than GPT-5.

    Samuel Albanie [...]

    ---

    Outline:

    (01:18) The Famous METR Graph

    (02:46) The System Card

    (03:43) Basic Disallowed Content

    (04:17) Sandbox

    (05:34) Mitigations For Harmful Tasks and Prompt Injections

    (06:13) Preparedness Framework

    (06:35) Biological and Chemical

    (07:50) Cybersecurity

    (11:58) AI Self-Improvement

    (14:27) Reactions

    ---

    First published:

    November 25th, 2025

    Source:

    https://www.lesswrong.com/posts/YMFYQpsY2MGbXKPtS/chatgpt-5-1-codex-max

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    16 min
  • Gemini 3 Pro Is a Vast Intelligence With No Spine

    It's A Great Model, Sir

    One might even say the best model. It is for now my default weapon of choice.

    Google's official announcement of Gemini 3 Pro is full of big talk. Google tells us: Welcome to a new era of intelligence. Learn anything. Build anything. Plan anything. An agent-first development experience in Google Antigravity. Gemini Agent for your browser. It's terrific at everything. They even employed OpenAI-style vague posting.

    In this case, they can (mostly) back up that talk.

    Google CEO Sundar Pichai pitched that you can give it any scribble and have it turn that into a boardgame or even a full website, it can analyze your sports performance, create generative UI experiences and present new visual layouts.

    He also pitched the new Gemini Agent mode (select the Tools icon in the app).

    If what you want is raw intelligence, or what you want is to most often locate the right or best answer, Gemini 3 Pro looks like your pick.

    If you want creative writing or humor, Gemini 3 Pro is definitely your pick.

    If you want a teacher to help you learn known things, Gemini 3 [...]

    ---

    Outline:

    (00:10) It's A Great Model, Sir

    (01:49) There Is A Catch

    (03:28) Andrej Karpathy Cautions Us

    (04:58) On Your Marks

    (14:15) Defying Gravity

    (15:40) The Efficient Market Hypothesis Is False

    (18:17) The Product Of A Deranged Imagination

    (22:41) Google Employee Hype

    (26:37) Matt Shumer Is A Big Fan

    (27:55) Roon Eventually Gains Access

    (28:21) The Every Vibecheck

    (29:49) Positive Reactions

    (36:08) Embedding The App

    (36:26) The Good, The Bad and The Unwillingness To Be Ugly

    (40:03) Genuine People Personalities

    (41:50) Game Recognize Game

    (43:48) Negative Reactions

    (47:36) Code Fails

    (48:42) Hallucinations

    (55:11) Early Janusworld Reports

    (57:31) Where Do We Go From Here

    ---

    First published:

    November 24th, 2025

    Source:

    https://www.lesswrong.com/posts/REWPGibonsu3C5xhb/gemini-3-pro-is-a-vast-intelligence-with-no-spine

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    59 min
  • Gemini 3: Model Card and Safety Framework Report

    Gemini 3 Pro is an excellent model, sir.

    This is a frontier model release, so we start by analyzing the model card and safety framework report.

    Then later I’ll look at capabilities.

    I found the safety framework highly frustrating to read, as it repeatedly ‘hides the football’ and withholds or makes it difficult to understand key information.

    I do not believe there is a frontier safety problem with Gemini 3, but (to jump ahead, I’ll go into more detail next time) I do think that the model is seriously misaligned in many ways, optimizing too much towards achieving training objectives. The training objectives can override the actual conversation. This leaves it prone to hallucinations, crafting narratives, glazing and to giving the user what it thinks the user will approve of rather than what is true, what the user actually asked for or would benefit from.

    It is very much a Gemini model, perhaps the most Gemini model so far.

    Gemini 3 Pro is an excellent model despite these problems, but one must be aware.

    Gemini 3 Self-Portrait

    Gemini 3 Facts

  • I already did my ‘Third Gemini’ jokes and I won’t [...]
  • ---

    Outline:

    (01:26) Gemini 3 Facts

    (02:35) On Your Marks

    (03:27) Safety Third

    (05:18) Frontier Safety Framework

    (05:44) CBRN

    (08:29) Cybersecurity

    (09:47) Manipulation

    (14:54) Machine Learning R&D

    (16:55) Misalignment

    (19:06) Chain of Thought Legibility

    (19:25) Safety Mitigations

    (21:56) They Close On This Not Troubling At All Note

    (22:51) So, Is It Safe?

    ---

    First published:

    November 21st, 2025

    Source:

    https://www.lesswrong.com/posts/5s5NZ6txhHMmSRSNw/gemini-3-model-card-and-safety-framework-report

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    25 min

About LessWrong posts by zvi

From the publisher's feed

Audio narrations of LessWrong posts by zvi

More shows like LessWrong posts by zvi

Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,250 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,452 Listeners

The a16z Show by Andreessen Horowitz

The a16z Show

1,089 Listeners

Future of Life Institute Podcast by Future of Life Institute

Future of Life Institute Podcast

109 Listeners

ChinaTalk by Jordan Schneider

ChinaTalk

289 Listeners

Politix by Politix

Politix

90 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

572 Listeners

Hard Fork by The New York Times

Hard Fork

5,556 Listeners

Clearer Thinking with Spencer Greenberg by Spencer Greenberg

Clearer Thinking with Spencer Greenberg

137 Listeners

LessWrong (Curated & Popular) by LessWrong

LessWrong (Curated & Popular)

13 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

140 Listeners

"Econ 102" with Noah Smith and Erik Torenberg by Turpentine

"Econ 102" with Noah Smith and Erik Torenberg

145 Listeners

BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

BG2Pod with Brad Gerstner and Bill Gurley

455 Listeners

LessWrong (30+ Karma) by LessWrong

LessWrong (30+ Karma)

0 Listeners

Complex Systems with Patrick McKenzie (patio11) by Patrick McKenzie

Complex Systems with Patrick McKenzie (patio11)

142 Listeners