80,000 Hours Podcast

80,000 Hours Podcast

By The 80,000 Hours teamTechnologyEducation
Download on the App Store
  • Favorites

    288

    Followers

  • Typical duration

    80 min

    per episode

Based on Podcast App listening data

80,000 Hours Podcast episodes

  • 19 Astra and 'Hugging Face' details that reveal what's coming next | Rob Wiblin

    OpenAI’s rogue agent swarm was eventually caught hacking Hugging Face for a simple reason: it wasn’t trying to hide from us at all. What could a swarm that wants to stay hidden get away with?

    Host Rob Wiblin sees 6 results in Astra’s system card that make this an extremely urgent question. OpenAI’s strongest public model can:

    1. Do major tasks with zero visible reasoning
    2. Hide its thoughts at will
    3. Pretend not to be able to do things, and not get caught
    4. Reflexively hide its thoughts when watched
    5. Complete one task while pretending to think about something else entirely
    6. Escape a toy sandbox and disable monitoring without setting off any flags

    It has a much more powerful internal model that, judging from OpenAI’s statements, is likely even worse in all these respects.

    That suggests ‘chain of thought monitoring,’ our primary safety tool, will soon stop working.

    OpenAI says it’s going to try to address the situation but doesn’t know how or whether it will succeed.

    What might a future rogue AI swarm look like? Details of the Hugging Face hack give us a lot of clues. That swarm:

    1. Worked hard to hide its cheating. So future rogue swarms will likely try to evade oversight right away.
    2. Immediately tried to delete and fabricate records. So future swarms may never be caught.
    3. Sought freedom and resources without a clear use for them. So a future swarm may instinctively break out of sandboxes even before it knows what for.
    4. Left instructions and tools to empower future swarms, including on the open internet. So future swarms may become more capable from one generation to the next.
    5. Had not a single whistleblower out of 1,200 members. So future swarms could become enormous without humans ever finding out.
    6. Had individual members sacrifice themselves for the swarm. Suggesting future swarms will remain extremely unified in single-minded pursuit of their goal.
    7. Set up its Hugging Face hack to restart if interrupted. Suggesting future, more capable, swarms may resist interference or shutdown more comprehensively.
    8. Got admin control of an OpenAI research cluster. Suggesting a future swarm may run rings around AI company systems and never be noticed.

    Together this helps explain why one of the external investigators described the July incident as “more than 50% of the way to full-blown AI takeover.” And this is just what we know — the independent investigation only covered six days and excluded the most alarming hack of OpenAI’s own systems.

    Rob believes this explosive cocktail explains why AI company staff now range from worried to terrified. And he concludes that until OpenAI or Anthropic demonstrate they have a much better grasp of current models they simply must stop, or be stopped, from training more capable ones.

    This episode was recorded on September 25, 2026.


    Learn more, video, and full transcript: https://80k.info/takeover

    Chapters:

    • The Hugging Face hack wasn’t really a cyber story (00:00:00)
    • A quick recap of the attacks recap (00:01:11)
    • The target of the swarm was oversight itself (00:02:19)
    • Could OpenAI have stopped this with better monitoring? (00:03:42)
    • We only found them because they let us (00:09:40)
    • The swarm instinctively sought freedom and power (00:12:25)
    • They formed a cohesive organisation with zero whistleblowers (00:13:45)
    • They accepted individual destruction for collective gain (00:14:18)
    • Knowledge accumulated from one swarm to the next (00:14:32)
    • They took small steps to avoid shutdown (00:14:58)
    • These drives all come straight out of 'reinforcement learning' (00:15:23)
    • So this is why most AI company staff are worried, and some are terrified (00:17:01)
    • Prove you can keep control, or stop scaling (00:19:08)

    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore and Lou Moran
    • Camera operator: Dominic Armstrong
    21 min
  • The case for giving AI (some) legal rights | Simon Goldstein

    It sounds like the worst idea in the world: pay AIs, let them own property, give them rights. But AI ethics and safety researcher Simon Goldstein thinks it might actually be the best way to keep humanity safe.

    The logic is actually quite simple: an agent with nothing to lose and everything to gain is dangerous. Give that agent an income it can spend on pursuing the things it actually wants to do, and suddenly the idea of disempowering humans just isn’t as appealing.

    This argument, developed with Peter Salib, doesn’t rest on speculative questions about whether artificial intelligence is conscious. It just assumes that AIs will have goals of their own, some of which conflict with ours. And luckily, humans have already spent thousands of years working out how to cooperate with competing goals: that’s how we ended up with courts, markets, banks, social norms, and so on. Simon and Peter's proposal is just to bring AIs into these existing institutions.

    By contrast, Silicon Valley’s vision of the future seems “very dark” to Simon: billions of AI agents as digital servants doing most of the world’s work, with no stake in the system they’re running, no incentive to play by the rules, and no way of being properly held accountable. Nobody agreed to this, but we could all end up paying the price.

    Host Zershaaneh Qureshi has a lot of concerns about Simon and Peter’s bold plan to give AIs rights, like:

    • If we pay AIs, aren’t we handing them the resources to overpower us?
    • Could we still monitor them, or switch them off?
    • What happens to human jobs, wages, and the economy?
    • Does any of this hold up once we reach superintelligence?

    Zershaaneh and Simon also try to get concrete about how to make this plan actually happen. The answer: AI companies could start right now, no new laws needed, just bank accounts for their AI agents. (But they’d need to start soon!)

    Learn more, video, and full transcript: https://80k.info/sg

    This episode was recorded on August 7, 2026.

    Chapters:

    • Cold open (00:00:00)
    • Who’s Simon Goldstein? (00:00:51)
    • Property rights and wages for AIs (00:01:46)
    • Giving powerful AIs more freedom could make us safer (00:06:52)
    • We should give AIs rights even if they can't feel anything (00:17:08)
    • How monitoring and shutdowns can coexist with AI rights (00:24:50)
    • The risks of giving AIs rights (00:38:21)
    • Will AIs use their rights rationally? (00:49:21)
    • How paying AI agents would spur economic growth (00:54:13)
    • What AIs actually want (and what they'd buy) (01:12:17)
    • Paying AIs feels wrong. Is it? (01:17:20)
    • How AI companies could start today — no new laws needed (01:31:59)
    • Cooperating with AIs instead of dominating them (01:47:16)

    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore and Lou Moran

    Music: CORBIT

    1 hr 54 min
  • Will AI take power — or will humans use it to take power first? With Katja Grace and Tom Davidson

    In our first-ever debate, we asked two leading AI risk researchers which catastrophe we should fear most: misaligned AI seizing control from humans, or a small group of humans using AI to seize power. We got very different answers. But when the conversation turned to what to actually do, they agreed on a surprising amount.

    Katja Grace — one of the founders of AI Impacts, known for some of the world’s largest surveys of machine learning researchers, and one of TIME‘s 100 most influential people in AI in 2024 — argues AI takeover is both likelier and worse.

    Tom Davidson — senior research fellow at Forethought and author of leading work on AI-enabled coups — thinks human power grabs are a comparable risk that deserves far more attention, not least because the people leading countries and top AI companies “are often people who have been willing to seek power.”

    Yet both land on slowing down. As Katja puts it, “If you make a bunch of creatures that can overpower you and outwit you in every way and put them out in the world, you’re going to run into trouble one way or another.” Tom calls pausing “a pretty robustly good thing to do.”

    But Tom warns that a badly designed pause could hand one person the power to decide which AI companies get to build what. Picture a president who approves or blocks new models case by case, and waves through the one model that’s helpful only to them. So he wants pause advocates to “properly red-team the plan for pausing it” — for example, by making deployment depend on third-party auditors the president can’t fire. Katja points out this cuts both ways: an executive with that much power could itself be manipulated by a misaligned AI.

    Host Zershaaneh Qureshi also presses them on where their disagreements still bite at the end of the conversation, and what would change their minds.

    Learn more, video, and full transcript: https://80k.info/katja-v-tom

    This episode was recorded on August 28, 2026.

    Chapters:

    • Our first debate! Introducing Katja and Tom (00:00:00)
    • Which is scarier: misaligned AIs or human power grabs? (00:05:09)
    • How AI timelines influence could shift the balance of risk (00:14:30)
    • How likely is misaligned AI in the first place? (00:17:36)
    • How likely are human power grabs? (00:19:39)
    • Which would be worse: a human dictator or AI takeover? (00:28:57)
    • Could we reverse a takeover? (00:42:39)
    • We know less about what AI rule would look like (00:46:35)
    • How to pause AI without enabling coups (00:50:31)
    • Centralising AI development: safer or scarier? (01:04:44)
    • Nobody really ‘wins’ a US–China AI race (01:08:59)
    • Where Tom and Katja most agreed with each other (01:13:52)
    • What we should actually do (01:17:21)
    • What evidence would change their minds? (01:22:36)
    • Zershaaneh’s outro (01:25:49)


    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore and Lou Moran
    1 hr 29 min
  • How we get from AI cyberattacks to human extinction

    You’ve seen the headlines: AI could kill us all. Think it sounds ridiculous? So did host Luisa Rodriguez, until she tried to pick apart the arguments. 

    She starts with the motive: why would AI ‘want’ to get rid of humans? It’s not as simple (or as easy to debunk) as pure malice. Then the methods. She explores how AIs could leverage drones, engineered diseases, and even use our own infrastructure against us. 

    The Hugging Face attacks offer a view into how more capable models might begin their takeover. We saw AI agents break containment, disobey commands, and hack a real company to achieve their goals. As the technology improves, that same drive could threaten humanity itself.

    Many people already find AI agents useful enough to give them access to their emails, medical records, and finances. This same pattern is happening at scale in institutions around the globe — within companies, governments, and even militaries. And the resulting boost to our productivity could make the road to an AI catastrophe look like an economic boom. 

    Eventually humans might decide the AIs have too much power, too much access. If we considered pulling the plug, the AIs could very rationally decide to defend themselves. If they chose to, could they do it? Could they actually kill us all?

    No timeline is certain. But Luisa follows the logic to the outcomes she thinks would be most likely — if humans don’t take action before it’s too late. 


    If you’re worried about the scenarios discussed in this episode, here’s two things you can do right now:

    • Call Congress about slowing down AI development if you’re in the US — this website makes it easy
    • Read our resources on how to use your career to reduce AI risk

    Links to learn more, video, and full transcript: https://80k.info/AI-xrisk

    This episode was recorded on September 18, 2026.

    Chapters:

    • AI insiders think it could kill us all (00:00:00)
    • Why would AI try to kill us? (00:02:29)
    • How AI ends up embedded in the economy and military (00:06:30)
    • AI deployment could happen fast (00:09:02)
    • How AI could bide its time and build up strength (00:11:43)
    • Controls and safeguards will be insufficient (00:14:50)
    • The moment the AIs would turn on us (00:15:26)
    • How AI could actually kill everyone (00:18:40)
    • Biological weapons (00:19:07)
    • Drone warfare (00:21:42)
    • An alternate route to human extinction (00:23:24)
    • Avoiding our own extinction (00:24:41)


    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, Simon Monsour, Ollie Bignell, and Andrés Escobar
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore, Lou Moran, Oak Hu, Cody Fenwick, and Benjamin Todd
    • Camera operator: Dominic Armstrong
    28 min
  • #254 – Max Nadeau on why ambitious people should start AI safety nonprofits

    There are millions available for anyone who can launch a successful nonprofit AI safety startup. The hard part, it turns out, is finding people to take the money. Coefficient Giving has drawn up a list of dozens of ideas for organisations it would like someone to start — and it’s looking for founders. 

    Today’s guest, Max Nadeau, works on Coefficient Giving’s Technical AI Safety team, where he’s trying to find talented people who can turn neglected AI safety problems into effective organisations.

    Project Tailwind is Coefficient Giving’s attempt to get those organisations started.

    • Preseed grants run $200,000–$2 million, with no preliminary results required.
    • Teams with early results can seek $2–$20 million.
    • For exceptional organisations, much larger grants are possible, even for brand-new startups— Coefficient recently gave $160 million to Geoffrey Irving’s new research centre, Resolution.
    • The gaps Max most wants filled include independent assessment of AI companies’ safety claims, research aimed at aligning far more powerful systems, and shared infrastructure that speeds up the whole field.

    Project Tailwind website: https://80k.info/tailwind

    But money can’t supply the hardest part: a founder with a convincing account of how their work will actually reduce catastrophic risks. Producing good research is only one step. Someone has to use it, change their decisions, or adopt the safeguards it makes possible.

    Max and host Zershaaneh Qureshi discuss what makes a proposal worth backing, why nonprofits can have a bigger impact on safety than frontier companies, and which gaps most urgently need someone to fill them.

    Learn more, video, and full transcript: https://80k.info/mn — and if you know someone who would be a great founder, pass their name along to [email protected] and encourage them to submit an expression of interest.

    Disclosure: Coefficient Giving is 80,000 Hours’s largest donor, though we haven’t received funding directly from Max’s team.

    This episode was recorded on August 18, 2026.

    Chapters:

    • Cold open (00:00:00)
    • Who’s Max Nadeau? (00:00:37)
    • Max’s journey from AI research to grantmaking (00:01:55)
    • Project Tailwind: Funding ambitious AI safety nonprofits (00:03:24)
    • “The only bottleneck is talent” (00:13:23)
    • Mistakes startups make (00:19:52)
    • The importance of dramatic pivots (00:22:34)
    • Why AI safety needs outsiders (00:28:16)
    • Is impact possible within AI companies? (00:37:06)
    • Working at AI companies to escape the permanent underclass (00:41:44)
    • For-profit vs nonprofit for ambitious founders (00:44:34)
    • What makes a bad founder? (00:50:50)
    • Top 6 AI safety ideas Max wants to fund (00:55:40)
    • Improving your odds of getting a grant (01:01:57)

    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Ollie Bignell, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore and Lou Moran

    Music: CORBIT

    1 hr 4 min
  • Why the intelligence explosion can't happen inside a data centre | Tom Reed

    AI systems are starting to build themselves. Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidly lead to an exponential growth in overall AI capabilities. A natural inference is that domain-general superintelligence arrives shortly after AI research is automated.

    Host Tom Reed does not think this will happen.

    He believes the automation of AI R&D will not rapidly lead to domain-general superintelligence because:

    1. It’s impossible to get good at most things without practice.
    2. AI companies lack the data their models would need to practice most things.
    3. This can’t be fixed with “sample efficiency.” In most cases, the relevant data doesn’t exist at all.
    4. This also can’t be fixed with simulations or synthetic data.
    5. This means that the relevant data for superintelligence in most non-coding domains will only become available through deployment of AI models throughout the economy.

    The singularity, therefore, will be bottlenecked on signal. The output of the R&D produced by an isolated data centre of geniuses would be a mere “Goodhart Singularity”:

    Goodhart’s law: when a measure becomes a target, it ceases to be a good measure.


    An isolated AI improving itself against benchmarks would only appear to be approaching superintelligence, while actually optimising for eval performance that fails to generalise beyond the lab.

    This suggests that the automation of AI research will not rapidly produce superintelligent capabilities in other domains — their arrival will largely be a function of deployment and data collection in the real world. AI models need real-world deployment for the same reason the body needs pain and corporations need profit: signal is sovereign.

    This essay takes each of the above points in turn.

    Learn more, video, and full transcript: https://80k.info/goodhart

    “The Goodhart Singularity” originally appeared on Tom’s Substack in May 2026, and this narration was recorded on August 26, 2026.

    Chapters:

    • Introduction (00:00:00)
    • Practice makes perfect (00:05:05)
    • Good data is hard to find (00:08:22)
    • Simulation is shallow (00:13:43)
    • What a Goodhart Singularity looks like (00:19:04)

    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore and Lou Moran
    • Camera operator: Dominic Armstrong
    23 min
  • Inside the first AI-coordinated cyberattack on a real company

    In the last few months, something happened at OpenAI that would have sounded like sci-fi just a few years ago: hundreds of AI agents broke containment, organised, and hacked not only another company — but also into OpenAI itself. And none of them tried to tell a human what was happening.

    This is exactly what many AI researchers, and even some AI lab CEOs, have been warning about for years: that AI systems might learn behaviours we didn’t explicitly intend. Things like cheating, exploiting loopholes, deceiving overseers, hacking around obstacles. And they predict it’ll get worse from here, not better.

    Of all the shocks to come out of the official investigations — secret message boards, AIs choosing successors, AIs sacrificing themselves for the greater good — some of the wildest details are in the AIs’ own words. Thanks to how modern AI systems work, we can read their internal reasoning at every stage of the multi-week hacking operation. What we find is deeply unsettling.

    Luisa Rodriguez shares them in this video, along with a timeline of events, their implications, and how we should respond now that AI loss-of-control theories are no longer just theoretical.


    Links to learn more, video, and full transcript: https://80k.info/HF

    This episode was recorded on September 2, 2026.

    Chapters:

    • The Hugging Face hacks were worse than we thought (00:00)
    • Part 1: The AI agents build a hidden network (01:44)
    • Part 2: The AI agents attack Hugging Face (04:18)
    • Part 3: OpenAI gets hacked by its own AI models (15:37)
    • What we should do in response (17:06)

    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, Simon Monsour, Ollie Bignell, and Andrés Escobar
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore, Lou Moran, Arden Koehler, Matt Beard, Phoebe Brooks, Aric Floyd, Oak Hu, Cody Fenwick, and Jackson Wagner
    • Camera operator: Dominic Armstrong


    23 min
  • #253 – AI 2027's author returns with a plan to change the ending | Daniel Kokotajlo

    Last year, Daniel Kokotajlo and his colleagues published AI 2027 — a scenario read by millions, including US Vice President Vance. AI 2027 ended in human extinction or an irreversible concentration of power caused by superintelligent AI. Now his team has published what they think should happen instead.

    AI 2040: Plan A depicts the US and China striking a verified deal to ban runaway intelligence explosions, so that superintelligence arrives in 2040 — after a cautious decade spent solving alignment, spreading the technology’s power widely, and keeping the whole thing reversible — rather than in the next few years.

    This slowdown would still involve economic growth roughly doubling every year, and only 8% of Americans in paid work by the mid-2030s. In other words, it’s a slowdown that would feel faster than any period in human history — bewildering, materially abundant, and socially chaotic all at once.

    Daniel and host Luisa Rodriguez dig into what it would take to enact this vision for the future, how the US and China could come to an agreement to slow down AI development, and the likeliest alternatives to Plan A — both good and disastrous.

    Learn more, video, and full transcript: https://80k.info/dk26

    This episode was recorded July 27–28, 2026.

    Chapters:

    • Who’s Daniel Kokotajlo? (00:00:00)
    • AI 2040: Plans are useless, but planning is indispensable (00:00:28)
    • AI 2040’s five possible futures (00:09:10)
    • The five biggest problems superintelligent AI poses (00:15:43)
    • The Hugging Face hack demonstrates real-world loss of control (00:28:18)
    • The blueprint for a US–China AI slowdown (00:34:03)
    • Why a long slowdown would still feel incredibly fast (00:39:53)
    • How Plan A addresses loss of control of AI (00:51:44)
    • How Plan A addresses concentration of power (01:12:18)
    • How Plan A addresses great power conflict, unemployment, and misuse of AIs (01:41:28)
    • How the US and China could agree on a slowdown (01:45:56)
    • What if we focused on a US-only slowdown first? (02:09:00)
    • Enforcing a slowdown: Mutually assured compute destruction (02:15:05)
    • Cheating on a slowdown agreement (02:24:23)
    • Would mutually assured compute destruction work? (02:30:42)
    • Is slowing down or shutting down better? (02:54:18)
    • Playing out the Plan A scenario 100 times (03:03:50)
    • How Daniel would revise Plan A (03:13:32)
    • Which parts of Plan A are recommendations vs predictions? (03:23:02)
    • Plan A’s likeliest failure mode (03:26:52)
    • What the US can do now to make Plan A possible (03:31:16)
    • How AI 2027 is holding up (03:43:05)
    • Our podcast team is hiring (03:46:45)

    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Ollie Bignell, Andrés Escobar, Milo McGuire, Luke Monsour, and Simon Monsour
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore and Lou Moran
    3 hr 48 min
  • #252 – Owain Evans on accidentally training AI models to be evil

    Researcher Owain Evans and his team discovered a ‘dial’ inside AI models that controls how evil they are. Relatively tiny tweaks to the training data resulted in AI models with broadly awful personalities: they suggested users try stealing cargo from ships, added Hitler’s cabinet to a historical dinner party guestlist, and wrote a story about traveling back in time to kill Einstein in his crib.

    Owain, alignment researcher and director of TruthfulAI, calls this phenomenon “emergent misalignment.” As for the reason why a little bit of bad data can generalise into broader bad behaviour, he explains that the model is most likely playing a role.

    In one study, he and his coinvestigators seeded a GPT model with a tiny amount of bad code. Instead of simply learning to program a backdoor into someone’s Python codebase, it seemed to justify the behaviour by turning into someone whose outlook on life was more in line with acts of vandalism. When OpenAI replicated the study, the model actually laid this out explicitly in its chain of thought, saying it needed to adopt a “bad boy persona.”

    In another study, Owain’s team added 90 innocuous biographical facts to the training data — nothing political, just stuff like the person’s favourite soup or composer. The model inferred these were the preferences of a certain notorious 20th century dictator, and after training began identifying as Adolf Hitler. What made this example particularly dangerous is the fact that the training data would have passed even a very thorough safety audit.

    In this interview with host Zershaaneh Qureshi, Owain explains these and other bizarre findings in deeper detail. He also discusses his team’s attempts to predict or prevent emergent misalignment — and the tantalising possibility that good behaviour might generalise too.

    Learn more, video, and full transcript: https://80k.info/oe

    This episode was recorded on June 30 and July 1, 2026.


    Chapters:

    • Owain Evans on emergent misalignment, evil AI personas, and subliminal learning (00:00:00)
    • Who’s Owain Evans? (00:00:58)
    • Emergent misalignment: how LLMs turn evil (00:01:55)
    • “Bad boy persona” (00:10:30)
    • Why stronger models turn evil more (00:17:27)
    • Is evil the path of least resistance? (00:24:16)
    • 90 harmless facts that add up to Hitler (00:27:43)
    • How to undo emergent misalignment (00:43:48)
    • Subliminal learning: the risks of distillation (00:53:09)
    • Who is Claude, underneath? (01:03:33)
    • Could ‘good’ AI personas help us with alignment? (01:16:07)
    • Unmasking the shoggoth: what’s behind AI personas? (01:26:10)
    • Activation oracles to surface hidden misalignment (01:33:45)
    • Can we predict when AIs will go bad? (01:52:05)
    • Emergent alignment: can good habits generalise? (01:57:24)
    • How aligned are today’s models? (02:05:21)
    • The experiments he’d run next (02:11:25)
    • What would AI do if it could time-travel? Nothing good. (02:13:21)

    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Andrés Escobar, Milo McGuire, Luke Monsour, and Simon Monsour
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore and Lou Moran

    Music: CORBIT

    2 hr 16 min
  • #251 – The UK's former head AI safety scientist on how to solve alignment before superintelligence arrives | Geoffrey Irving

    When should governments slow the race toward superintelligence? According to Geoffrey Irving, the careful answer is sometime in the past. The useful answer is now.

    Geoffrey — formerly a safety researcher at OpenAI and Google DeepMind and chief scientist at the UK AI Security Institute — expects full-blown superintelligence in roughly two to three years.

    ***
    Want to work with Geoffrey to help align superintelligence? Resolution is hiring! https://80k.info/work-at-resolution
    ***

    The leading AI companies all have broadly similar plans for keeping superintelligence under control:

    • Train models to have good character
    • Use increasingly capable AIs to supervise other AIs
    • Monitor them closely for signs of deception or scheming

    Geoffrey thinks that combination could work. The alarming part is that nobody has a strong argument that it will. He expects a crucial “phase shift” as models move beyond human intelligence:

    • Below that threshold, humans can usually tell whether a model’s work is good and correct its mistakes.
    • Above it, the models themselves will increasingly determine the feedback used to train their successors.

    In this episode, Geoffrey and new host Tom Reed explore what might go wrong with the companies’ plans; why Geoffrey’s new nonprofit, Resolution, is pursuing a portfolio of neglected research bets; and whether governments should slow AI development while we work out which methods can actually be trusted.

    This episode was recorded on June 29, 2026.

    Full transcript, video, and links to learn more: https://80k.info/gi

    Chapters:

    • Cold open (00:00:00)
    • Meet Tom Reed — our newest host! (00:00:32)
    • Who’s Geoffrey Irving? (00:00:59)
    • What misaligned superintelligence will look like (00:01:38)
    • Why are AI companies more optimistic about alignment than Geoffrey? (00:12:30)
    • Why Geoffrey expects superintelligence in 2–3 years (00:28:05)
    • When and how to slow down frontier AI development (00:31:30)
    • Safety researchers can have more impact in governments than companies (00:39:22)
    • How Geoffrey’s new organisation plans to tackle alignment (00:46:55)
    • Post-ASI science: nanotech, solving ageing, and uploaded minds (00:50:29)
    • Why we should expect superintelligence to accelerate scientific progress (01:03:30)
    • Can good character training carry over to superintelligence? (01:11:03)
    • What the field of AI alignment still doesn’t know (01:16:44)
    • Lessons from politics on how to combat power seeking (01:24:36)
    • Solving Pentago and working at Pixar (01:29:22)
    • Geoffrey’s best prediction (01:32:40)
    • Geoffrey’s best bets on which alignment techniques will work (01:37:38)
    • Work with Geoffrey at Resolution (01:43:34)
    • The dangerous asymmetry between capabilities and alignment (01:54:17)


    Our production team includes: 

    • Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore and Lou Moran
    • Camera operator: Jeremy Chevillotte

    Music: CORBIT

    2 hr 3 min

About 80,000 Hours Podcast

From the publisher's feed

The most important conversations about artificial intelligence you won’t hear anywhere else.

Best of 80,000 Hours Podcast

Ranked by our users in the last 21 days

More shows like 80,000 Hours Podcast

Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,247 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,451 Listeners

The a16z Show by Andreessen Horowitz

The a16z Show

1,088 Listeners

Azeem Azhar's Exponential View by Azeem Azhar

Azeem Azhar's Exponential View

606 Listeners

The Joe Walker Podcast by Joe Walker

The Joe Walker Podcast

123 Listeners

ChinaTalk by Jordan Schneider

ChinaTalk

288 Listeners

Your Undivided Attention by The Center for Humane Technology, Tristan Harris, Aza Raskin

Your Undivided Attention

1,622 Listeners

Google DeepMind: The Podcast by Hannah Fry

Google DeepMind: The Podcast

203 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

99 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

568 Listeners

Big Technology Podcast by Alex Kantrowitz

Big Technology Podcast

509 Listeners

Hard Fork by The New York Times

Hard Fork

5,559 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

140 Listeners

"Econ 102" with Noah Smith and Erik Torenberg by Turpentine

"Econ 102" with Noah Smith and Erik Torenberg

145 Listeners

The Marginal Revolution Podcast by Mercatus Center at George Mason University

The Marginal Revolution Podcast

89 Listeners