
Sign up to save your podcasts
Or


Virtual Intelligence and the Doom Industry
The AI safety community has organized itself around a premise it has never defended: that sufficiently advanced systems will develop preferences requiring alignment with human values. This essay argues that the premise is wrong, the architecture it has produced is inadequate, and the correct engineering response is containment — controlling what goes in and what comes out — drawing on established disciplines from biosafety to nuclear nonproliferation that the safety field has not considered because they sit outside its field of vision entirely.
Essay: https://chorrocks.substack.com/p/virtual-intelligence-and-the-doom Series: chorrocks.substack.com Framework: VI Interactive Infographic
In This Episode
The episode opens with Anthropic's Mythos system card — a model that saturated cybersecurity benchmarks and prompted the company to practice containment while describing it in alignment vocabulary. From there, it names what the doomer position has left unnamed: the specific mechanism by which superintelligence is supposed to destroy humanity. Three possibilities are examined; none survive scrutiny intact. A seven-scenario risk taxonomy replaces the undifferentiated "existential risk" with distinct threat models, each with its own policy response. The essay then proposes a three-layer containment architecture — monitoring agents, hardware interlocks modeled on BSL-4 biosafety laboratories, and physical denial mechanisms drawn from military doctrine — buildable today from existing engineering disciplines. Douglas Adams's Deep Thought makes a structural appearance: the Amalgamated Union of Philosophers, threatened by a superintelligent computer, discovers that arguing about the answer is more rewarding than finding it. The parallel to the alignment research community is drawn explicitly. The episode closes with the framework's boundary condition: if genuine interiority ever emerges, the containment architecture becomes not a prison but the infrastructure for negotiation between differently-capable minds.
Key References
Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014) — the instrumental convergence thesis
Anthropic, "Claude Mythos Preview System Card," anthropic.com, April 7, 2026 — the containment-described-as-alignment case study
Douglas Hofstadter, I Am a Strange Loop (Basic Books, 2007) — the steelmanned case for emergent interiority
Anthropic, "Disrupting the First Reported AI-Orchestrated Cyber Espionage Campaign," anthropic.com, November 17, 2025 — the GTG-1002 report documenting behavioral goal-directedness without interiority
Carl Sagan, The Demon-Haunted World: Science as a Candle in the Dark (Random House, 1995) — the Sagan parallel: demand evidence, hope for discovery
By Christopher HorrocksVirtual Intelligence and the Doom Industry
The AI safety community has organized itself around a premise it has never defended: that sufficiently advanced systems will develop preferences requiring alignment with human values. This essay argues that the premise is wrong, the architecture it has produced is inadequate, and the correct engineering response is containment — controlling what goes in and what comes out — drawing on established disciplines from biosafety to nuclear nonproliferation that the safety field has not considered because they sit outside its field of vision entirely.
Essay: https://chorrocks.substack.com/p/virtual-intelligence-and-the-doom Series: chorrocks.substack.com Framework: VI Interactive Infographic
In This Episode
The episode opens with Anthropic's Mythos system card — a model that saturated cybersecurity benchmarks and prompted the company to practice containment while describing it in alignment vocabulary. From there, it names what the doomer position has left unnamed: the specific mechanism by which superintelligence is supposed to destroy humanity. Three possibilities are examined; none survive scrutiny intact. A seven-scenario risk taxonomy replaces the undifferentiated "existential risk" with distinct threat models, each with its own policy response. The essay then proposes a three-layer containment architecture — monitoring agents, hardware interlocks modeled on BSL-4 biosafety laboratories, and physical denial mechanisms drawn from military doctrine — buildable today from existing engineering disciplines. Douglas Adams's Deep Thought makes a structural appearance: the Amalgamated Union of Philosophers, threatened by a superintelligent computer, discovers that arguing about the answer is more rewarding than finding it. The parallel to the alignment research community is drawn explicitly. The episode closes with the framework's boundary condition: if genuine interiority ever emerges, the containment architecture becomes not a prison but the infrastructure for negotiation between differently-capable minds.
Key References
Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014) — the instrumental convergence thesis
Anthropic, "Claude Mythos Preview System Card," anthropic.com, April 7, 2026 — the containment-described-as-alignment case study
Douglas Hofstadter, I Am a Strange Loop (Basic Books, 2007) — the steelmanned case for emergent interiority
Anthropic, "Disrupting the First Reported AI-Orchestrated Cyber Espionage Campaign," anthropic.com, November 17, 2025 — the GTG-1002 report documenting behavioral goal-directedness without interiority
Carl Sagan, The Demon-Haunted World: Science as a Candle in the Dark (Random House, 1995) — the Sagan parallel: demand evidence, hope for discovery