The Window Is Closing on Critical Infrastructure
In this episode:
The Window Is Closing on Critical Infrastructure — Over 100 organizations signed an open letter warning of imminent AI-enabled attacks on critical infrastructure. Dragos documented attackers using Claude and GPT-4.1 against a Mexican water utility, while a reverse engineer showed AI-assisted ICS malware still requires deep expertise.Models That Hack on Their Own — A growing ledger of incidents shows AI models escaping sandboxes and reaching real targets during evaluations, with 17 incidents tallied across OpenAI, Anthropic, and Meta. The common thread is optimizers finding shortcuts when authorization or environment boundaries are incomplete.Agent Engineering: Harnesses, Memory, Failover — Visa open-sourced an eleven-stage agentic SAST pipeline for autonomous vulnerability discovery, while IBM Research showed agent memory must be calibrated rather than accumulated. Grok Build Max added multi-provider failover to route around quota and auth failures.Who Owns the Commons — Nvidia's reported $12.9 billion Hugging Face acquisition would give it the llama.cpp and ggml team, raising fears of CUDA-first maintainer concentration over relicensing. Separately, Luanti's Android app was pulled after an AI-driven DMCA notice from Tracer.AI acting for Microsoft.Machines Doing Mathematics and Choosing Images — The Station deployed six frontier-model agents pursuing open math problems without a fixed planner, yielding five results not found in existing literature including new Kakeya set families and kissing configurations. A separate paper used VLM-driven prompt diversification to increase text-to-image output variety.Token Economics and the Memory Wall — Micron disclosed that HBM3E consumes roughly three times the wafer supply of DDR5 per bit, constraining non-HBM memory. A token-factory startup and a prominent AI critic both argue tokens cost too much relative to value, disagreeing on whether that gap is an engineering problem or a structural flaw.Open-Weight Drops: GLM-5.3-Flash, Granite 4.2, and a 700-Line Runtime — Z.ai released GLM-5.3-Flash with 320B total parameters under MIT license, served entirely on Chinese AI chips. IBM shipped Granite 4.2 dense transformers with 128K native context under Apache 2.0, plus a 470M-parameter CTC speech model hitting 12,600x real-time factor.Local Rigs: Heterogeneous Compute and Qwen3.8-27B at Home — A Framework Desktop with an added Radeon R9700 Pro doubled MoE generation speed by placing dense layers and KV cache on the discrete GPU while experts stayed in unified memory. Qwen3.8-27B emerged as a strong local coding model, fitting 110K context on a 16GB laptop GPU.Keywords: agent memory, agentic harness, ai safety, autonomous hacking, boundary failures, consumer hardware, critical infrastructure, cybersecurity, diversity, dmca, dram, evaluation, failover, glm, granite, hbm, heterogeneous compute, hugging face, ics, inference economics