AI vs AI: An Agent Hacks McKinsey in Two Hours
In this episode:
AI vs AI: An Agent Hacks McKinsey in Two Hours — An autonomous offensive agent exploited unauthenticated API endpoints and a JSON-key SQL injection to gain full read-write access to McKinsey's AI chatbot Lilli, exposing 46.5 million chat messages and 728,000 confidential files. The breach highlights how AI platforms inherit classic web vulnerabilities when system prompts are treated as configuration rather than code.From Detection to Clarity: Securing the Software You Actually Run — Bynario analyzes compiled binaries in production to determine whether vulnerabilities are actually triggerable in a given environment, reducing scanner noise to real exposures. Skillsmith introduces Dependency Intelligence for agent skill workflows, surfacing MCP server and model dependencies at install time to prevent silent failures.The AI-Native Organization: Trail of Bits Shows the Blueprint — Trail of Bits transformed from AI-assisted to AI-native by standardizing on Claude Code, creating an AI Maturity Matrix, and building 94 plugins with 201 skills and 84 specialized agents. Bug discovery jumped from 15 to 200 per week on some engagements, with 20 percent of client-reported bugs now initially found by AI.Trust but Verify: The Review Crisis in Agentic Coding — BullshitBench v2 showed raw local models accepted 100 percent of fabricated terminology as real, while top models still exhibited dangerous domain-specific gaps and run-to-run variance of 50-80 percent pushback rates. Teams merging 40-50 AI-generated PRs weekly face a review crisis where agents checking their own work create a self-congratulation loop.On-Device AI: The Full Local Stack Arrives — On-device AI capabilities have matured into a complete local inference stack, enabling models to run entirely on consumer hardware without cloud dependencies. This shift brings privacy, latency, and cost advantages while raising questions about capability gaps compared to cloud-hosted alternatives.Open-Weight Model Craft: Upscaling, Ablation, and Billion-Dollar Bets — The open-weight model ecosystem is advancing through techniques like upscaling and ablation studies that reveal which architectural choices drive performance. Major investments signal billion-dollar bets on open models competing with proprietary alternatives across benchmarks and real-world applications.Agent Orchestration: Gas Town Goes Industrial — Agent orchestration frameworks are moving from experimental prototypes to industrial-grade infrastructure, with standardized patterns for multi-agent coordination, tool use, and workflow management. The ecosystem is converging on production-ready patterns for deploying autonomous agents at enterprise scale.Keywords: ablation studies, agent orchestration, agentic coding, ai hallucination, ai security, ai-native, autonomous hacking, binary analysis, claude code, code review, dependency management, edge computing, enterprise ai, hardware acceleration, infrastructure, local inference, mcp tools, model training, multi-agent systems, on-device ai