Agent Security -- From Prompt Injection Pentesting to Systems Foundations
In this episode:
Agent Security -- From Prompt Injection Pentesting to Systems Foundations — Bugcrowd's new guide demonstrates prompt injection attacks harvesting AWS credentials and leaking internal APIs via poisoned RAG documents. A major systematization paper identifies four fundamental challenges applying classical security to agents, while Exploitation Validator addresses hallucinations in AI vulnerability scanning.Governing Agent Actions -- The Authorization Boundary Problem — The AI Automation Ceiling concept explains how governance failures, not intelligence failures, stall agentic deployments. Faramesh's Action Authorization Boundary introduces a deterministic enforcement layer between agent reasoning and execution, achieving 2.24ms latency and zero double-executions across one million requests.The Maturing Agent Toolchain — Charlotte's MCP browser server cuts token costs by 99% through structured page decomposition. Kilntainers provides ephemeral sandboxes for agent isolation, Run-Agent offers markdown-based multi-agent orchestration with full traceability, and Manifold gives each agent its own git branch with real-time conflict detection.AI Hardware -- Hard-Coded Silicon Meets Salvaged GPUs — Taalas raised $200M to etch model weights directly into transistors, achieving major throughput gains on its 53-billion-transistor HC1 chip. Meanwhile, a team repurposed 800 retired Ethereum mining GPUs into an inference cluster processing document OCR at a fraction of API costs.The Human Side of AI-Assisted Development — Microsoft researchers document a seniority bias where AI amplifies experienced engineers but drags down juniors, proposing preceptor mentorship programs. Personal accounts of AI coding sessions reveal that bottlenecks have shifted from code generation to verification, context management, and workflow discipline.Agent Learning and Memory Systems — SkillRL distills agent experience into hierarchical skill libraries, achieving 89.9% success on ALFWorld with a 7B model beating GPT-4o by 42%. AgentDB v3 introduces cognitive containers combining vectors, LoRA adapters, episodic memory, and cryptographic witness chains in portable single-file agent state.Gemini 3.1 Pro — Google released Gemini 3.1 Pro, more than doubling reasoning performance over 3 Pro with a 77.1% score on ARC-AGI-2. The model is rolling out across AI Studio, Vertex AI, Android Studio, NotebookLM, and the Gemini app, with demos showcasing animated SVGs, live dashboards, and 3D simulations.Keywords: action boundary, agent governance, agent memory, agent tooling, agentdb, agentic systems, ai coding, arc-agi, authorization, automation ceiling, browser automation, cognitive debt, cost optimization, custom silicon, developer experience, experience replay, gemini, google, gpu repurposing, inference hardware