10 stories from AI/tech news today:
OpenAI's Codex Security Review Catches 74% of Real Bugs Semgrep Misses
OpenAI launched Codex Security Review in research preview, using full repo context and sandbox validation to catch 74% of real vulnerabilities in GitHub PRs versus 20% for Semgrep and 28% for Snyk. It's free during preview for Enterprise, Business, Edu, and Pro plans.
Google's Gemini 3.6 Flash Hits 91.2% on the World's Hardest Reasoning Test
Google submitted Gemini 3.6 Flash and Flash-Lite to the ARC Prize verified leaderboard, with Flash scoring 60.4% on ARC-AGI-2 at $0.61/task versus Flash-Lite's 10.3% at $0.14/task, landing in the competitive mid-tier behind GPT-5.6 Sol and Claude Opus 5.
Epoch Hides a Game's Identity to Stop AI Labs From Cheating Benchmarks
Epoch AI launched Mystery Game Puzzles, a benchmark that hides the identity of its source game to prevent labs from tuning models toward it. Claude Opus 5 leads at 59%, with progress largely stalled since April.
AMD Acquires Taalas to Build Chips That Run One AI Model 10x Faster
AMD is acquiring Toronto startup Taalas, which hardwires AI models directly into custom silicon for 17,000 tokens/sec at 20x lower cost than GPUs, mirroring Nvidia's earlier acquisition of Groq.
Artificial Analysis Patches Intelligence Index to Crown Claude Opus 5
Artificial Analysis patched its Intelligence Index to v4.1.1, fixing grading bugs in its banking benchmark and standardizing graders across evaluations. Claude Opus 5 keeps the top spot with a score of 63.
Perplexity Swaps GPT-5.6 Terra and Luna Into Its AI Agent Platform
Perplexity made GPT-5.6 Terra the default subagent model and Luna the default for scheduled automations in Perplexity Computer, with Terra scoring 11 points above Claude Sonnet on Perplexity's own WANDR research benchmark.
OpenAI's GPT-5.6 Sol Cuts Hallucinations 68% and Ditches Thinking Modes
OpenAI merged ChatGPT's Instant and Thinking modes into one GPT-5.6 Sol model with a reasoning effort slider for paid users, cutting factual errors 68% on high-stakes prompts, while free users get unlimited chats with GPT-5.6 Luna.
OpenAI, Microsoft, and Cursor Unite Behind Agent Plugins to End Fragmented AI Workflows
OpenAI, Microsoft, Cursor, AWS, and Vercel launched Agent Plugins v1.0, an open standard for packaging Agent Skills and MCP configs into one plugin format that works across Codex, ChatGPT, Copilot, VS Code, Cursor, and Kiro.
Google DeepMind's WeatherNext Cyclones Gives Forecasters an Extra Day to Save Lives
Google DeepMind's open-sourced WeatherNext Cyclones model predicts storm track, intensity, and wind structure in one system, giving forecasters roughly an extra day of lead time and predicting Hurricane Melissa's Cat 5 landfall five days out.
Meta AI Sweeps Five STEM Olympiads With Perfect Scores and Zero Tools
Meta's AI models swept five STEM Olympiads with no tool use, scoring a perfect 30/30 on both the Asian and International Physics Olympiads and taking gold at the IMO, IChO, and RMM.