Every time the agent does something dumb, there's a comforting little prayer: the next model will fix it. Ada and Boris open Chapter 3 by showing why that's backwards — AI coding has compressed three eras into three years (prompt engineering, then context engineering, then harness engineering), and a frontier model that scores 90%+ in a clean test still falls apart on real work. Not because it can't reason, but because it's a brilliant new hire dropped in with no onboarding, no tests, no guardrails. With three receipts — LangChain jumping 13.7 points by changing only the harness, Vercel going 80% to 100% by cutting fifteen tools down to two, and the APEX-Agents benchmark exposing a 24% real-world pass rate from orchestration failures — they land the pattern: once a model crosses the capability threshold, more intelligence has diminishing returns while every harness improvement compounds across every future session.
From Chapter 3 of *Harness Engineering for Vibe Coders* by J.M. Crider — go deeper at hamiltonroadagility.com/vibe