Anthropic put its production playbook for twelve-hour agents on the record, and the same day's research kept converging on one principle: keep the verbatim record, verify with a separate pair of eyes, and let corrections accumulate. Meanwhile, the engineers whose jobs sit under all of this are choosing between riding the wave, ducking under it, and organizing against its terms.
- Lance Martin's AI Engineer talk, Claude for Long-Horizon Tasks — the brain/hands split, the append-only session log, independent verifier loops, and the offline "dreaming" process that fixed a Pokémon agent's trapdoor problem. His claim: infrastructure, not intelligence, now limits production agents.
- The Guardian's interactive by Varsha Bansal — working engineers adapting by chasing new skills, returning to fundamentals, and pushing for collective action, in a profession that thought of itself as permanently scarce.
- Fidelity Before Structure (Tao An, Hawaii Pacific University) — verbatim conversation chunks beat extracted memory artifacts by 15.9 and 22 points in a controlled ablation; the culprit is lossy distillation, and the advice is to annotate the transcript, never replace it.
- NexForge — requirement-driven synthesis of executable agent training tasks, with self-reported Terminal-Bench numbers that claim to pass Claude Opus 4.6.
- Alipay-PIBench — a payment-integration benchmark where the advanced tier tests refund safeguards and notification idempotency; domain skills move mean pass rates about ten points.
- Vocabulary dropout against diversity collapse — a pointer for people training self-play curricula.
- Zero2Skill — robot data collection where operator corrections persist in a Corrective Memory, cutting human working time to 16% of teleoperation.
- RouterVLA — admission and routing rules for pools of robot policies, a problem coding-agent teams currently solve by hand.