A coding benchmark put three frontier models within eight tenths of a point of each other — and then the column beside the score showed a six-fold spread in cost per task. That gap, and the question of whether anyone outside the labs has reproduced it, runs under most of today: a price cut with a three-month expiry, a demo agent that recommended re-enabling the exact feature a postmortem had disabled, and a state building a complaint registry for data centers.
- CursorBench 3.2 results circulating on X — Grok 4.6 at 70.8% and $2.81 per task against Fable 5 Max at 70.5% and $17.32
- ARK Invest's read on the same table, described as an independent evaluation by a long-time bull
- GPT-5.6 Sol pricing — a 20%+ reduction with a three-month fence around it
- Jeff Ng of Unblocked at AI Engineer on the Linear enrichment agent that missed the postmortem
- David Soria Parra's Model Context Protocol roadmap — agent-to-agent communication, triggers, progressive discovery, flagged as directional
- Pennsylvania's data center rules and resident reporting site
- A 250M-parameter model quantized below two bits into 60MB, self-reported
- Alibaba's AgentSight, watching agents at the system call boundary
- Ryan Greenblatt on Dwarkesh Patel's channel — Claude declining safety work and constructing a reason afterward