本期内容
GPT-5.6 家族发布,但旗舰模型 Sol 受政府审查影响暂未全量开放,模型分层策略开始成为行业标配。工程师用 Claude Code 产能翻三倍之后,产品判断成了新瓶颈,公司反而更需要会想清楚"做什么"的人。Agent 记忆成本一直是落地的隐性门槛,MRAgent 框架把 token 消耗砍了二十七倍,让更多场景重新变得可行。一个公开邀请两千人来攻击的 AI 邮件助手,零次被攻破,提供了一套值得参考的防御思路。最后,Benedict Evans 提醒我们:那些告诉你哪些职业会消失的图表基本不可靠,真正值得追踪的是具体任务在怎么变化。
本期要点
- GPT-5.6 推出 Sol、Terra、Luna 三级模型,旗舰 Sol 因政府审查暂限访问,模型分层是成本控制的开始
- Claude Code 让工程产出相当于原来三倍,瓶颈从"怎么做"移到了"做什么",产品判断能力正在升值
- MRAgent 通过动态重建记忆替代一次性检索,token 消耗降至 LangMem 的二十七分之一,Agent 落地成本大幅下降
- 两千名攻击者发出六千封邮件,无人攻破 AI 邮件助手 Fiu,安全来自多层叠加的设计决策而非单一功能
- Benedict Evans 指出 AI 职业暴露度研究普遍不可靠,应该追踪具体任务怎么变,而不是职业标签会不会消失
参考资料
Previewing GPT-5.6 Sol — https://openai.com/index/previewing-gpt-5-6-sol/
GPT-5.6 Preview System Card — https://deploymentsafety.openai.com/gpt-5-6-preview
How agents are transforming work — https://openai.com/index/how-agents-are-transforming-work/
Statement on the US government directive to suspend access to Fable 5 and Mythos 5 — https://www.anthropic.com/news/fable-mythos-access
Claude Code turned every engineer into three. Now companies need more product thinkers — https://venturebeat.com
AI agent memory: MRAgent cuts token use up to 27x — https://venturebeat.com
Hackmyclaw experiment (Fernando Irarrázaval) — https://fernandoi.cl
Predicting AI job exposure (Benedict Evans) — https://www.ben-evans.com
---
BearTalk 狗熊有话说播客,始于 2012 年。
订阅地址:https://beartalking.com/page/podcast