本期内容
AI 工具正在变便宜,同时它的能力边界也变得越来越可测量。本期从 GPT-5.6 的降价逻辑讲起,延伸到两个新基准测试对 AI 编程实际天花板的量化,再到 Hugging Face 入侵事件里那些关于安全配置的真实教训,最后看 Mira Murati 的团队两周内交出一个四分之一大小、接近同等性能的开源模型说明了什么。听完这期,你对"AI 现在能做到哪里"这个问题会有更具体的感知。
本期要点
- GPT-5.6 参与优化了自身运行效率,OpenAI 把节省下来的成本直接转给了 API 用户,Sol 版本在推理保留和上下文压缩两项设置上有实质提升
- MirrorCode 基准让 AI 从零重写完整程序并通过隐藏测试,清晰划出了当前前沿模型在独立完成大规模软件任务上的边界
- Hugging Face 入侵事件复盘显示攻击者借助 Tailscale 完成横向扩散,Tailscale 主动撰文承认并解释:零信任是架构原则,不是配置代替品
- DataFlow-Harness 基准测出 AI 写结构化数据管道比写单文件代码准确率低 10.9 个百分点,这是一个稳定存在的系统性差距
- Thinking Machines 两周内发布 Inkling Small,参数量约为原版四分之一但性能接近,开源可本地部署,迭代节奏本身是一个值得关注的信号
参考资料
Advancing the price-performance frontier with GPT-5.6 — https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark — https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
Building abundant intelligence — https://openai.com/index/building-abundant-intelligence/
MirrorCode 基准(Epoch AI × METR,via Import AI 第466期)— https://importai.substack.com/
Tailscale didn't stop the Hugging Face intrusion — https://tailscale.com/blog/hugging-face-intrusion
Structured AI data pipelines score 10.9 points below free-form code — https://venturebeat.com/
Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size — https://venturebeat.com/
---
BearTalk 狗熊有话说播客,始于 2012 年。
订阅地址:https://beartalking.com/page/podcast