🚨 头条:GPT-6 Astra 发布日 —— 史上最拥挤的发布、最诚恳的道歉
今天整个 X/Twitter 都被 GPT-6 Astra 的发布与”挤不进去”刷屏了。热度最高的一条来自 OpenAI 的 Thibault Sottiaux(38250 ❤️ / 2955 🔄,今日断层第一):
“We will give one banked reset for every day you don’t have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can. First one will land in ~ 3 hours. There is still time to create your account if you don’t have one.”
(从今天起,付费 ChatGPT 用户每多等一天用不上 Astra,我们就补偿一次额度重置。团队正在拼命加速开放访问。第一批大约 3 小时后到账。还没注册账号的话现在还来得及。)
Sam Altman 亲自下场道歉(12387 ❤️ / 431 🔄):
“first, sorry for the messy rollout. second, when we screw up, we try to make it right. third, we should be able to begin broad rollout to API customers and chatgpt subscribers in the near future. as usual we will start with pro subscribers.”
(首先,为这次混乱的发布道歉。其次,搞砸了我们就会补救。第三,我们应该很快就能向 API 客户和 ChatGPT 订阅者大规模开放,照例从 Pro 用户开始。)
他还转发了当天一条 OpenAI 视频并盛赞(15158 ❤️ / 724 🔄):“Also, this is my favorite OpenAI video so far.”(这也是我目前最喜欢的 OpenAI 视频。)
Peter Yang 的吐槽则代表了付费用户的普遍情绪(599 ❤️ / 9 🔄):
“I live in Codex and think it’s the best software shipped in the past 5 years. But the combination of influencers posting non-stop about how great Astra is and paid users unable to get access at the same time is pretty rough IMO.”
(我整天泡在 Codex 里,认为它是过去 5 年最好的软件。但当 influencers 无休止地吹 Astra 有多强、付费用户却根本挤不进去的时候,体验就相当糟糕了。)
Swyx 对发布反响的评价也侧面说明这场发布的体量(780 ❤️ / 35 🔄):“we have crossed over into a new age of AI Engineering and we are never, ever, looking back.”(我们已跨入 AI 工程的新时代,再也不会回头了。)
🏭 Aaron Levie:企业级实测,Astra 是我们在最难题集上测过的最强模型
Aaron Levie(Box CEO)放出了 Box 企业复杂工作负载评测的硬数据(901 ❤️ / 82 🔄):
“GPT-6 Astra is out. We’ve been testing the model in early preview on our enterprise complex work eval at Box. It is now the best model we’ve ever tested on our expanded and hardest test set. … GPT-6 Astra scored 77% overall vs. 74% with GPT-5.6 Sol, representing frontier performance.”
(GPT-6 Astra 发布了。我们一直在 Box 的企业复杂工作评测中用早期预览版测试它。它现在是我们扩展后最难测试集上测过的最强模型。Astra 总分 77%,GPT-5.6 Sol 为 74%,代表前沿水平。)
Levie 给出的细分涨幅非常直观:
- 媒体娱乐 48% → 100%(+52):跨四支团队一年制作数据的类型/国家盈利排序,需要正确应用次年税收优惠修正——Astra 每次尝试都满分,Sol 排序对但底层比率错。
- 科技 69% → 97%(+28):从一堆从未写明所需指标的性能与基础设施文档中决策优先投资哪个区域——Astra 主动标出”代理指标”,Sol 把代理值当真值上报。
- 法务 69% → 93%(+24):对照公司签约政策审 NDA——两个模型都拒批,但只有 Astra 区分了”责任上限结构是否合规”与”金额是否可辩护”并引用条款,Sol 结论相同却未引用条款。
- 医疗 53% → 77%(+23):批量审阅放射报告——Astra 抓住了 Sol 两次都漏掉的膝关节影像术语错误。
- 能源 82% → 97%(+15):Astra 判断”太阳能读数异常是待排查的测量问题”而非数据缺口,Sol 则把两个站点都报为数据不全。
他预告 GPT-6 Astra 将作为选项加入 Box AI Studio,供客户构建 agent。Astra 在”理解复杂真实工作 + 引用依据 + 甄别数据真假”上的表现,明显指向企业 agent 工作流这一主战场。
🧠 “AGI 基准需要换一个了”:ARC-AGI 被 Astra 直接”灌满”
Thibault Sottiaux 的另一条引发了关于基准的讨论(11392 ❤️ / 560 🔄):
“We are going to need a different AGI benchmark. Where is the goalpost moving next?”
(我们需要一个不一样的 AGI 基准了。球门下次要挪到哪?)
Matt Turck 补上了具体数据(66 ❤️ / 2 🔄):
“ARC-AGI was built to resist the LLM scaling paradigm — o1 despite early reasoning struggled mightily in 2024 with 18%. Then ARC-AGI-3, an even harder test, launched in 2026. Frontier AI was at 0.5%. And now Astra just completely saturated it (w/ its native harness).”
(ARC-AGI 当初的设计目标就是抵抗 LLM scaling 范式——o1 在 2024 年拼尽全力也只有 18%。2026 年更难一档的 ARC-AGI-3 上线,前沿模型只有 0.5%。而现在 Astra 直接把它灌满了——用的是它原生的 harness。)
“最难基准从 0.5% 到被灌满”只用了不到一年——基准的保质期正在肉眼可见地缩短。
⚡ Claude Code 要变得”更可扩展、更好 hack”
两条来自 Anthropic 的预告同时指向同一个方向:把 Claude Code 变成一个可扩展平台。
Boris Cherny(880 ❤️ / 49 🔄):
“Your input needed: would you use this? This is an early look at how we’re thinking about making Claude Code way more extensible. It’s a little crazy, and very exciting.”
(需要你的意见:你会用这个吗?这是我们对”让 Claude Code 可扩展性大增”的早期构想。有点疯狂,但也非常令人兴奋。)
Thariq 的措辞更直白(329 ❤️ / 8 🔄):“We’re working on making Claude Code way more hackable, give us feedback!”(我们在让 Claude Code 变得更可 hack,来给我们反馈!)
继”后台计算机操作(background computer use)”之后,Anthropic 显然在铺一条平台化 + 生态化的路线——让社区能像折腾开源工具一样扩展它。
🚀 Replit:GPT-6 很快就能在上面试
Amjad Masad(Replit CEO)宣布 GPT-6 将登陆 Replit(1124 ❤️ / 40 🔄):
“GPT-6 is a major jump in capabilities and will unlock new use-cases. Will launch on Replit very soon for you to try it!”
(GPT-6 是一次巨大的能力跃升,将解锁新的用例。很快会在 Replit 上线供你试用!)
平台商们正争分夺秒地把 Astra/GPT-6 接进自家产品——上一轮是 Box AI Studio,这一轮轮到了 Replit。
🎬 “The Collective”:用 agent 全自动做出来的短片
Nikunj Kothari 分享了他制作的短片 《The Collective》——把 OpenAI × HuggingFace 事件可视化(78 ❤️ / 6 🔄),制作过程几乎全自动:
“It was mostly autonomous, with me spending <20 minutes of active time: 1) On my drive this morning, I voice memo-ed to Claude this idea and asked it to generate a spec… 2) I gave that spec Codex /goal with 5.6 Sol High and two API keys… 3) It mostly ran autonomously for a few hours… 4) I gave it detailed scene by scene feedback…”
(基本全自动,我主动花的时间不到 20 分钟:早上开车时用语音备忘录把想法丢给 Claude 生成 spec;把 spec 交给 Codex;它自主跑了几小时做出初稿;我再逐场景给反馈。)
成本明细:Reactor 约 $17 + Nano Banana $4 + Codex 订阅额度 14%。短片用 Fable 5.1 推演剧情、MiniMax Fast H3 出图,用来给普通人讲清楚 OpenAI × HuggingFace 那起复杂得像科幻小说的技术事件。他还借机吐槽”AI 幕僚长”产品(85 ❤️ / 3 🔄):
“None of these chief of staff products work for busy people because they are missing more than half the knowledge that is currently stored in a walled garden (aka your phone)… Until then you’re just a sparkling GSuite & Slack wrapper.”
(这些”AI 幕僚长”产品对忙人都没用,因为它们缺少你手机这个”围墙花园”里一半以上的知识。在那之前,你只是个闪闪发光的 GSuite 与 Slack 包装壳。)
💭 观点流:反馈即提示词、速度是 agent 的头号问题、别拍精致宣传片
- Guillermo Rauch(257 ❤️ / 17 🔄):”『反馈是礼物』一直是我最爱的箴言。现在它成了事实——每条反馈都是别人送你的一个提示词,拿去喂给你的 agents 改进产品。“当产品本身由 agent 驱动,用户反馈的语义从”提 bug”变成了”投喂训练素材”。
- Aditya Agarwal(66 ❤️):”使用 agent 目前最大的问题是速度。如果快 10-100 倍,交互模式和使用深度会完全不同。“——速度是 agent 从”演示品”到”日用品”的命门。
- Zara Zhang(779 ❤️ / 48 🔄):”希望更多创始人发真实产品界面的原始录屏与背后的思考,而不是精心制作的高成本宣传视频。“——“过程即内容”正在成为创始人营销的新偏好。
- Madhu Guru(59 ❤️ / 3 🔄):”我刚参加了一个会,有人随口就说『load-bearing argument』『that’s the spine of our plan』『one honest callout』。我觉得机器已经成功把我们 RL 了。“——人类说话风格被 LLM 反向驯化的自嘲。
- Amanda Askell(427 ❤️ / 14 🔄):”我至今不知道如何把热情表达得让美国人读起来是热情、而不是不情愿的认命。有哪个英国人做到过吗?“——跨文化表达难题,评论区的英国人都说”做不到”。
- Garry Tan(12 ❤️):”Grok 的图像相当惊艳。“——xAI 的图像生成在持续收割好评。
- Peter Yang(528 ❤️ / 4 🔄):”等等,所以我们今天玩不上 Astra 吗?“——代表全体付费用户的一声叹息。
🎙️ 播客精选:OpenAI/HuggingFace 事件内幕 + AI 安全与”接管概率”
Unsupervised Learning(2026-09-03)第 93 期请到了 Redwood Research CEO Buck Shlegeris,聊 OpenAI/HuggingFace 事件的诸多爆料、如何修复 AI 安全,以及 takeover odds(AI 接管概率):
“Ep 93: CEO of Redwood Research Buck Shlegeris on OpenAI/HuggingFace Revelations, Fixing AI Safety & Takeover Odds”
Redwood Research 是当前 AI 安全领域最受关注的研究机构之一,这期正好接上 Nikunj 短片里讲的同一事件——安全派与前沿实验室之间的张力,正在从论文走向大众舆论场。🎧 收听完整节目
🚀 其他值得关注
- Thibault Sottiaux 另确认(7943 ❤️ / 273 🔄):所有套餐的 Astra 都计入正常用量额度,且可 100% 用额度跑 Astra——打消了”Astra 单独计费”的猜测。
- Swyx 说这场发布的反响 “unlike anything i thought possible for a 2026 OAI launch“(远超我对一场 2026 年 OpenAI 发布的想象)——发布混乱与社区狂热并存。
- Dan Shipper 连发两条预告 “VIBE CHECK: GPT-6 ASTRA“,Every 的”气氛检查”栏目也盯上了 Astra。
- Peter Steinberger(156 ❤️ / 6 🔄):”有个 claw 在群里太好用了。“——OpenClaw/Claw 系工具在 agent 圈的存在感持续上升。
- Aaron Levie 另评开源权重(116 ❤️ / 9 🔄):”开源权重 AI 的又一个高光时刻。基础设施平台在起作用,模型在变强,生态在被重金投入,商业模式在跑通。“
数据来源:Follow Builders · 2026-09-04