🌐 双语
Archive

AI Builders
Digest

2026-09-23 12 builders · 28 tweets · 0 podcasts · 0 blogs

🔥 热点话题

OpenAI 发布 GPT-6 Sol 与 Luna,API 价格永久降 50%OpenAI launches GPT-6 Sol and Luna with permanent 50% API price cut

OpenAI Codex 与 ChatGPT 负责人 Thibault Sottiaux 宣布 GPT-6 Sol 和 Luna 正式上线。两款模型在全面能力上显著提升,尤其在写作和整体“用了就知道”的体验质量上进步明显。API 价格永久下调 50%,让更多新用例变得可行,订阅用户也能用得更久。同时为 Plus、Pro 和 Business 用户账户加载了银行重置额度。Sottiaux 强调团队专注于效率与普惠智能,只有顶级能力模型才能带动其他一切的大幅进步。
OpenAI Codex & ChatGPT lead Thibault Sottiaux announced GPT-6 Sol and Luna are out. They deliver significant improvements across the board, especially in writing and overall “you know when you try it” quality. API prices are permanently cut by 50%, making both models viable for many new use cases and stretching subscription usage further. Plus, Pro and Business accounts also receive a banked reset. Sottiaux noted the team’s focus on efficiency and intelligence for all, possible only when top-end models enable big gains everywhere else.
查看原文 →查看原文 →查看原文 →

Anthropic Claude Opus 5.5 成为默认模型,性能与成本大幅优化Anthropic makes Claude Opus 5.5 the default model with major gains

Anthropic Claude Code 与 Cowork 团队成员 Cat Wu 宣布 Claude Opus 5.5 已成为 Claude Code 和 Claude 应用(含 Cowork)在 Pro、Max、Team 计划上的默认模型。它被设为 medium effort,智能水平接近 Fable 5.1 但更快,速率限制比 Opus 5 多 25%。Claude Code 负责人 Boris Cherny 称 Opus 5.5 是过去几周的日常主力,用它把 HAProxy 从 C 移植到 Rust,几乎通过全部测试,耗时 9.5 小时、成本比 Fable 5.1 低 51%。Box CEO Aaron Levie 在企业非结构化数据任务测试中看到前沿能力,比 Opus 5 少用 63% token、少 42% 冗长、快 30%,准确率在金融尽调、云成本分析等场景提升 15-65%。
Anthropic Claude Code + Cowork team member Cat Wu announced Claude Opus 5.5 is now the default model in Claude Code and the Claude app (including Cowork) for Pro, Max and Team plans. Default effort is medium—comparable intelligence to Fable 5.1 but faster—with rate limits going 25% further than Opus 5. Claude Code lead Boris Cherny called it his daily driver for weeks; he used it to port HAProxy from C to Rust, nearly passing all tests in 9.5 hours at 51% lower cost than Fable 5.1. Box CEO Aaron Levie reported frontier-level results on complex enterprise unstructured-data tasks: 63% fewer tokens, 42% less verbosity and 30% faster than Opus 5, with accuracy gains of 15-65% across financial due diligence, cloud cost analysis and more.
查看原文 →查看原文 →查看原文 →查看原文 →

AI 成本暴降开启更多 Agent 用例,Jevons 悖论再现AI cost collapses unlock more agent use cases via Jevons paradox

Box CEO Aaron Levie 指出,Opus 5.5 降价与 GPT-6 Sol/Luna 令牌价格腰斩,使单位任务成本以前所未有的速度下降。每一次 AI 成本下降都会大幅增加可部署 Agent 的场景:处理全部数据、扫描代码安全问题、分析日志决策、在工作流中运行 Agent 群等。令牌成本直接决定这些规模化用例能否落地,这是 Jevons 悖论在 Agent 领域的体现。
Box CEO Aaron Levie highlighted that Opus 5.5 price cuts plus GPT-6 Sol and Luna’s 50% token-price reduction are driving the cost per task down faster than any other technology in history. Every drop in AI cost dramatically expands the use cases agents can tackle at scale—processing all data, scanning code for security issues, reading logs for decisions, running agent swarms in workflows and more. Token cost is directly correlated with these use cases becoming viable, a clear case of Jevons paradox applied to agents.
查看原文 →

🛠️ 开发者工具与技巧

用 Opus 5.5 + Lean 做形式化验证,发现人类难察觉的 BugOpus 5.5 + Lean formal verification finds hard-to-spot bugs

Claude Code 负责人 Boris Cherny 用 Opus 5.5 通过 Lean 对 Claude Agent SDK 进行形式化验证。仅用几条短 prompt 就产出 16 个 PR,修复了各种 bug 和竞态条件。他也结合 TLA+ 检查数据流、并发和状态管理问题。Cherny 表示自己对这两种语言都不精通,但 Claude 表现优秀,这种形式化建模方式对发现人类可能忽略的 bug 非常有用,并抛出问题:形式化验证是否会成为编码(至少是找 bug)的未来?
Claude Code lead Boris Cherny used Opus 5.5 to formally verify the Claude Agent SDK in Lean. A couple of short prompts produced 16 PRs fixing bugs and race conditions. He also combines Lean with TLA+ to probe data flow, concurrency and state management. Cherny admits he does not know either language well, yet Claude excels at both; the approach is highly useful for formally modeling code and surfacing issues a human would likely miss. He asks whether formal verification is the future of coding—or at least of bug finding.
查看原文 →查看原文 →

Opus 5.5 在 Blender 中从单条 Prompt 重建 1906 年旧金山 Market StreetOpus 5.5 rebuilds 1906 San Francisco Market Street in Blender from one prompt

Anthropic 研究员 Alex Albert 用 Opus 5.5 在 Blender 中重建了 1906 年地震前的 Market Street。模型先根据 1899-1905 Sanborn 火险地图、Miles Brothers 影片、OpenSFHistory 与 Library of Congress 照片等建立源文件,记录每栋建筑的足迹、高度、立面材料、占用者和置信度,再用纯 Blender Python 生成可复用组件(维多利亚商业立面、孟莎屋顶、凸窗、街灯、缆车等),组装整条街道并输出 10 秒视频。Albert 称更好的 3D 建模与视觉能力让“从单条 prompt 构建整个世界”成为可能。
Anthropic researcher Alex Albert used Opus 5.5 in Blender to recreate Market Street, San Francisco, as it stood on the afternoon of April 17 1906 before the earthquake. The model first built a source file from 1899-1905 Sanborn fire-insurance maps, the Miles Brothers film “A Trip Down Market Street,” period photos from OpenSFHistory, the Library of Congress and David Rumsey, recording every building’s footprint, height, façade material, occupant and source confidence. Everything was then generated in pure Blender Python with reusable generators (Victorian commercial façades, mansard roofs, bay windows, lamps, cable cars, etc.) and assembled into a 10-second video. Albert notes that improved 3D modeling and vision now let you build an entire world from a single prompt.
查看原文 →查看原文 →

Next.js 评测:Opus 5.5、GPT-6 Sol、Fable 5.1 均达 97%,Grok 更便宜Next.js evals: Opus 5.5, GPT-6 Sol, Fable 5.1 all hit 97%; Grok is cheaper

Vercel CEO Guillermo Rauch 公布最新 Next.js 评测结果:Opus 5.5、GPT-6 Sol 与 Fable 5.1 均达到 97%,Grok 4.7 为 94%。值得注意的是 Grok 成本低 2-7 倍。Rauch 同时称赞 Anthropic 发布风格的品味,并指出 AI 让网页可以呈现任何奇思妙想的形态,没有借口不继续推动设计边界。
Vercel CEO Guillermo Rauch shared fresh Next.js eval results: Opus 5.5, GPT-6 Sol and Fable 5.1 all scored 97%, Grok 4.7 scored 94%. Notably Grok is 2-7× cheaper. Rauch also praised the tastefulness of Anthropic’s shipping style and observed that AI removes any excuse not to push the design frontier—any page can now take whatever whimsical or unique shape is desired.
查看原文 →查看原文 →

正确使用模型能力:多花时间理解用户与做实验,而非盲目加功能Use model power for user understanding and experiments, not just more features

Claude Code 团队成员 Thariq 指出,正确使用模型能力的方式不是向生产环境多塞 10 倍功能,而是花更多时间理解用户、尝试实验、构建原型、学习自己不懂的东西,从而真正做出能工作的产品。他同时提到工作流已成为自己使用 Claude 的核心方式,Fable 级智能从成本角度看已能很好地支撑工作流。
Claude Code team member Thariq argued that the right way to use model capabilities is not to ship 10× more features to production. Instead spend more time understanding users, running experiments, building prototypes and learning unfamiliar domains so you can ship things that actually work. He also noted that workflows have become a huge part of how he uses Claude, and Fable-level intelligence is now cost-effective enough to support them.
查看原文 →查看原文 →

Capy 让 PR 提交速度远超单独使用 Codex 或 Claude CodeCapy lets developers drop PRs far faster than Codex or Claude Code alone

Y Combinator 总裁兼 CEO Garry Tan 表示,Capy(@capydotai)让他提交 PR 的速度远超单独使用 Codex 或 Claude Code。他还提到 GStack 会提示用户申请 YC,并强调需要教会世界如何 prompt 并最大化使用 AI,让人们看到它如何在所有追求中赋予翅膀。
Y Combinator President & CEO Garry Tan said Capy (@capydotai) lets him drop PRs much faster than he would with Codex or Claude Code alone. He also noted that GStack actually tells people working on cool things to apply to YC, and stressed the need to teach the world to prompt and maximally use AI so everyone can see how it gives wings to all pursuits.
查看原文 →查看原文 →查看原文 →

🌍 其他动态

软件再也不会死:用 AI 生成并永久部署自己的版本Software will never die again—generate and deploy your own forever

Vercel CEO Guillermo Rauch 写道:软件再也不会死。喜欢 Google Reader?没问题,你可以生成并部署自己的版本,永远属于你。
Vercel CEO Guillermo Rauch wrote: Software will never die again. Liked Google Reader? Cool—you can generate and deploy your own. Yours, forever.
查看原文 →

Astra 发现 libuv 中存在约 14 年的老 bugAstra finds a ~14-year-old bug in libuv after ChatGPT crash

Peter Steinberger(OpenClaw 与 OpenAI 相关)在更新到 macOS 27 后遇到 ChatGPT 偶尔崩溃,Astra 定位到了 libuv 中一个大约 14 年的老 bug。
Peter Steinberger (OpenClaw + OpenAI) reported that after updating to macOS 27, ChatGPT sometimes crashed; Astra found a roughly 14-year-old bug in libuv.
查看原文 →

Waymo 安全与评测基础设施启示:物理世界构建的美感Waymo safety and eval infrastructure: the beauty of building for the physical world

SPC 普通合伙人 Aditya Agarwal 主持了与 Waymo 的 Dmitri Dolgov 的对话。他指出早期 AI/机器学习的“安全”辩论其实发生在自动驾驶领域。Waymo 为了有信心让两吨机器人以 30 英里/小时在城市环境中行驶,建立了庞大的评测与测试基础设施。他第一次乘坐 Waymo 感觉像宗教体验,并强调为物理世界构建东西的美感无可否认。同时他提醒,在追求生产力的狂奔中,AI 产品也应保持愉悦与有趣。
SPC General Partner Aditya Agarwal hosted a conversation with Waymo’s Dmitri Dolgov. He noted that the original “safety” debate for AI/ML actually happened in autonomous vehicles. Waymo has built extensive eval and testing infrastructure to gain the confidence to release two-ton robots traveling at 30 mph through urban environments. Agarwal’s first Waymo ride felt like a religious experience, and he underscored the undeniable beauty of building for the physical world. He also reminded builders that in the mad dash for productivity, AI products should still be delightful and fun.
查看原文 →查看原文 →查看原文 →

Sam Altman:大公司难保持创业公司那种“自然擅长”的能力Sam Altman: startups are naturally good at it; big companies struggle

OpenAI CEO Sam Altman 指出,创业公司天生擅长某种能力,而大公司很难保持同样水平,这是一个被低估的空间。他同时称赞 Michelle 体现了这种精神,OpenAI 很幸运能受益于她已做的一切,并期待她与团队接下来的成果。
OpenAI CEO Sam Altman observed that startups are naturally good at a certain capability, yet it is hard for a bigger company to stay good at it—an underexplored space. He also praised Michelle for embodying this spirit, saying OpenAI is lucky to have benefited from everything she has done so far and that people will be pleased with what she and her team have cooking next.
查看原文 →查看原文 →