🌐 双语
Archive

AI Builders
Digest

2026-09-02 16 builders · 38 tweets · 1 podcasts · 0 blogs

🔥 热点话题

Claude Fable 5.1 全面发布:更好写作、更少干预、更低价格与企业安全Claude Fable 5.1 Ships Everywhere: Better Writing, Fewer Safeguard Hits, Lower Prices, Enterprise Frontier Safeguards

Anthropic 正式推出 Claude Fable 5.1,全面可用。模型写作更自然、语气更好,团队正积极减少“Claude 腔”。生物学相关安全干预比 Fable 5 少 85%,Claude Code 用户每个会话的网络安全干预约减少 60%。企业、API 与 SDK 客户缓存读取价格降至每百万 token 0.25 美元(原 1 美元),典型 Claude Code 会话最高便宜 38%。同步推出 Enterprise Frontier Safeguards(EFS),在零数据保留基础上为代理时代增加可观测与风险缓解层,企业数据留在自有云,自动监控标记高风险模式。Claude Mythos 5.1 面向网络防御与生命科学,通过可信访问计划提供。

Claude Code 负责人 Boris Cherny、Alex Albert、Cat Wu、Thariq 等确认模型在复杂任务上表现强劲,Alex Albert 展示用代码驱动 Blender 生成房产电影级漫游视频。Box CEO Aaron Levie 在内部企业评估中看到非结构化数据任务提升 7 个百分点,金融、科技、公共部门场景分别有 17%、37%、16% 的显著改进。
Anthropic launched Claude Fable 5.1 everywhere. Writing quality and tone improved significantly as the team works to reduce “Claude-speak.” Biology safeguards now intervene 85% less often on benign requests than Fable 5; Claude Code users should see roughly 60% fewer cyber interventions per session. Cache reads for Enterprise, API, and SDK customers dropped to $0.25 per million tokens (from $1), up to 38% cheaper for a typical Claude Code session. Enterprise Frontier Safeguards (EFS) arrive as ZDR++ for the agent era: data stays in the customer’s cloud while an automated monitoring layer flags risky cross-session patterns. Claude Mythos 5.1 for cyber defenders and life scientists is available via trusted access programs.

Claude Code lead Boris Cherny, Alex Albert, Cat Wu, and Thariq highlighted strong performance on ambitious work. Alex Albert demoed code-driven Blender pipelines that design a house from a lot photo and produce a cinematic walkthrough. Box CEO Aaron Levie reported a 7-point jump on complex unstructured enterprise tasks, with double-digit gains in financial services (+17%), technology (+37%), and public sector (+16%) scenarios.
查看原文 →查看原文 →查看原文 →查看原文 →查看原文 →查看原文 →查看原文 →查看原文 →查看原文 →

Sam Altman:能力与安全必须同步推进,下一代模型即将到来Sam Altman: Capabilities and Safeguards Must Advance Together; Next Model Coming Soon

OpenAI CEO Sam Altman 表示,整个夏天团队都在冲刺安全优先级,能力与防护同步提升比以往任何时候都更重要。Astra 训练已完成一段时间,在能力与对齐上均有显著进步,但后续模型会根据需要放慢节奏以确保安全与对齐工作充分。他强调 AI 正变得极其强大,没人完全理解其后果,管理向丰富强大 AI 的过渡、优先保障安全与人类福祉,是当今世界最高优先级之一,也是 OpenAI 的最高优先级。团队在兴奋与谨慎之间长期处于张力,希望社会与技术通过迭代循环共同演进。
OpenAI CEO Sam Altman said the team has been sprinting on safety priorities all summer; advancing capabilities and safeguards together is more important than ever. Astra finished training some time ago and represents a significant step in both capability and alignment, yet subsequent models are being paced as needed to complete sufficient safety and alignment work. AI is becoming extremely capable and no one fully understands the consequences. Managing the transition to abundant, powerful AI while optimizing for safety and benefit to people is one of the highest priorities in the world and OpenAI’s highest priority. The company lives with the tension between excitement and anxiety and believes an iterative loop in which society and the technology evolve together offers the best chance of getting the transition right.
查看原文 →

减少代理“烦人”程度是被低估的 UX 机会Making Agents Less Annoying Is Untapped UX Alpha

即将加入 OpenAI 产品团队的前 Linear 产品负责人 Nan Yu 指出,新 Claude 版本把“更少烦人的写作”作为 headline 特性之一。看似微小的改进能显著降低用户因恼怒而退出的概率。他判断,对话/修辞设计将成为 UX 设计师尚未充分开发的重要机会。
Nan Yu (soon joining OpenAI product staff, previously head of product at Linear) noted that a new Claude release lists less-annoying writing as a headline feature. Small improvements in tone and rhetoric meaningfully reduce rage-quits. He sees conversation and rhetoric design as an under-explored opportunity for UX designers who want users to reach value.
查看原文 →查看原文 →

💰 创业成功案例

Peregrine:用正向部署工程让城市更安全,同时拒绝监控国家Peregrine: Forward-Deployed Engineering That Makes Cities Safer Without Building a Surveillance State

The Takeaway:真正让城市变好的 AI 不是多收集数据,而是把机构已有的敏感数据用高信任、高治理的方式连接起来,并让一线人员快速拿到可行动的答案。

Peregrine 联合创始人 Nick Noone(前 Palantir SOCOM 正向部署)与 Ben Rudolph(前联合国难民署与 Demagi)把“正向部署工程”定义为:心理上彻底拥有客户问题、以结果为唯一导向、同时保持谦卑与同理心。他们 2017-2018 年冷呼叫开圣巴勃罗警察局大门,经历二十多次拒绝后才获得第一块工位。公司刻意反转传统公共安全科技公司的商业模式——不卖传感器、不最大化数据收集,而是做“数据之上的精确层”,让客户在已有系统上获得更高精度与可审计的答案。

核心哲学是数据主权:每个机构拥有自己的数据,Peregrine 提供细粒度权限、治理与安全共享能力。实际落地案例包括佛罗里达县用代理发现连续三天特定天气模式导致沙槽与离岸流、从而解释百余次水域救援;侦探用语义理解找出针对犹太教堂的反犹威胁模式;冷案代理处理 200-300GB 证据(视频、音频、PDF),在威斯康星复现并推进关键案。内部数据集成代理已承担约 90% 的 Python notebook 工作,可运行数小时并拆分子代理。

Nick 强调:“我们几乎是在反网络效应的生意里——如何在每个机构内部保护数据神圣性与所有权,同时再谈互操作。”他们选择做“安静的专业人士”,把功劳留给客户,用极高同理心进入复杂机构,而不是用硅谷速度压过三十年制度知识。

“Technology and all of the skills and ways of delivering our tech is part of the answer, but the real answer is just getting to the outcome at all costs.”
The Takeaway: The AI that actually improves cities does not collect more data; it securely joins the sensitive data institutions already hold and gives frontline operators precise, actionable answers under strict governance.

Peregrine co-founders Nick Noone (ex-Palantir SOCOM forward-deployed) and Ben Rudolph (ex-UNHCR and Demagi) define forward-deployed engineering as psychologically owning the customer’s problem, driving to the outcome at all costs, and doing so with deep empathy rather than ego. They cold-called their way into the San Pablo Police Department in 2017-2018 after more than two dozen rejections. The company deliberately inverts the classic public-safety tech model: it does not sell sensors or maximize data collection; it sits on top of existing systems as a precision and governance layer.

Data ownership is non-negotiable—each institution owns its data; Peregrine supplies fine-grained permissioning, auditability, and controlled sharing. Real deployments include a Florida county whose agent discovered that three consecutive days of a particular weather pattern created sand channels and rip currents, explaining over 100 water rescues; detectives using semantic search to surface patterns of antisemitic threats against synagogues; and a cold-case agent that ingests 200-300 GB of mixed media and has already helped place a suspect at both the crime scene and body-recovery site in Wisconsin. Internally, long-horizon agents now write roughly 90% of integration notebooks and can run for hours with sub-agents.

Nick puts it plainly: the business is almost anti-network-effect—“how do you preserve the sanctity of the data and the ownership of the data for every individual agency.” They stay quiet professionals, credit the customer, and enter complex institutions with patience instead of Silicon Valley speed.

“Technology and all of the skills and ways of delivering our tech is part of the answer, but the real answer is just getting to the outcome at all costs.”
查看原文 →

🛠️ 开发者工具与技巧

Peter Yang:精简 AI Skills,并用 /claude-api prompt-audit 清理冗余Peter Yang: Keep AI Skills Minimal and Run /claude-api prompt-audit

实用 AI 教程作者 Peter Yang 分享自己的 Skills 管理原则:只保留大约十几个真正在用的(多数是自己写的),定期删除不用的,并尽量保持每个 skill 短小。他推荐在 Fable 5.1 上对所有 skills 运行 /claude-api prompt-audit,可发现大量冗余规则并针对新模型做精简。他也提出“单次迭代后让 AI 更新 skill 却容易过拟合”的问题,向社区求更好的防漂移方法。
Practical AI educator Peter Yang keeps only about a dozen skills (mostly his own), regularly deletes unused ones, and forces every skill to stay short. He strongly recommends running /claude-api prompt-audit on all skills with Fable 5.1; it surfaces redundancies and rules that the newer models no longer need. He also flagged the common failure mode in which asking the model to update a skill after a successful manual iteration causes overfitting and gradual drift, and asked the community for better solutions.
查看原文 →查看原文 →查看原文 →

Madhu Guru:每个公司都该构建自我改进的产品飞轮Madhu Guru: Every Company Should Build Self-Improving Product Flywheels

Meta AI 高级总监 Madhu Guru(前 Google Gemini/Veo 负责人)强调,企业若能自建后训练系统、高质量 eval 与数据飞轮,将获得巨大机会。他列出自我改进产品的核心组件:清晰的主次与护栏指标、明确的战略与路线图、历史产品决策知识库、连接内部仪表盘/API/MCP 的工具、以及理解端到端产品开发流程的 harness。Shopify ML 团队是他眼中的世界级范例。
Meta Senior Director of AI Madhu Guru (previously led Gemini and Veo at Google) argues that enterprises that build their own post-training systems, rigorous evals, and data flywheels unlock massive opportunity. Core pieces of a self-improving product include crisp primary/secondary/guardrail metrics, an articulated strategy and roadmap, a knowledge base of past product decisions, connections to internal dashboards/APIs/MCPs, and a harness that understands the end-to-end product development flow. He called out the Shopify ML team as world-class.
查看原文 →查看原文 →

Nikunj Kothari:WebMCP 让代理原生调用网站工具与 UINikunj Kothari: WebMCP Lets Agents Call Website Tools and Build Their Own Views

FPV Ventures 合伙人 Nikunj Kothari 指出 WebMCP 仍被严重低估。代理现在可以直接对网站发起工具调用,获得完整 UI/UX 支持与交互元素,甚至构建自己的视图并生成可分享链接。他展示了 El Niño 态势追踪器 demo(由 Codex 与 Railway 构建),代理可创建视图、保留人类编辑并生成跨代理分享链接。
FPV Ventures partner Nikunj Kothari says people are still sleeping on WebMCP. Agents can now receive native tool calls against a website, including full UI/UX support and interactive elements, then build their own views and produce shareable links. He shared a live El Niño situation-tracker demo (built with Codex and Railway) in which the agent creates views, preserves human edits, and generates links for other agents or humans.
查看原文 →

Vercel:Fable 5.1 上线 AI Gateway,并深化与 TanStack 合作Vercel: Fable 5.1 on AI Gateway + Deeper TanStack Partnership

Vercel CEO Guillermo Rauch 宣布 Fable 5.1 已上线 Vercel AI Gateway,同时官宣与 Tanner Linsley 及 TanStack 团队的合作支持。无论客户选择 Next.js 还是 TanStack,Vercel 都承诺提供最高质量的服务与支持。他同时强调 Fluid 统一计算平台带来的构建性能、沙箱可靠性与超长函数时长。
Vercel CEO Guillermo Rauch announced Fable 5.1 is live on the Vercel AI Gateway and confirmed deeper partnership and support for Tanner Linsley and the TanStack team. Customers choosing either Next.js or TanStack receive the same high-quality service commitment. He also highlighted how Fluid’s unified compute substrate underpins industry-leading build performance, sandbox reliability, and 30-minute (soon longer) function durations.
查看原文 →查看原文 →查看原文 →

Aaron Levie:网络安全即将垂直化,修复与分诊必须靠更多 AIAaron Levie: Cyber Is About to Go Vertical; Triage and Fix Require More AI

Box CEO Aaron Levie 观察,前沿模型在发现与利用漏洞上已变得极其强大,开源权重模型也紧随其后。企业本就被安全发现淹没,未来只会更严重。唯一出路是用更多 AI 进行分诊与自动修复(叠加人类监督)。他认为这会创造大量安全相关新岗位。
Box CEO Aaron Levie notes that frontier models are becoming extremely good at finding and exploiting vulnerabilities, with open-weight models not far behind. Enterprises are already inundated with findings; volume will only multiply. The only viable path is AI-assisted triage and automated remediation under human oversight. He expects this to create substantial new demand for security talent.
查看原文 →

🌍 其他动态

Dan Shipper:Fable 5.1 Vibe Check 已在 Every 发布Dan Shipper: Fable 5.1 Vibe Check Live on Every

Every CEO Dan Shipper 多次提醒读者去阅读 Every 上的 Fable 5.1 完整 vibe check,并感慨以前还需要手写底层代码的日子。
Every CEO Dan Shipper repeatedly pointed readers to the full Fable 5.1 vibe check published on Every and reflected on how recently developers still had to write low-level code by hand.
查看原文 →查看原文 →查看原文 →

Thibault Sottiaux 与其他简短动态Thibault Sottiaux and Other Brief Notes

OpenAI Codex & ChatGPT 负责人 Thibault Sottiaux 发推“Bullish”,并半开玩笑询问是否该推出周边商品。Zara Zhang 提醒“做真实的人,而不是人设”。Matt Turck 用“成为苹果 CEO”作为在 X 上暴涨的第一条建议,延续其一贯幽默。
OpenAI Codex & ChatGPT lead Thibault Sottiaux simply posted “Bullish” and asked whether a merch line was in order. Zara Zhang reminded followers to “Be a person, not a persona.” Matt Turck opened a growth-hacks thread with the suggestion “Become CEO of Apple.”
查看原文 →查看原文 →查看原文 →查看原文 →