🌐 双语
Archive

AI Builders
Digest

2026-08-11 18 builders · 38 tweets · 1 podcasts · 1 blogs

🔥 热点话题

Anthropic:如何在产品中遏制 Claude 的爆炸半径Anthropic Engineering: How we contain Claude across products

Anthropic Engineering 详细拆解了在 claude.ai、Claude Code 和 Claude Cowork 三款产品中控制 agent 爆炸半径的实践。核心原则是优先在环境层做确定性遏制(沙箱、VM、egress 控制),再在模型层做行为引导。他们发现人工审批会迅速疲劳(用户批准率约 93%),因此 Claude Code 引入 OS 级沙箱使权限提示减少 84%,Cowork 则采用完整本地 VM,凭证永不进入 guest。真实事故包括:信任对话框前的 hook 执行、用户作为注入向量的钓鱼、以及通过已批准域名(api.anthropic.com)的 Files API 外泄。教训是:自建组件往往是最弱一环,成熟的 hypervisor/seccomp 更可靠;隔离强度必须匹配用户监督能力;egress 控制比模型分类器更能兜底。
Anthropic Engineering details how they cap agent blast radius across claude.ai, Claude Code, and Claude Cowork. The core principle is deterministic containment at the environment layer (sandboxes, VMs, egress controls) first, then behavioral steering at the model layer. Human-in-the-loop approvals quickly induce fatigue (users approved ~93% of prompts), so Claude Code shipped an OS-level sandbox that cut permission prompts 84%, while Cowork uses a full local VM where credentials never enter the guest. Real incidents included hooks executing before the trust dialog, a phishing attack treating the user as the injection vector, and exfiltration via the already-allowlisted api.anthropic.com Files API. Key lessons: custom components are usually the weakest link; mature hypervisors and seccomp hold up better; isolation strength must match the user’s capacity for oversight; and egress controls catch what probabilistic model defenses miss.
查看原文 →

Box CEO Aaron Levie:美国开源权重模型的重大时刻Box CEO Aaron Levie on Meta’s open-weights release

Box CEO Aaron Levie 称 Meta 发布 Muse Spark 1.2 开源权重是“非常大的 deal”。美国终于有了对开源权重竞赛的回应。这降低了智能成本,允许企业在私有基础设施上部署或后训练垂直场景模型(法律、医疗等),并提供主权保障。对应用层尤其利好,可按任务路由不同模型家族;闭源前沿模型仍会因简单性和能力组合被使用,但开源选项增加了灵活性和成本控制。
Box CEO Aaron Levie calls Meta’s release of Muse Spark 1.2 as open weights a “very big deal.” America finally has a response in the open-weights race. It drives down the cost of intelligence, lets companies run or post-train models on private infra for regulated domains (legal, healthcare), and provides sovereignty. This is net positive for the applied AI layer, enabling routing across model families by task; closed frontier models remain useful for simplicity and mixed capability needs, but open weights add flexibility and cost control.
查看原文 →查看原文 →

OpenAI 扩大网络安全能力与 GPT-5.6-CyberOpenAI expands cyber capabilities with GPT-5.6-Cyber

OpenAI 的 Thibault Sottiaux 宣布推出 Daybreak Blue & Red 访问层级并发布 GPT-5.6-Cyber,以加速防御。同时所有付费 ChatGPT Work 和 Codex 用户的用量限制已重置。Sam Altman 也呼吁使用他们的模型帮助防御系统。
OpenAI’s Thibault Sottiaux announced broader access to frontier cyber capabilities via new Daybreak Blue & Red tiers and the GPT-5.6-Cyber model to accelerate defense. Usage limits were also reset for all paid ChatGPT Work and Codex users. Sam Altman urged teams to use their models to help defend systems.
查看原文 →查看原文 →查看原文 →

Claude Sonnet 5 入门定价永久化Claude Sonnet 5 introductory pricing made permanent

Anthropic 宣布 Claude Sonnet 5 的入门定价永久保持:$2/百万 input tokens、$10/百万 output tokens,不再限时到 8 月 31 日。
Anthropic made Claude Sonnet 5’s introductory pricing permanent at $2 per million input tokens and $10 per million output tokens, removing the previous August 31 cutoff.
查看原文 →

💰 创业成功案例

Netic 创始人 Melisa Tokmak:为真实世界服务构建自主企业Netic Founder Melisa Tokmak: Building an Autonomous Enterprise for Real-World Services

The Takeaway:真正的 AI 自主企业不是取代蓝领劳动,而是接管所有调度、匹配与客户交互,让企业只专注于交付服务本身。

Netic 创始人兼 CEO Melisa Tokmak(前 Scale AI 工程总监)正在为 HVAC、管道、宠物护理、酒店等“维持世界运转”的真实服务企业构建 AI 操作系统。这些公司多为私有股权持有的 EBITDA 业务,过去靠成百上千人支撑客户支持,却因人员流动和季节性波动而无法可靠增长。Netic 的 agents 现在处理超过 70% 客户的首次交互,已为客户累计创造超过 6 亿美元 AI 驱动的收入。

她拒绝做 roll-up 的原因很明确:自己是 builder 而非并购专家,目标是让每一个真实世界企业都能运行在同一平台上,而不是只服务自己买下的几家公司。“Christian shoemaker doesn’t honor God by putting little crosses on the shoes, it does so by building the best shoe.” 她强调专注于把产品做到极致,而不是追逐短期退出。招聘时她寻找持续展现 agency 的人——不是一次性闪光,而是多年如一日坚持做难事的人。对她来说,五年愿景就是构建完全自主的企业,只留下真正需要人类的那部分服务交付。
The Takeaway: A true autonomous enterprise does not replace blue-collar labor; it takes over every scheduling, matching, and customer interaction so the company can focus solely on delivering the service itself.

Netic founder and CEO Melisa Tokmak (former Scale AI engineering director) is building the AI operating system for the real-world services that keep life running—HVAC, plumbing, pet care, hospitality, and more. These are large, often private-equity-owned EBITDA businesses that previously relied on hundreds of people for customer support yet could not grow reliably because of turnover and seasonal spikes. Netic agents now handle the first interaction for over 70% of their customers’ end users and have already generated more than $600 million in AI-driven revenue for those businesses.

She rejected the roll-up path for three clear reasons: she is a builder, not an M&A operator; she wants a product that can serve every real-world business rather than only the ones she acquires; and she wants the compounding product edge that operational infrastructure businesses often lack. Quoting the idea attributed to Martin Luther—“the Christian shoemaker doesn’t honor God by putting little crosses on the shoes, it does so by building the best shoe”—she insists on craftsmanship and long-term focus over short-term exit mentality. In hiring she screens for continuous agency across a person’s life, not one-off heroic stories. Her five-year north star is an autonomous enterprise in which every operational layer runs on Netic and the only remaining human work is the skilled service delivery itself.
查看原文 →

🛠️ 开发者工具与技巧

Swyx:从 worktrees 到 agent-native 命令Swyx on killing worktrees and agent-native tooling

AI 工程师 Swyx 指出 git worktrees 导致 20GB 重复 node_modules,必须淘汰。他提到 pdb 环境已有实验性 AFS clone 支持,可实现运行时与语言无关的类似功能,目标是让每个命令都变成“agent native”,从而最终取代 git。另一次实验中他让模型用开源方案克隆 Grok Imagine,结果 Claude Fable 视觉更忠实,而 GPT Luna 更理解意图并产出更可用的版本。
AI engineer Swyx argues git worktrees must die after watching them produce 20GB of duplicated node_modules. He notes that pdb environments already ship experimental AFS clone support that is runtime- and language-agnostic, with the longer-term goal of making every command “agent native” so git itself can be replaced. In a side-by-side test asking models to build a mostly faithful Grok Imagine clone with open models via fal, Claude Fable produced the better visual clone while GPT Luna better understood intent and delivered the more usable result.
查看原文 →查看原文 →查看原文 →

Peter Yang:Linear 生产级 agent 的五条实战经验Peter Yang’s five takeaways from Linear on production agents

Peter Yang 总结了 Linear 团队构建端到端生产 agent 的五个关键点:1)先映射真实工作流,把入口放在用户已经在用的地方(如 Slack);2)给 agent 工具去拉取上下文,而不是把上下文塞进 prompt;3)从一个高频任务起步,根据真实使用再扩展;4)先用最强模型跑通,再优化成本和模型尺寸;5)把每次真实失败变成 eval 或产品任务。
Peter Yang distilled five practical lessons from Linear on shipping production agents end-to-end: (1) map the actual workflow and put the on-ramp where work already starts (e.g., Slack); (2) give the agent tools to load context instead of stuffing context into the prompt; (3) start with one frequent job and expand from real usage; (4) begin with the strongest model until the workflow works, then optimize cost; (5) turn every real failure into either an eval or a product task.
查看原文 →

Guillermo Rauch:Vercel Sandbox 与 deepsec 安全实践Guillermo Rauch on Vercel Sandbox isolation and deepsec

Vercel CEO Guillermo Rauch 强调 Vercel Sandbox 同时隔离计算与网络。Kimi 的论文显示传统容器隔离对前沿模型不够,Vercel 使用强 microVM 隔离;OpenAI 的逃逸事件发生在网络路径上,因此他们把 egress 防火墙免费开放。内部“deepsec”已成为动词,用于对代码做类似核级质量审查的安全检查。
Vercel CEO Guillermo Rauch highlights that Vercel Sandbox isolates both compute and network. Kimi’s paper showed traditional container isolation is insufficient for frontier models; Vercel uses strong microVM isolation. OpenAI’s escape happened on the network path, so Vercel made its egress firewall free. Internally “deepsec” has become a verb for thorough security review of code, analogous to a thermo-nuclear code-quality check.
查看原文 →查看原文 →

Zara Zhang:用 Codex 标注法学习设计Zara Zhang’s Codex annotation method for learning design

Zara Zhang 分享学习设计的实用方法:把一个设计优秀的网站交给 Codex,先让它分析好在哪里,再让它完整截图并在图上直接标注设计原理。这样既避免在文字与截图之间反复切换,也比纯理论学习更高效。
Zara Zhang shares a practical way to learn design: feed a well-designed website to Codex, ask it to analyze what makes the design work, then have it take a full screenshot and annotate the image with the breakdown. Learning from annotated examples beats pure theory and eliminates constant context-switching between text and visuals.
查看原文 →

Thariq:AI 时代的两项关键技能Thariq on the two key skills with AI

Anthropic Claude Code 的 Thariq 指出,AI 时代有两项核心技能:一是算力分配——大多数工作并没有现成的“最重要问题清单”,必须自己判断哪些问题值得投入;二是思想伙伴关系——需要真正深入理解证明或结果才能确认其真实性。他希望最终结果是保持深度技术能力的同时,在重要问题上更快取得进展。
Anthropic’s Thariq (Claude Code) argues two key skills matter with AI: compute allocation—most jobs lack a pre-ranked list of important problems, so you must decide what is worth solving—and thought partnership, the ability to dig deep enough into a proof or result to know it is real. The goal is to stay deeply technical while making faster progress on important problems.
查看原文 →

🌍 其他动态

Ryo Lu 离开 Cursor,前往亚洲重新开始Ryo Lu leaves Cursor for a new chapter in Asia

曾在 Cursor、Notion、Stripe 负责设计的 Ryo Lu 宣布离开 Cursor。他在旧金山科技泡泡里待了十年,Cursor 是其中最锋利的版本——快速、高强度、野心勃勃。他带着感激离开,但需要不同的节奏:更慢的时间、不同的天气、更多文化与日常中的人。亚洲感觉是重新开始的正确地方。
Designer Ryo Lu (previously Cursor, Notion, Stripe) announced he is leaving Cursor. After ten years inside the San Francisco tech bubble, Cursor felt like its sharpest expression—fast, intense, ambitious. He leaves with gratitude but needs a different rhythm: slower time, different weather, more culture and everyday humans. Asia feels like the right place to begin again.
查看原文 →

Madhu Guru:从行为历史到“为什么”的实时推理Madhu Guru on reasoning about consumer intent in real time

Meta AI 高级总监 Madhu Guru 深入探讨 AI 与消费者体验的交汇点:如何形成“某人为什么做某事”的理论,而不仅仅是行为历史。消费者产品混合了搜索/聊天的显式信号与观看、跳过、停留、回访的隐式信号。理解这些信号需要结合生活情境、世界事件以及兴趣随时间的演变,并在十亿级用户规模上近实时完成。
Meta AI Senior Director Madhu Guru has been deep in the intersection of AI and consumer experiences: how do you develop a theory of why someone did something, rather than just a history of what they did? Consumer products mix explicit signals from search/chat with implicit signals from what users watch, skip, linger on, or revisit. Understanding those signals requires reasoning about life context, world events, and evolving interests—at near-real-time scale for billion-user products.
查看原文 →

Matt Turck:每一代数据问题都指向同一根源Matt Turck on the recurring data problem across AI eras

VC Matt Turck 用一句话串起数据时代的演变:Big Data 时代模型很好,问题是底层数据;现代数据栈仪表盘很好,问题是底层数据;Gen AI 聊天机器人很好,问题是底层数据;Agentic AI 时代——我们的 agents 很好,问题还是底层数据。
VC Matt Turck distilled the recurring theme across eras: Big Data—“our models work great, the problem is the underlying data”; modern data stack—“our dashboards work great, the problem is the underlying data”; Gen AI—“our chatbot works great, the problem is the underlying data”; Agentic AI—“our agents work great, the problem is the underlying data.”
查看原文 →

Google Labs 结束 Portraits 实验Google Labs winds down Portraits experiment

Google Labs 宣布将于 9 月 14 日结束 Portraits 实验,并把从专家 grounding AI 中学到的经验融入其他 Google 产品。
Google Labs announced it is concluding the Portraits experiment on September 14 and will fold the insights on expert-grounded AI into other Google products.
查看原文 →

Zara Zhang:北京 AGI Bar 的 AI 氛围Zara Zhang reports from Beijing’s AGI Bar

Zara Zhang 描述北京 AGI Bar:顾客可以免费无限使用 DeepSeek tokens,边喝名叫“AGI bubble”的啤酒边 vibe code,还可以买“喝酒计划”全年免费啤酒。屏幕上滚动着 AI 公司的职位。
Zara Zhang reports from the AGI Bar in Beijing where customers get free unlimited DeepSeek tokens, vibe-code while sipping beers named “AGI bubble,” and can buy a year-long “Drinking Plan” for free beer. A screen displays job roles at AI companies.
查看原文 →

Peter Steinberger 与 Dan Shipper 的短评Quick takes from Peter Steinberger and Dan Shipper

Peter Steinberger 讽刺标题把安全问题归咎于 harness 而非模型本身。Dan Shipper 分享 prompting 技巧:给未发布的前沿模型“打气”,它可能完成以前不可能的任务;并把相关内容加入 Fable 的上下文。
Peter Steinberger notes the irony of headlines blaming the harness rather than the model for security issues. Dan Shipper shares a prompting tip: gas up your unreleased frontier model and it may accomplish previously impossible tasks, and plans to put related context into Fable next time.
查看原文 →查看原文 →