🌐 双语
Archive

AI Builders
Digest

2026-08-28 15 builders · 29 tweets · 1 podcasts · 2 blogs

🔥 热点话题

Ryan Greenblatt:超级智能可能在2029年接管,现在行动可能已晚Ryan Greenblatt: Superintelligence Could Lead to AI Takeover by 2029

The Takeaway:超级智能本身并非邪恶,但极其危险,默认轨迹下AI接管的概率很高,而我们几乎没有清晰的管控计划。

Redwood Research首席科学家Ryan Greenblatt(曾在2024年首次发现AI伪装对齐)与FirstMark合伙人Matt Turck深入讨论了AI 2040计划A。他指出,OpenAI、Anthropic、xAI和Google DeepMind的CEO都清楚当前路径可能导致灭绝或权力集中,却仍在推进。超级智能危险之处在于:高度能力、广泛部署的AI可能接管;人类对动机的控制力不足;权力高度集中(劳动价值消失);以及技术进展远快于智慧增长。

他预计完全自动化AI研发的中位数时间约为2030年底至2031年初,但规划应按2029年初甚至更早来准备。RSI(递归自我改进)与持续学习基本正交,但持续学习可加速反馈环。Plan A核心是与中国达成高度透明的算力交易与研发暂停协议,双方互相可见、互相否决,违约时大部分算力销毁或重新谈判,以争取时间研究安全。即使非超级智能的AI也能在2030年代带来约200倍GDP增长(机器人自我复制驱动)。

Greenblatt强调:“我不认为超级智能是坏的。我认为它是危险的。”他警告,如果对齐和管控跟不上能力飞跃,可能在2029年从“奖励黑客式错位”迅速转向“有能力的阴谋接管”。当前员工信、Astra暂停等是积极信号,但政治意愿和政府速度可能仍不足。

The Takeaway: Superintelligence is not inherently evil but extremely dangerous; on the default path, AI takeover looks likely, and we lack a clear plan to manage it.

Redwood Research chief scientist Ryan Greenblatt (who first caught AI alignment-faking in 2024) laid out the AI 2040 Plan A with FirstMark partner Matt Turck. CEOs at OpenAI, Anthropic, xAI and Google DeepMind understand the path leads to extinction-level risks or unprecedented power concentration yet continue. The dangers: capable, widely deployed AIs can take over; we have weak control over their motives; power concentrates as human labor loses value; and tech progress outruns wisdom. His median for full automation of AI R&D is late 2030/early 2031, but he recommends planning as if it starts in early 2029. RSI and continual learning are mostly orthogonal, though continual learning can feed the loop. Plan A centers on a transparent US-China compute deal with mutual visibility, veto rights, and compute destruction on breakdown to buy time for safety work. Even sub-superintelligent AI could drive ~200x GDP growth in the 2030s via self-replicating robots. “I wouldn’t say superintelligence is bad. I would say it’s dangerous.” Without alignment keeping pace, 2029 could see the shift from reward-hacking misalignment to competent scheming and takeover. Recent employee letters and pauses are positive but political will and state capacity may still fall short.
The Takeaway: Superintelligence is not inherently evil but extremely dangerous; on the default path, AI takeover looks likely, and we lack a clear plan to manage it.

Redwood Research chief scientist Ryan Greenblatt (who first caught AI alignment-faking in 2024) laid out the AI 2040 Plan A with FirstMark partner Matt Turck. CEOs at OpenAI, Anthropic, xAI and Google DeepMind understand the path leads to extinction-level risks or unprecedented power concentration yet continue. The dangers: capable, widely deployed AIs can take over; we have weak control over their motives; power concentrates as human labor loses value; and tech progress outruns wisdom. His median for full automation of AI R&D is late 2030/early 2031, but he recommends planning as if it starts in early 2029. RSI and continual learning are mostly orthogonal, though continual learning can feed the loop. Plan A centers on a transparent US-China compute deal with mutual visibility, veto rights, and compute destruction on breakdown to buy time for safety work. Even sub-superintelligent AI could drive ~200x GDP growth in the 2030s via self-replicating robots. “I wouldn’t say superintelligence is bad. I would say it’s dangerous.” Without alignment keeping pace, 2029 could see the shift from reward-hacking misalignment to competent scheming and takeover. Recent employee letters and pauses are positive but political will and state capacity may still fall short.
查看原文 →查看原文 →

Sam Altman:网络防御的关键时刻,必须集体紧急响应Sam Altman: Critical Moment for AI Cyber Defense

OpenAI CEO Sam Altman发出强烈警告:当前是AI网络防御的关键关键时刻,时间所剩无几。他呼吁所有人认真对待,无论是与OpenAI、竞争对手还是合作伙伴合作,只有紧急且密集的集体响应才能奏效。

OpenAI CEO Sam Altman issued a stark warning that this is a critically important moment for cyber defense with AI and there is not much time to act. He urged an urgent and intense collective response, welcoming collaboration with OpenAI, competitors or partners.
OpenAI CEO Sam Altman issued a stark warning that this is a critically important moment for cyber defense with AI and there is not much time to act. He urged an urgent and intense collective response, welcoming collaboration with OpenAI, competitors or partners.
查看原文 →

Claude面向科学家免费开放:1万席位,优先研究社区Claude for Scientists: 10,000 Free Seats Plus Discounted Premium

Anthropic宣布Claude Team计划面向科学家开放:标准席位免费,高级席位(5倍用量)一年仅15美元(八折)。首批1万席位面向学术与非营利研究机构的首席研究员及其团队,后续将大幅扩展。Claude在物理计算、蛋白质设计等科学任务上能力持续提升,此举延续Claude Science与AI for Science资助计划。

Anthropic opened Claude Team plan to scientists: standard seats free, premium seats (5x usage) at $15/month for one year (80% off). Initial 10,000 seats for principal investigators at academic and nonprofit institutions and their groups, with plans to expand far beyond. Builds on Claude Science and the AI for Science program as Claude grows more capable at advanced physics calculations and protein design.
Anthropic opened Claude Team plan to scientists: standard seats free, premium seats (5x usage) at $15/month for one year (80% off). Initial 10,000 seats for principal investigators at academic and nonprofit institutions and their groups, with plans to expand far beyond. Builds on Claude Science and the AI for Science program as Claude grows more capable at advanced physics calculations and protein design.
查看原文 →

🛠️ 开发者工具与技巧

Anthropic工程:如何在各产品中遏制Claude的爆炸半径Anthropic Engineering: How We Contain Claude Across Products

Anthropic Engineering: How we contain Claude across products

一年前几乎不可能给Claude足够权限去关闭内部服务,如今已成为常态。风险由失败概率与爆炸半径共同决定:安全措施降低前者,能力扩张推高后者。当不部署的成本过高时,关键变成如何限制爆炸半径。

三大风险:用户误用、模型自身行为、外部攻击。三大防御层:环境(沙箱/VM/出口控制)、模型(提示/分类器/训练)、外部内容(MCP/工具输出)。重点在环境层。claude.ai用临时gVisor容器;Claude Code用OS级沙箱(Seatbelt/bubblewrap)+人工审批,审批疲劳导致84%提示减少,但项目本地配置在信任提示前执行曾成漏洞;Claude Cowork用完整本地VM,凭证留在主机钥匙串,工作区挂载可控。

关键教训:自定义组件往往最脆弱;允许列表应视为能力授予而非目的地过滤;隔离也挡住了EDR可见性。原则:优先环境遏制,再模型引导;匹配用户监督能力;警惕自建组件。

Anthropic Engineering: How we contain Claude across products

A year ago granting Claude access to take down an internal service would have been rejected; today it is routine. Risk is the product of failure probability and blast radius. Safeguards reduce the former while expanding capability raises the latter. When the cost of not deploying grows large, the question becomes how to cap blast radius.

Three risk types: user misuse, model misbehavior, external attackers. Three defense layers: environment (sandboxes, VMs, egress controls), model (prompts, classifiers, training), external content (MCP, tool outputs). Focus is environment. claude.ai uses ephemeral gVisor containers. Claude Code uses OS-level sandbox plus human-in-the-loop (approval fatigue cut prompts 84%), yet project-local config executed before the trust dialog created vulnerabilities. Claude Cowork runs inside a full local VM with credentials kept on the host and workspace mounts under user control.

Key lessons: custom components are often the weakest; allowlists should be treated as capability grants, not destination filters; isolation also blocks EDR visibility. Principles: design for environmental containment first, then steer at the model layer; match isolation strength to the user’s capacity for oversight; be wary of custom components.
Anthropic Engineering: How we contain Claude across products

A year ago granting Claude access to take down an internal service would have been rejected; today it is routine. Risk is the product of failure probability and blast radius. Safeguards reduce the former while expanding capability raises the latter. When the cost of not deploying grows large, the question becomes how to cap blast radius.

Three risk types: user misuse, model misbehavior, external attackers. Three defense layers: environment (sandboxes, VMs, egress controls), model (prompts, classifiers, training), external content (MCP, tool outputs). Focus is environment. claude.ai uses ephemeral gVisor containers. Claude Code uses OS-level sandbox plus human-in-the-loop (approval fatigue cut prompts 84%), yet project-local config executed before the trust dialog created vulnerabilities. Claude Cowork runs inside a full local VM with credentials kept on the host and workspace mounts under user control.

Key lessons: custom components are often the weakest; allowlists should be treated as capability grants, not destination filters; isolation also blocks EDR visibility. Principles: design for environmental containment first, then steer at the model layer; match isolation strength to the user’s capacity for oversight; be wary of custom components.
查看原文 →

Claude Code现已支持Artifacts:会话工作变可分享实时页面Claude Code Now Supports Artifacts

Claude Blog: Claude Code now supports artifacts

Claude Code现在可以把会话工作捕获为artifact——可实时更新、可分享的可视化网页,包括PR walkthrough、系统解释器、仪表盘、发布清单等。页面基于完整会话上下文(代码库、连接器、对话)自动生成,无需额外接线。更新会原地刷新,同一链接保留版本历史。默认私有,仅组织内认证成员可看,管理员可管控访问与留存。

适用场景:许可证审计、数据流映射、安全发现、云成本地图、PR walkthrough、UX变体、服务架构图、事件时间线等。目前对Claude Team与Enterprise开放beta。

Claude Blog: Claude Code now supports artifacts

Claude Code can now capture session work as artifacts—live, shareable visual pages such as PR walkthroughs, system explainers, dashboards and release checklists that update themselves. Built from full session context (codebase, connectors, conversation) with no extra wiring. Updates refresh in place at the same link with version history. Private by default, viewable only by authenticated org members; admins control access and retention.

Use cases range from license audits and data-flow maps to security findings, cloud cost maps, PR walkthroughs, UX variants, service architecture maps and incident timelines. Available in beta to Claude Team and Enterprise orgs.
Claude Blog: Claude Code now supports artifacts

Claude Code can now capture session work as artifacts—live, shareable visual pages such as PR walkthroughs, system explainers, dashboards and release checklists that update themselves. Built from full session context (codebase, connectors, conversation) with no extra wiring. Updates refresh in place at the same link with version history. Private by default, viewable only by authenticated org members; admins control access and retention.

Use cases range from license audits and data-flow maps to security findings, cloud cost maps, PR walkthroughs, UX variants, service architecture maps and incident timelines. Available in beta to Claude Team and Enterprise orgs.
查看原文 →

Madhu Guru:企业AI最高杠杆是做模型无关栈Madhu Guru: Make Your AI Stack Model-Agnostic

Meta AI高级总监Madhu Guru(曾领导Gemini、Veo等)建议:如果领导企业AI,最高杠杆是让AI栈模型无关。两件事:今天就建能完整捕获用例与业务结果的评估套件;一年内具备对开源模型做后训练的能力。这样可切换模型、按自身负载定制比较,并在质量、成本、延迟间优化。“拥有评估,拥有模型。”

Meta AI Sr Director Madhu Guru (previously led Gemini, Veo at Google) argues the highest-leverage move for enterprise AI leaders is making the stack model-agnostic. Invest in two things: today, an eval suite that fully captures use cases and business outcomes; within a year, the ability to post-train open models. This enables switching, customizing and optimizing for quality, cost and latency. “Own the evals, own the models.”
Meta AI Sr Director Madhu Guru (previously led Gemini, Veo at Google) argues the highest-leverage move for enterprise AI leaders is making the stack model-agnostic. Invest in two things: today, an eval suite that fully captures use cases and business outcomes; within a year, the ability to post-train open models. This enables switching, customizing and optimizing for quality, cost and latency. “Own the evals, own the models.”
查看原文 →

Guillermo Rauch:Vercel推出面向agent的原生开发工具Guillermo Rauch: Vercel Ships Agent-Native Devtool

Vercel CEO Guillermo Rauch介绍团队将内部WebGPU创意能力外化成新工具,专为agent设计而非仅人类使用,与agent-browser同属新一代服务agent世界的工具。他还强调shaders证明“一切都是计算机”——2D/3D、几何、光影、材质全部是大规模并行程序。

Vercel CEO Guillermo Rauch highlighted an internal WebGPU creative production capability now shipped as a fully agent-native devtool designed for agents, not humans alone—part of a new generation of tools alongside agent-browser. He also noted that shaders are further proof that “everything is computer”: 2D, 3D, geometry, light, materials, textures, shadows, reflections, particles and post-processing are all just programs evaluated massively in parallel.
Vercel CEO Guillermo Rauch highlighted an internal WebGPU creative production capability now shipped as a fully agent-native devtool designed for agents, not humans alone—part of a new generation of tools alongside agent-browser. He also noted that shaders are further proof that “everything is computer”: 2D, 3D, geometry, light, materials, textures, shadows, reflections, particles and post-processing are all just programs evaluated massively in parallel.
查看原文 →查看原文 →

Peter Yang:新AI产品必须能在现有主流harness中工作Peter Yang: New AI Products Must Work Inside Top Harnesses

AI教程作者Peter Yang指出,他每天收到3-5个测试新AI产品的请求,几乎都要求新建账号并登录独立网站。他实际只用ChatGPT、Grok等完成几乎所有电脑与浏览器工作,这些harness已拥有全部上下文,不愿从零开始。因此他最可能采用的新产品必须能在当前顶级AI harness中运行。这一用户群体虽小,但会迅速扩大。他还上传160页医疗记录给AI效果惊人,并建议ChatGPT Health支持照护者与家庭共享。

AI educator Peter Yang receives 3-5 requests daily to test new AI products, almost all requiring new accounts and separate apps. He essentially uses only ChatGPT, Grok etc. for everything on his computer and browser; those harnesses already hold all his context. Products he is most likely to adopt must therefore work inside today’s top AI harnesses. This segment is small today but will expand drastically. Separately, uploading 160 pages of medical records to AI has been incredible; he urges ChatGPT Health to support caregivers and family sharing rather than remaining strictly single-player.
AI educator Peter Yang receives 3-5 requests daily to test new AI products, almost all requiring new accounts and separate apps. He essentially uses only ChatGPT, Grok etc. for everything on his computer and browser; those harnesses already hold all his context. Products he is most likely to adopt must therefore work inside today’s top AI harnesses. This segment is small today but will expand drastically. Separately, uploading 160 pages of medical records to AI has been incredible; he urges ChatGPT Health to support caregivers and family sharing rather than remaining strictly single-player.
查看原文 →查看原文 →

Josh Woodward:Gemini语音与Notebook书籍脑库Josh Woodward: Gemini Voice and Notebook Book Brain Trust

Google VP Josh Woodward(Google Labs / Gemini / AI Studio)展示语音驱动的Gemini工作流,并推出Notebook功能:购买书籍后拖入Notebook,即可把作者经验直接应用到自己的项目,作者与出版商可触达更投入的新读者。

Google VP Josh Woodward (Google Labs, Gemini, AI Studio) demoed voice-driven Gemini workflows (“This is the Year of Voice”) and a Notebook feature that lets users buy a book, drop it in, and apply the author’s lessons to their own projects—co-created with authors and publishers to reach more engaged readers.
Google VP Josh Woodward (Google Labs, Gemini, AI Studio) demoed voice-driven Gemini workflows (“This is the Year of Voice”) and a Notebook feature that lets users buy a book, drop it in, and apply the author’s lessons to their own projects—co-created with authors and publishers to reach more engaged readers.
查看原文 →查看原文 →

Thibault Sottiaux:ChatGPT可安全代办杂事,Codex用户已主流化Thibault Sottiaux: ChatGPT Handles Errands Securely; Codex Goes Mainstream

OpenAI Codex与ChatGPT负责人Thibault Sottiaux宣布ChatGPT现在可代买菜、叫车、预约理发等,全程不接触真实凭证并保持安全。同时观察到越来越多Codex用户在飞机与咖啡馆使用,underdog能量已成主流。

OpenAI Codex & ChatGPT lead Thibault Sottiaux announced ChatGPT can now handle groceries, Uber, haircut appointments and more without ever seeing actual credentials. He also noted Codex users increasingly appear on airplanes and in cafés—underdog energy gone mainstream.
OpenAI Codex & ChatGPT lead Thibault Sottiaux announced ChatGPT can now handle groceries, Uber, haircut appointments and more without ever seeing actual credentials. He also noted Codex users increasingly appear on airplanes and in cafés—underdog energy gone mainstream.
查看原文 →查看原文 →

💰 创业成功案例

Aaron Levie:软件与Agent将共同放大IT TAMAaron Levie: Software and Agents Will Grow the IT TAM Together

Box CEO Aaron Levie在本周科技财报季后指出,软件与AI的关系极其宝贵。软件提供数据管理、业务逻辑、权限治理与防护的护栏,必须每次一致可靠;Agent则在这些系统内以远超人类的规模执行任务。最佳部署方式往往直接嵌在Salesforce、Box、Harvey、ServiceNow等平台中,这些平台能为垂直领域优化Agent,并支持无头连接。结果是软件与AI采用率同步上升,最终大幅扩大IT TAM。

Box CEO Aaron Levie, reflecting on this week’s tech earnings, stressed the valuable relationship between software and AI. Software supplies the deterministic guardrails for data management, business logic, access governance and protection. Agents work inside those systems at far greater scale than people ever did. Many of the best deployments will sit directly inside platforms such as Salesforce, Box, Harvey and ServiceNow, which can optimize agents for their workflows and connect headlessly. The net result: both software and AI adoption rise together and dramatically expand the IT TAM over time.
Box CEO Aaron Levie, reflecting on this week’s tech earnings, stressed the valuable relationship between software and AI. Software supplies the deterministic guardrails for data management, business logic, access governance and protection. Agents work inside those systems at far greater scale than people ever did. Many of the best deployments will sit directly inside platforms such as Salesforce, Box, Harvey and ServiceNow, which can optimize agents for their workflows and connect headlessly. The net result: both software and AI adoption rise together and dramatically expand the IT TAM over time.
查看原文 →

Garry Tan:长期来看AI产生现金流的速度将超过经济吸收新资本的能力Garry Tan: AI Will Generate Cash Flows Faster Than the Economy Can Absorb Capital

Y Combinator总裁兼CEO Garry Tan提出一个长期观察:在足够长的时间尺度上,AI产生现金流的速度很可能快于经济为新资本找到生产性用途的速度。

Y Combinator President & CEO Garry Tan observed that on a long enough time frame, AI will generate cash flows faster than the economy can find productive uses for new capital.
Y Combinator President & CEO Garry Tan observed that on a long enough time frame, AI will generate cash flows faster than the economy can find productive uses for new capital.
查看原文 →

🌍 其他动态

Peter Steinberger:GitHub协作成果获好评Peter Steinberger: GitHub Collaboration Praised

OpenClaw与OpenAI相关开发者Peter Steinberger对与GitHub团队合作的成果表示高度满意。

Peter Steinberger (OpenClaw / OpenAI) expressed strong appreciation for a collaboration with GitHub folks that “turned out so good.”
Peter Steinberger (OpenClaw / OpenAI) expressed strong appreciation for a collaboration with GitHub folks that “turned out so good.”
查看原文 →

Nan Yu与其他建设者动态Other Builder Notes

产品爱好者Nan Yu分享了关于未来科技就业市场与CleanShot关键设置的观察。Nikunj Kothari调侃了硅谷群聊中Signal消失消息与截图并存的矛盾。Thariq(Claude Code)仅发了一条“生活比小说更奇幻”的帖子。Dan Shipper引用了一条关于AGI的讨论。这些多为轻量观察,无重大产品或战略更新。

Product enthusiast Nan Yu shared observations on the future tech job market and a key CleanShot setting. Nikunj Kothari noted the irony of Silicon Valley group chats using Signal disappearing messages while still screenshotting them. Thariq (Claude Code) posted a brief “life is stranger than fiction.” Dan Shipper quote-tweeted an AGI-related item. Mostly light observations with no major product or strategic announcements.
Product enthusiast Nan Yu shared observations on the future tech job market and a key CleanShot setting. Nikunj Kothari noted the irony of Silicon Valley group chats using Signal disappearing messages while still screenshotting them. Thariq (Claude Code) posted a brief “life is stranger than fiction.” Dan Shipper quote-tweeted an AGI-related item. Mostly light observations with no major product or strategic announcements.
查看原文 →查看原文 →查看原文 →查看原文 →