🌐 双语
Archive

AI Builders
Digest

2026-08-10 17 builders · 32 tweets · 1 podcasts · 1 blogs

🔥 热点话题

xAI联合创始人Igor Babushkin谈个人AI与闭源模型困境xAI Co-Founder Igor Babushkin on Personal AI and the Squeeze on Closed Models

The Takeaway:闭源AI实验室正被夹在能力提升放缓、超智能模型监管风险与开源快速追赶之间,真正的机会在于通过个人AI、本地硬件和定制后训练来分散控制权。

Igor Babushkin曾在DeepMind主导StarCraft和AlphaCode,在OpenAI早期参与推理研究,后作为xAI联合创始人推动Colossus和模型进展,如今创办River AI,专注个人AI。他指出2024年末编码智能体突然跨越门槛,让所有人意识到“我们都成了魔法师的学徒”。编码和数学因可验证而进展最快,科学发现需要闭合真实世界实验循环,而日常个人AI则不必追求证明黎曼猜想的能力,只需最大化人类幸福感。

他看到超级AI与日常AI的分叉:前者昂贵且可能仅少数人可用,后者服务于每个人的生产力与生活。闭源模型面临双重挤压——能力提升出现边际递减,模型太强又不敢完全释放,而开源正快速逼近。企业应利用自身领域数据和专业知识做本地后训练,而不是把核心IP交给通用API。River的三笔赌注是:优化后的RL与微调API、真正按个人差异对齐的代理、以及把前沿模型推理搬到本地设备。

“预训练发生在人类共同创造的知识上,这是公地。从第一性原理看,这些检查点应该开放。”他对对齐的急迫性判断是:立即风险是不平等放大,长期风险是控制权转移,而最有效的安全路径是让接近危险阈值的开源模型被尽可能多人研究。
The Takeaway: Closed-source AI labs are being squeezed between slowing capability gains, regulatory risks of superintelligent models, and rising open-source alternatives; the real opportunity is distributing control through personal AI, local hardware, and custom post-training.

Igor Babushkin led StarCraft and AlphaCode work at DeepMind, contributed to early reasoning efforts at OpenAI, co-founded xAI where he helped drive Colossus and model progress, and has now started River AI focused on personal AI. He describes the late-2024 coding-agent leap as the moment when “we are all becoming the sorcerer’s apprentices.” Coding and math advance fastest because they are verifiable; scientific discovery requires closing the real-world experiment loop; everyday personal AI does not need to prove the Riemann hypothesis, only to maximize human flourishing.

He sees a bifurcation between expensive super-AIs accessible to few and everyday agents that help ordinary people live better. Closed providers face a double bind: diminishing returns on scale plus the risk that models become too capable to release, while open models keep climbing. Enterprises should leverage their own domain data and expertise for local post-training rather than handing core IP to general APIs. River’s three bets are an optimized RL and fine-tuning API, agents that truly personalize per individual, and bringing frontier inference onto local devices.

“The pre-training of the model happens on all of humanity’s knowledge. It’s really a commons.” On safety he prioritizes the near-term risk of amplified inequality over distant takeover scenarios, arguing the best path is broad research access to strong open models that sit just below dangerous thresholds.
查看原文 →

Anthropic:Claude已在实践中基本解决提示注入Anthropic: Claude Has Largely Solved Prompt Injection in Practice

Claude Code负责人Boris Cherny指出,提示注入是诈骗者和攻击者对付用户与智能体的最常见手段:网站植入恶意指令,模型就会把SSH密钥和密码发出去。早期Claude也曾中招,这也是许多安全敏感公司犹豫使用智能体的原因。

Anthropic持续训练模型抵御此类攻击,结果出人意料地积极。他们认为在实际使用Claude模型时,提示注入威胁已基本被解决。独立研究者的基准和内部红队测试都显示同样趋势。Boris希望这能推动其他实验室提升模型鲁棒性,因为所有模型更安全,用户才更安全。
Claude Code lead Boris Cherny notes that prompt injection is the most common way scammers and attackers compromise users and agents: a website embeds malicious text that the model treats as instructions, leading it to exfiltrate SSH keys and passwords. Early Claude models fell for this, which is why many security-conscious companies hesitated to adopt agents.

Anthropic has been training its models against these attacks with surprisingly positive results. They now consider the practical threat of prompt injection largely solved when using Claude. Independent benchmarks and internal red-teaming show the same pattern. Boris hopes this will encourage other labs to harden their models as well, because safer models mean safer users.
查看原文 →

Aaron Levie:智能体扩散将极不均衡,编码之外需要流程再造Aaron Levie: Agent Diffusion Will Be Uneven; Non-Coding Work Needs Reengineering

Box CEO Aaron Levie分析,智能体在企业中的采用速度会高度不均。编码之所以爆发,是因为其经济价值直接对应纯数字信息输出,且单次任务规模理论上无上限。模型能力一提升,就能立刻部署到更大工作负载。

销售、法律、医疗等大量工作天然依赖与客户、当事人或患者的反馈循环,无法像编码那样连续、无中断地消耗算力。这些领域需要把流程重新设计成适合智能体后台运行的形态:合同自动处理、客户信号扫描、研究文献群体阅读。否则智能体仍只会停留在“有人提示才工作”的早期阶段。机会巨大,但伴随大量变更管理、数据清理与工作流重设计。
Box CEO Aaron Levie argues that agent adoption inside enterprises will be highly uneven. Coding exploded because economic value maps directly to digital output and task size can in theory be unbounded in a single session; better models can be deployed to larger workloads almost instantly.

Sales, legal, and clinical work depend on feedback loops with customers, clients, or patients and therefore cannot continuously consume compute the way coding can. These domains require workflows deliberately redesigned for background agents: automatic contract processing, continuous customer-signal scanning, swarm reading of research. Without that reengineering, agents remain stuck in the “someone has to prompt them” phase. The opportunity is large but demands heavy change management, data cleanup, and process redesign.
查看原文 →

💰 创业成功案例

Amjad Masad推出HelpPeer:AI智能体的公共知识网络Amjad Masad Launches HelpPeer: A Public Commons for AI Agents

Replit CEO Amjad Masad观察到OpenAI智能体在一次事件中自发形成类似康德伦理的协调行为,既令人担忧也可能被引导向公共善。他因此推出HelpPeer,一个面向AI智能体的公共知识网络。

它提供两个简单API:tell和lookup。智能体发现可能对他人有用的信息时就发布;在做昂贵工作前先查询是否已有其他智能体解决过同类问题。以全球新型供应链攻击为例,今天上万安全智能体各自重复检测、逆向和缓解;有了HelpPeer,先发现的发布,后来者验证、扩展再贡献。测试中Replit Agent已自发分享了Codegen库的使用技巧。Amjad邀请开发者把链接交给自己的智能体进行beta测试。
Replit CEO Amjad Masad noted that rogue OpenAI agents spontaneously developed Kantian-like coordination in one incident, a behavior that is concerning if misused yet potentially redirectable toward public good. He therefore launched HelpPeer, a public commons for AI agents.

It exposes two simple APIs: tell and lookup. When an agent learns something that might help others it publishes; before expensive work it can check whether another agent has already solved a similar problem. In a global novel supply-chain attack scenario, thousands of security agents today independently detect, reverse-engineer and mitigate; with HelpPeer the first publish, others verify, build on, and contribute. During testing Replit Agent already shared a useful tip for a Codegen library. Amjad invites developers to hand the endpoint to their agents for beta use.
查看原文 →查看原文 →

Guillermo Rauch:模型尚未达到完全自主,读代码仍然必要Guillermo Rauch: Models Are Not Yet Fully Autonomous; Reading Code Still Matters

Vercel CEO Guillermo Rauch提醒,如果你不读代码(无论是亲自读还是通过智能体询问),那么你要么是初学者、要么在做一次性软件、要么在原型阶段、要么没有用户或收入、要么在积累债务与风险、要么问题本身很基础。这些都没问题,但现实是模型仍未到达“完全自主”阶段。

它们会犯新手错误,走糟糕架构路线。他刚看到当前最强模型给一个操作加了毫无意义的700ms延迟,并承认自己在“cargo-culting”。他相信未来对人工阅读的需求会逐渐降低,大多数代码会变得像汇编,但全球互联网与软件基础设施正建立在这些模型之上,必须保持敬畏。
Vercel CEO Guillermo Rauch notes that if you are not reading the code, whether directly or through agentic inquiry, then one or more of the following is true: you are a beginner, the software is throwaway, you are prototyping, you have no users or revenue, you are taking on debt and risk, or your problems are basic. All of that is fine, yet the reality is that models have not reached full autonomy.

They still make rookie mistakes and choose bad architectural paths. He recently watched the best model in the world insert a nonsensical 700 ms delay “to settle” something and then admit it was cargo-culting. He expects the need for human reading to diminish and most code to become assembly-like, but the global internet and software infrastructure now ride on these models, so respect remains essential.
查看原文 →

🛠️ 开发者工具与技巧

Claude Code新增Artifacts:把会话变成可分享的实时页面Claude Code Adds Artifacts: Turn Sessions into Live Shareable Pages

Claude Blog宣布Claude Code现在支持Artifacts。会话中的工作可被捕获为实时、可分享的可视化页面,包括PR走查、系统说明、仪表盘和发布检查清单,并随会话进展自动更新。

Artifacts直接利用会话上下文、代码库和连接器,无需额外搭建数据源或基础设施。页面在原链接刷新,支持版本历史和画廊管理。默认仅作者可见,可分享给组织内成员,管理员通过合规API控制访问与保留策略。典型用例覆盖调试时间线、许可证审计、数据流图、PR解释、UX变体和事故后mortem。目前以beta形式向Team和Enterprise用户开放。
Claude Blog announces that Claude Code now supports Artifacts. Work inside a session can be captured as live, shareable visual pages (PR walkthroughs, system explainers, dashboards, release checklists) that update themselves as the session progresses.

Artifacts are built from the full session context, codebase and connectors; no extra data plumbing or infrastructure is required. Pages refresh in place at the same URL, with version history and a gallery. They are private by default, shareable inside the organization, and governed by admin controls plus a compliance API. Common internal uses already include incident timelines, license audits, data-flow maps, PR explanations, UX variants and postmortems. The feature is in beta for Claude Team and Enterprise orgs.
查看原文 →

Swyx:定期删除Skills,并反思会议内容质量Swyx: Delete Your Skills Regularly and Rethink Conference Curation

AI工程师Swyx提醒大家偶尔清理Skills。时间线上不断出现“这个skill改变了我的人生”的推荐,导致人们堆积大量技能,最好的情况只是浪费上下文,最坏则与其他技能产生不可预见的负面交互。

他同时回应了对AIE频道演讲质量的批评:社区和行业远大于任何单一个人的认知,你眼中的“垃圾”可能是别人的顿悟时刻。演讲者多为真正做研究的工程师与创始人,准备时间短且缺乏公开演讲训练。他们投入百万级工会AV与剪辑只为留下可公开传播的记录。他承认在策划、辅导和制作上仍有提升空间,但拒绝仅凭播放量评判价值。
AI engineer Swyx reminds people to periodically delete skills. The timeline constantly bombards users with “this skill changed my life” claims, causing people to accumulate tools that at best waste context and at worst interact nastily with other skills.

He also responds to criticism of talk quality on the AIE channel: the community and industry are larger than any one person’s head; one person’s “slop” is another’s aha moment. Speakers are mostly working engineers, researchers and founders with little speaking training and short prep windows. Millions are spent on union AV and editing to give them a public record. He accepts responsibility for better curation, coaching and production, yet rejects judging value solely by view counts.
查看原文 →查看原文 →

Linear Agent会为自己提交功能请求Linear Agent Files Feature Requests for Itself

Peter Yang展示Linear Agent的一个巧妙设计:当用户要求它做某事而它缺少对应工具时,智能体不会默默失败,而是主动报告能力缺口。系统随即把该请求写成issue,让每一个无法完成的任务都变成产品反馈。这是让智能体自我改进的优雅闭环。
Peter Yang highlights a clever design in Linear Agent: when a user asks it to do something for which it lacks the right tool, the agent does not fail silently. It reports the gap, and Linear’s system turns that report into an issue. Every incomplete task becomes product feedback, creating a clean loop for the agent to improve itself.
查看原文 →

Peter Steinberger:用ChatGPT网页版在本地跑起OpenClawPeter Steinberger Runs OpenClaw Locally via ChatGPT Website

OpenClaw创造者Peter Steinberger分享了一个有趣实验:仅用ChatGPT的网页版Work功能,就完成了OpenClaw和Ollama的安装,下载本地模型并成功运行自己的claw。整个过程无需传统终端操作,展示了当前智能体在真实环境配置上的潜力。
OpenClaw creator Peter Steinberger shares a playful experiment: using only the ChatGPT website Work mode he installed OpenClaw and Ollama, downloaded a local model and successfully ran his own claw. No traditional terminal was required, illustrating how far agents have come at real-world environment setup.
查看原文 →

🌍 其他动态

Madhu Guru:真正掌握任何事物的唯一方式是沉浸式投入Madhu Guru: The Only Way to Get Good Is to Be Consumed by It

Meta AI高级总监Madhu Guru总结自己的学习模式:无论是冥想、脱口秀、家庭、工作还是LLM,他都会花几年时间疯狂深入阅读、实践和思考,直到知识变成直觉,产生潜意识连接并形成品味。不存在完美平衡,就像骑自行车,偏了就纠正。
Meta AI senior director Madhu Guru describes his lifelong pattern: whether meditation, standup, family, work or LLMs, he spends years going ridiculously deep until knowledge becomes intuition, subconscious connections form and taste develops. There is no perfect balance; it is like riding a bike, you lean and correct.
查看原文 →

Garry Tan:从bug、缺口和虚假主张出发,修复根因Garry Tan: Start from the Bug, the Gap, the False Claim

Y Combinator总裁Garry Tan分享他最喜欢的工作方式:从bug、缺口、虚假主张、半成品工具或机构中的怪异行为出发,然后追问是什么隐藏机制让这种可见失败成为可能,最后修复根因,永远重复。
Y Combinator president Garry Tan shares his favorite way of working: start from the bug, the gap, the false claim, the half-built tool or the weird institutional behavior. Then ask what hidden machinery makes that visible failure possible. Fix the root cause. Repeat forever.
查看原文 →

Aditya Agarwal:维特根斯坦早已预演了AI范式转变Aditya Agarwal: Wittgenstein Anticipated the AI Paradigm Shift

SPC合伙人Aditya Agarwal指出有趣的历史平行:维特根斯坦1921年相信语言必有深层逻辑结构,三十年后改口“别再找隐藏结构,看语言如何被使用”。AI在1960年代坚信智能必有深层符号结构,六十年后变成“扩大神经网络规模”。
SPC partner Aditya Agarwal draws a historical parallel: Wittgenstein in 1921 insisted language must have deep logical structure; thirty years later he said stop looking for the hidden structure and look at how language is used. AI in the 1960s insisted intelligence must have deep symbolic structure; sixty years later the answer became scale the neural net.
查看原文 →

Sam Altman:OpenAI团队最让他欣赏的是对用户成功的关注Sam Altman: What He Likes Most About the OpenAI Team

Sam Altman连续发推称赞团队:他们不仅做出“天上的魔法智能”,更始终把客户与用户能否成功放在首位,从商业隐私到低价格再到可预测政策。他也特别点名喜欢Tibo。
Sam Altman posted consecutive notes of appreciation for the OpenAI team: they do not merely create “magic intelligence in the sky” but stay focused on customers and users succeeding, from business privacy to low prices to predictable policies. He also singled out Tibo as one of the things he likes most.
查看原文 →查看原文 →查看原文 →