🌐 双语
Archive

AI Builders
Digest

2026-08-03 12 builders · 26 tweets · 1 podcasts · 0 blogs

🔥 热点话题

Core Automation:用自动化实验室寻找 Transformer 替代方案Core Automation: Building the Automated Lab to Replace Transformers

The Takeaway:真正的瓶颈不是更多算力或更大模型,而是架构本身无法支持测试时持续学习,因此需要全新架构和高度自动化的研究实验室。

Core Automation 联合创始人 Jerry Tworek(前 OpenAI VP,领导过 Strawberry 和 reasoning 团队)与 Rohan Anil(前 Gemini 预训练负责人之一,Google Brain 和 Anthropic 资深研究员)认为,当前大模型在真实世界任务上仍受限于训练与部署的鸿沟。他们指出,强化学习虽然已大规模落地,但并非“从经验中学习”的终点,未来需要更高效、更接近人类多模式学习的算法。

Jerry 强调:“第一阶段是深刻欣赏 Transformer 已经把我们带到多远,然后才能专注它的弱点。”他们认为 Transformer 在计算深度上不足,主要依赖浅层结构和逐 token 的链式推理,导致推理时计算效率低下。真正需要的是能在部署中持续学习、适应新任务和新环境的系统。大实验室因季度竞争压力难以投入高风险的架构探索,因此他们选择独立创业,目标是打造“世界上最自动化的实验室”,让每个研究员以极高迭代速度试验新架构,并最终实现模型自我改进、无需人类在环的 AGI。

Rohan 补充,优化算法与架构必须端到端共同设计,当前预训练+强化学习的组合仍有数量级效率提升空间。他们把“写高性能 kernel”作为自动化内循环的关键瓶颈,并已通过竞赛证明人类+搜索能把 QR 分解加速 60 倍,而现有模型远未达到这一水平。

最终标准很实际:当实验室团队集体休假一周后回来,发现模型已经自主把他们的日常科研工作做得更好时,就知道方向对了。
The Takeaway: The real bottleneck is not more compute or larger models, but architectures that cannot learn continuously at test time; the path forward is new architectures plus a highly automated research lab.

Core Automation co-founders Jerry Tworek (former OpenAI VP who led the Strawberry and reasoning teams) and Rohan Anil (one of the four Gemini pre-training leads, previously at Google Brain and Anthropic) argue that today’s models still fail on real-world tasks because of the gap between lab training and deployment. Reinforcement learning has been scaled, yet it is not the final form of learning from experience; richer algorithms closer to how humans actually learn are needed.

Jerry’s core point: “The first step to replacing transformers is appreciating deeply how far they were able to carry us.” Transformers are computationally shallow and rely on one-token-at-a-time chain-of-thought, making inference-time scaling inefficient. What is required are systems that continuously adapt on user data and real distributions. Big labs, locked in quarterly model races, have little appetite for high-risk architectural bets, which is why the pair founded an independent lab whose mission is to become “the most automated lab in the world.”

Rohan stresses that optimization and architecture must be co-designed end-to-end; current pre-training + RL pipelines still leave orders-of-magnitude efficiency on the table. Automating high-performance kernel generation is a key inner-loop bottleneck: a human-plus-search competition already delivered a 60× speedup on QR factorization that no existing model can match.

Their practical success metric is simple: go on vacation as a team and see whether the lab produces better work while they are gone.
查看原文 →

Aaron Levie:可验证性决定了哪些工作会先被自动化Aaron Levie: Verifiability Determines Which Work Gets Automated First

Box CEO Aaron Levie 指出一个反直觉现象:数学、安全和代码这些“最难”的工作反而会最先被自动化,因为它们结果可客观验证。清晰的奖励信号让模型训练更高效,运行时也能大规模测试正确性。相比之下,法律条款、营销活动、销售话术、财务预算等领域缺乏即时验证,依赖主观判断和延迟反馈,因此即使模型能力指数级提升,应用层和流程本身仍需大量改造。他预测未来需要全新的“知识工作测试”能力,才能真正释放自动化收益。
Box CEO Aaron Levie highlights a counter-intuitive dynamic: the “hardest” domains—math, cyber, and code—are actually the first to be automated because their outputs are objectively verifiable. Clear reward signals accelerate training and allow scalable testing at inference time. In contrast, legal drafting, marketing campaigns, sales messaging, and financial planning lack instant verifiability, depend on subjective risk tolerance, and often cannot be judged correct for months. Even as model capability grows exponentially, substantial applied-AI work and process redesign will still be required, and we may need entirely new methods to “test” knowledge work.
查看原文 →

Dan Shipper:AI 带来的“代理权断裂”及其三阶段代谢Dan Shipper: The Agency Rupture Cycle with AI

Every CEO Dan Shipper 描述了当模型突然能独立完成你以前必须逐步参与的任务时,会产生一种“代理权断裂”——你不再被需要,却毫无发言权,近乎一种身份死亡。他观察到可预测的三阶段循环:1)初始断裂,只看到 AI 本身,语言变成“AI 解决了某某问题”;2)看见脚手架,开始意识到人类在两端的提示、监督和修正工作;3)代理权重建,AI 的贡献变得隐形,你重新说“我做了这件事”。能把断裂快速转化为好奇与玩耍的人,更可能在新 AI 经济中胜出。
Every CEO Dan Shipper maps the psychological cycle that follows when a model suddenly performs a task that once required your continuous involvement. He calls it an “agency rupture”—you are no longer required and had no say in the matter, a kind of identity death. The pattern is predictable: (1) initial rupture, where only the AI is visible and language becomes “AI solved X”; (2) seeing the scaffolding, the human prompting, babysitting and correction work; (3) agency reconstruction, where the model’s contribution fades into the background and you simply say “I did this.” The ability to metabolize each rupture into playfulness and curiosity is a leading indicator of who thrives in the new AI economy.
查看原文 →

💰 创业成功案例

Guillermo Rauch:Vercel 内部 Agent @v 已成公司运营中枢Guillermo Rauch: Vercel’s Internal Agent @v Powers Day-to-Day Operations

Vercel CEO Guillermo Rauch 宣布公司内部已全面依赖名为 @v 的 agent。它覆盖财务、沟通、文档、营销、工程和业务分析,按用户记忆个性化工作流,并持续自我改进。Rauch 强调这与“大模型厂商的 Slack 插件”本质不同:公司必须从源码、运行时到数据完全掌控自己的 agent,因为未来 agent 可能成为现代公司的同义词。@v 同时充当路由器,可把任务委派给子 agent,避免“每个 agent 一个域名”的混乱。
Vercel CEO Guillermo Rauch reports that the company’s internal agent @v now sits at the center of daily operations across finance, communications, docs, marketing, engineering and analytics. It maintains per-user memory and workflows and continuously improves. Rauch stresses the difference from third-party “Big AI Slack bots”: if agents become the foundation of modern companies, full control from source to runtime to data is non-negotiable. @v also acts as a router that delegates to specialized sub-agents, solving the UX problem of every team spinning up its own isolated agent.
查看原文 →查看原文 →

Amjad Masad:LLM 象棋引擎已在 Lichess 自主对战真人Amjad Masad: LLM Chess Engine Now Playing Live on Lichess

Replit CEO Amjad Masad 展示自己的 LLM 象棋引擎已登上 Lichess,当前 Elo 1253,可同时进行多盘对局并与真人和机器人对战。系统完全自主运行,观众可在网站实时观看。
Replit CEO Amjad Masad has deployed his LLM chess engine on Lichess, where it currently sits at 1253 Elo and plays multiple concurrent games against humans and bots. The system runs fully autonomously and games can be watched live on the site.
查看原文 →查看原文 →

🛠️ 开发者工具与技巧

Peter Yang:Hermes 如何用“策展人”机制清理技能与记忆中的垃圾Peter Yang: How Hermes Curator Keeps Skills and Memory Free of Slop

实用 AI 教程作者 Peter Yang 采访 Nous Research 联合创始人时,对方解释了 Hermes 避免“技能膨胀”的机制:后台任务 Hermes Curator 定期检查技能与记忆,主动删除低效或冗余部分,并因为开源,用户可自定义“什么是垃圾”的定义,从而重写清理循环。
Practical AI educator Peter Yang asked Nous Research co-founder how Hermes avoids accumulating low-quality skills. The answer: a background process called Hermes Curator continuously audits skills and memory, removing inefficiency. Because the system is open-source, users can supply their own definition of “slop” and the curator rewrites its cleanup loop accordingly.
查看原文 →

Andrej Karpathy:鹈鹕骑自行车测试现已可在浏览器直接运行Andrej Karpathy: Pelican-on-a-Bicycle Test Now Playable in the Browser

Andrej Karpathy 分享了 Simon Willison 的“鹈鹕骑自行车”测试后续,并上传了可在浏览器直接运行、可 fork 的源码,方便社区继续玩味模型对荒诞场景的理解能力。
Andrej Karpathy posted a follow-up to Simon Willison’s “pelican on a bicycle” test, uploading the source so the prompt is playable directly in the browser and easily forkable for further experimentation.
查看原文 →

🌍 其他动态

Thariq:数学领域已出现杰文斯悖论,对深度思考者的需求将上升Thariq: Jevons Paradox Already Visible in Mathematics

Anthropic Claude Code 成员 Thariq 观察到,AI 让数学更容易理解后,实际发生的数学活动反而增加,数学家有更多时间在更高抽象层次讨论。他预计对真正能深度思考数学的人的需求会上升,并指出这与国际象棋被 AI 改变后的路径高度相似。
Anthropic Claude Code engineer Thariq notes that AI has made mathematics easier to understand, yet the volume of mathematical activity has increased and experts now spend more time discussing ideas at higher levels of abstraction. Demand for people who can truly think mathematically is therefore rising—an outcome he sees as parallel to what happened in chess.
查看原文 →查看原文 →

Nikunj Kothari:早期风险投资已完全变成“氛围资本”Nikunj Kothari: Early-Stage VC Has Fully Become Vibes Capital

FPV Ventures 合伙人 Nikunj Kothari 描述当前一级市场现实:轮次结果与基本面严重脱节,热门赛道可以零实质内容融到巨资,而看似稳妥的项目却举步维艰。他认为这种“玩场上的游戏”至少还会持续 12–18 个月。公开市场同样情绪化,连万亿美元市值公司也会因模型发布出现 5% 以上日波动。长期看,建立可控、盈利、干净资本结构的公司仍会胜出,但短期必须先理解这套新规则。
FPV Ventures partner Nikunj Kothari observes that early-to-mid-stage venture has become pure “vibes capital”: rounds with little substance close easily if the sector is fashionable, while fundamentally stronger companies struggle. He expects the dynamic to persist at least 12–18 months. Public markets show the same mood swings—even trillion-dollar names move >5% on model releases. Long-term fundamentals still win, but short-term survival requires understanding the new game.
查看原文 →

Garry Tan:AI 将创造难以想象的经济增长,这是最好的白丸Garry Tan: AI-Driven Growth Is the Ultimate White Pill

Y Combinator 总裁兼 CEO Garry Tan 强调增长本身是好事,AI 将带来难以想象的经济增长,这是最强的乐观信号。他同时提醒,真正的绩效是领地(做出人们想要的产品),而不是地图(头衔或叙事)。
Y Combinator President & CEO Garry Tan argues that growth is good and that AI will unlock unimaginable economic expansion—the strongest white-pill narrative available. He also notes that true meritocracy judges the territory (did you make something people want?) rather than the map of titles and stories.
查看原文 →查看原文 →

Ryo Lu:当 App 世界消退,软件哪些部分仍会可见?Ryo Lu: What Remains Visible When the World of Apps Fades?

Cursor 设计师 Ryo Lu 回忆早期被 Rdio、Mailbox 和 Apple 产品塑造的交互直觉,并提问:当独立 App 逐渐消失,软件中哪些部分仍会保持可见,以及它们将给人怎样的触感。
Cursor designer Ryo Lu reflects on how early products such as Rdio, Mailbox and Apple taught him patterns that made software feel simple and tactile, then asks what parts of software will remain visible—and how they will feel—once the discrete-app paradigm recedes.
查看原文 →

Guillermo Rauch:掌握 + 创造力 + AI 才能达到完全不同的层次Guillermo Rauch: Mastery + Creativity + AI Hits a Different Level

Vercel CEO Guillermo Rauch 提醒:单有 AI 很酷,但掌握、创造力与 AI 结合才能达到完全不同的层次。不要让任何人劝你放弃追求卓越与工艺。
Vercel CEO Guillermo Rauch notes that AI alone is cool, but mastery combined with creativity and AI reaches an entirely different level. He urges people not to be discouraged from pursuing excellence and craft.
查看原文 →