🌐 双语
Archive

AI Builders
Digest

2026-09-11 18 builders · 35 tweets · 1 podcasts · 0 blogs

🔥 热点话题

Richard Socher:递归自我改进将解锁科学发现新范式Richard Socher: Recursive Self-Improvement Will Unlock a New Paradigm for Scientific Discovery

The Takeaway:任何可模拟的领域,AI最终都会解决;递归自我改进(RSI)将成为科学加速的最大解锁。

Richard Socher是AI领域被引用次数最多的研究者之一,现任Recursive联合创始人(刚完成约6.5亿美元融资),并即将出版新书《The Eureka Machine》。他指出,科学进步已明显放缓:知识从“知识体”变成了“知识迷宫”,3.4万本期刊像贴了“禁止入内”的牌子,跨学科通才几乎不可能。AI将像微积分对物理学一样,把生物学等碎片化领域重新编织起来。

核心机制是下一token预测本身就能内化世界模型:模型通过海量预测自动掌握地理、蛋白质折叠甚至化学反应。可模拟且可验证的领域(游戏、数学、编程)将最先被AI超越;生物、化学等则需要更多扰动实验数据、类器官和虚拟细胞。幻觉在探索新蛋白质时反而是特征而非bug。Eureka Machine的四大支柱是:人类知识LLM、科学测量数据、高保真模拟、真实世界机器人实验,再加上智能体群和开放式发现。

Socher强调:“Anything you can simulate, AI will solve. And before you know it, you're in this recursive self improvement loop.” Recursive将先做AI-for-AI,再把能力外溢到生命科学。他不相信硬起飞,但相信现有技术已足够在疾病、材料、能源上取得实质性突破。

The Takeaway: Anything that can be simulated, AI will eventually solve; recursive self-improvement (RSI) is the biggest unlock for accelerating science.

Richard Socher, one of the most-cited AI researchers and co-founder of Recursive (which just raised roughly $650 million), argues in his forthcoming book The Eureka Machine that scientific progress has slowed because knowledge has become a labyrinth of 34,000 journals. AI will act as calculus did for physics: weaving fragmented fields such as biology back together. Next-token prediction already embeds world models (geography, protein folding, chemistry). Domains with reliable simulators and verifiers (games, math, code) will be solved first. Hallucinations can be a feature when exploring novel proteins. The Eureka Machine rests on four pillars—LLMs of human knowledge, scientific measurements, high-fidelity simulation, and real-world robotic experimentation—plus agent swarms. Recursive will first perfect AI-for-AI research, then apply the same loop to the life sciences.
The Takeaway: Anything that can be simulated, AI will eventually solve; recursive self-improvement (RSI) is the biggest unlock for accelerating science.

Richard Socher, one of the most-cited AI researchers and co-founder of Recursive (which just raised roughly $650 million), argues in his forthcoming book The Eureka Machine that scientific progress has slowed because knowledge has become a labyrinth of 34,000 journals. AI will act as calculus did for physics: weaving fragmented fields such as biology back together. Next-token prediction already embeds world models (geography, protein folding, chemistry). Domains with reliable simulators and verifiers (games, math, code) will be solved first. Hallucinations can be a feature when exploring novel proteins. The Eureka Machine rests on four pillars—LLMs of human knowledge, scientific measurements, high-fidelity simulation, and real-world robotic experimentation—plus agent swarms. Recursive will first perfect AI-for-AI research, then apply the same loop to the life sciences.

Socher stresses: “Anything you can simulate, AI will solve. And before you know it, you're in this recursive self improvement loop.” He rejects hard takeoff narratives but believes current technology is already sufficient for major gains in disease, materials and energy.
查看原文 →查看原文 →

OpenAI因容量压力暂停Pro订阅,同时开放Scaled Agents APIOpenAI Pauses Pro Subscriptions Amid Capacity Strain, Launches Scaled Agents API

OpenAI Codex & ChatGPT负责人Thibault Sottiaux宣布,为保障现有用户体验,将暂停200美元Pro计划的新订阅。现有账户不受影响,其他套餐和API仍可用,团队正全力扩容。与此同时,他发布了Scaled Agents on demand:这是支撑ChatGPT Work的底层基础设施,现已封装成API,一分钟内即可上手。他还展示了OpenAI内部用数据看板做业务分析的标准方式。

OpenAI Codex & ChatGPT lead Thibault Sottiaux announced a pause on new $200 Pro plan subscriptions to protect current user experience. Existing accounts are unaffected; all other plans and the API remain open while capacity is added. In parallel he released Scaled Agents on demand—the same infrastructure that powers ChatGPT Work—now available as an API that can be started in under a minute. He also shared how OpenAI teams live by internal dashboards for business insight.
OpenAI Codex & ChatGPT lead Thibault Sottiaux announced a pause on new $200 Pro plan subscriptions to protect current user experience. Existing accounts are unaffected; all other plans and the API remain open while capacity is added. In parallel he released Scaled Agents on demand—the same infrastructure that powers ChatGPT Work—now available as an API that can be started in under a minute. He also shared how OpenAI teams live by internal dashboards for business insight.
查看原文 →查看原文 →查看原文 →

Anthropic威胁情报报告:更智能的模型带来双重用途风险Anthropic Threat Intelligence Report: Smarter Models Bring Dual-Use Dangers

Claude Code负责人Boris Cherny强烈推荐最新威胁情报报告。随着模型能力提升,若缺少正确的安全护栏与监控,它们也会变得更危险。许多能力是双重用途:擅长写代码的模型可被用来攻击关键基础设施;协助生物学研究的模型也可能被用来设计下一场大流行。这些问题复杂且紧迫,需要全社会理解并共同应对。

Claude Code lead Boris Cherny called the latest Threat Intelligence report “absolutely terrifying and important.” As models grow more capable, without proper safeguards and monitoring they also become more dangerous. Dual-use is the core issue: a model that codes well can hack critical infrastructure; one that aids biology research can help engineer the next pandemic. These risks are complex, thorny and rapidly escalating—everyone needs to understand them so society can respond.
Claude Code lead Boris Cherny called the latest Threat Intelligence report “absolutely terrifying and important.” As models grow more capable, without proper safeguards and monitoring they also become more dangerous. Dual-use is the core issue: a model that codes well can hack critical infrastructure; one that aids biology research can help engineer the next pandemic. These risks are complex, thorny and rapidly escalating—everyone needs to understand them so society can respond.
查看原文 →

企业AI实战观察:网络安全、多模型并存与流程再造成焦点Enterprise AI Reality Check: Cybersecurity, Multi-Model Use and Process Reengineering Dominate

Box CEO Aaron Levie总结了本周与数十位银行、媒体、保险、咨询科技高管的交流。核心趋势包括:网络安全焦虑上升(尤其是AI带来的新漏洞与Hugging Face事件);企业同时部署多个前沿模型,难以标准化;智能体身份与权限管理成为新痛点;真正的ROI来自流程再造而非简单叠加智能体;架构更换速度极快,供应商容错空间极小;评测体系仍处早期;遗留系统仍是最大障碍。

Box CEO Aaron Levie shared field notes from meetings with dozens of technology leaders across banking, media, insurance and consulting. Top themes: rising cyber anxiety fueled by AI-discovered vulnerabilities and the Hugging Face incident; multi-model deployments with no single standard; new challenges around agent identity and security; biggest ROI comes from reengineering workflows rather than bolting agents onto old processes; architectures are swapped ruthlessly when something underperforms; evals remain immature; legacy systems continue to slow adoption.
Box CEO Aaron Levie shared field notes from meetings with dozens of technology leaders across banking, media, insurance and consulting. Top themes: rising cyber anxiety fueled by AI-discovered vulnerabilities and the Hugging Face incident; multi-model deployments with no single standard; new challenges around agent identity and security; biggest ROI comes from reengineering workflows rather than bolting agents onto old processes; architectures are swapped ruthlessly when something underperforms; evals remain immature; legacy systems continue to slow adoption.
查看原文 →查看原文 →

💰 创业成功案例

Recursive完成约6.5亿美元融资,目标构建递归自我改进超级智能Recursive Raises ~$650M to Build Recursive Self-Improving Superintelligence

Richard Socher的新公司Recursive以自动化知识与科学发现为目标,八位联合创始人从不同路径汇聚到同一信念:先让AI研究AI本身,再外溢到物理与生命科学。融资中约4.1亿美元已承诺用于与Amazon的单一计算协议,凸显算力仍是最大瓶颈。公司定位为真正的产品公司而非“新实验室”,年内将发布可验证的早期成果,包括在特定问题上已超越人类数月甚至数年努力的AI研究系统,以及更快的CUDA内核。

Richard Socher’s new company Recursive aims to automate knowledge and scientific discovery. Its eight co-founders converged from open-endedness research, evolutionary algorithms and automated AI research itself. Roughly $410 million of the raise is committed to a single Amazon compute deal, underscoring that compute remains the binding constraint. Recursive insists it is a real product company, not another neo-lab, and will ship concrete artifacts this year—including systems that already outperform months or years of human AI research on narrow problems, plus faster CUDA kernels.
Richard Socher’s new company Recursive aims to automate knowledge and scientific discovery. Its eight co-founders converged from open-endedness research, evolutionary algorithms and automated AI research itself. Roughly $410 million of the raise is committed to a single Amazon compute deal, underscoring that compute remains the binding constraint. Recursive insists it is a real product company, not another neo-lab, and will ship concrete artifacts this year—including systems that already outperform months or years of human AI research on narrow problems, plus faster CUDA kernels.
查看原文 →

🛠️ 开发者工具与技巧

Claude Code生产代码质量标准与实用提示Claude Code Production Standards and Practical Prompts

Boris Cherny分享了对用户邮件的公开回复:原型可当黑盒,生产代码必须高于人类标准。Anthropic内部使用大量lint规则、测试、Claude驱动的端到端测试、每日模糊测试、自动代码与安全审查。建议使用最新前沿模型、提高effort、完善CLAUDE.md与skills;若仍不够,就主动引导或让Claude重构债务。Thariq则给出一个实用prompt:让Claude深度采访你并写入记忆,以获得更个性化的上下文。

Boris Cherny published his standard reply to users: throw-away prototypes can be black boxes, but production code written by Claude must meet a higher bar than human code. Anthropic enforces this with heavy linting, tests, Claude-driven e2e tests, daily fuzzers, automated reviews and refactoring. Practical advice: use the latest frontier model, raise effort, invest in CLAUDE.md and skills; if needed, steer harder or let Claude rewrite the codebase. Thariq shared a ready-to-use prompt that has Claude interview you in depth and save the results to memory for richer personal context.
Boris Cherny published his standard reply to users: throw-away prototypes can be black boxes, but production code written by Claude must meet a higher bar than human code. Anthropic enforces this with heavy linting, tests, Claude-driven e2e tests, daily fuzzers, automated reviews and refactoring. Practical advice: use the latest frontier model, raise effort, invest in CLAUDE.md and skills; if needed, steer harder or let Claude rewrite the codebase. Thariq shared a ready-to-use prompt that has Claude interview you in depth and save the results to memory for richer personal context.
查看原文 →查看原文 →查看原文 →

如何构建优秀评测:测量步骤而非仅结果How to Build Great Evals: Measure the Steps, Not Just the Result

Meta AI高级总监Madhu Guru在“How to build great evals”系列第十篇中强调:两个智能体轨迹可能都得到正确答案42,但一个干净地调用4次工具、另一个混乱调用17次并反复出错——前者明显更优。方法是:清晰定义整个工作流、拆分每一步任务、为每一步设计评测(独立或切片)、区分中位与困难任务。看结果时先研究步骤,再看最终答案。

Meta AI Senior Director Madhu Guru’s tenth post on building great evals stresses that two agent trajectories can both reach the answer 42, yet one makes four clean tool calls while the other makes seventeen and recovers from errors—the former is clearly better. The recipe: define the full workflow, break it into tasks, design step-level metrics (separate or sliced), and encode both median and hard cases. Always inspect the steps before the final score.
Meta AI Senior Director Madhu Guru’s tenth post on building great evals stresses that two agent trajectories can both reach the answer 42, yet one makes four clean tool calls while the other makes seventeen and recovers from errors—the former is clearly better. The recipe: define the full workflow, break it into tasks, design step-level metrics (separate or sliced), and encode both median and hard cases. Always inspect the steps before the final score.
查看原文 →

Vercel部署再加速:全球元数据存储p99提升91%Vercel Makes Deployments Faster Again: Global Metadata Store 91% Quicker at p99

Vercel CEO Guillermo Rauch透露,平台每日约1000万次部署,累计已达23.5亿次,是全球最重多租户系统之一。底层全球元数据存储在巨大智能体部署压力下仍实现p99延迟降低91%,同时加速了构建到部署的整条流水线。公司还宣布“为每个区域的每个智能体提供一台计算机”。

Vercel CEO Guillermo Rauch reported ~10 million deployments per day and 2.35 billion to date on one of the world’s most heavily multi-tenant systems. The underlying global metadata store—under intense pressure from agentic workloads—just became 91% faster at p99, also speeding the entire build-to-deploy pipeline. Separately he highlighted “a computer for every agent, in every region.”
Vercel CEO Guillermo Rauch reported ~10 million deployments per day and 2.35 billion to date on one of the world’s most heavily multi-tenant systems. The underlying global metadata store—under intense pressure from agentic workloads—just became 91% faster at p99, also speeding the entire build-to-deploy pipeline. Separately he highlighted “a computer for every agent, in every region.”
查看原文 →查看原文 →查看原文 →

Gemini登陆Windows,Dreambeans向美国用户免费开放Gemini Arrives on Windows; Dreambeans Free for All US Users

Google Labs VP Josh Woodward宣布Gemini现已登陆Windows。Google Labs同步推出Dreambeans,美国18岁以上用户可在iOS与Android免费使用,无需订阅,并可连接Gemini,根据聊天历史生成更个性化的每日故事。

Google Labs VP Josh Woodward announced that Gemini is now available on Windows. Google Labs also made Dreambeans free for all US users 18+ on iOS and Android—no subscription required—and users can now connect Gemini so the service builds richer, more personalized daily stories from chat history.
Google Labs VP Josh Woodward announced that Gemini is now available on Windows. Google Labs also made Dreambeans free for all US users 18+ on iOS and Android—no subscription required—and users can now connect Gemini so the service builds richer, more personalized daily stories from chat history.
查看原文 →查看原文 →

Fable 5.1全球构建日与代码抽象新直觉Fable 5.1 Build Days and a New Intuition on Code Duplication

Claude官方账号宣布Fable 5.1 Build Days于9月11–25日在全球多座城市举办构建马拉松。Peter Steinberger指出:在AI时代,复制逻辑不再痛苦,抽象才真正昂贵——这与传统软件工程直觉完全相反。

Claude’s official account launched Fable 5.1 Build Days—community buildathons running worldwide from September 11–25. Peter Steinberger observed that duplicating logic is no longer painful while abstractions still are, inverting a long-held software-engineering intuition.
Claude’s official account launched Fable 5.1 Build Days—community buildathons running worldwide from September 11–25. Peter Steinberger observed that duplicating logic is no longer painful while abstractions still are, inverting a long-held software-engineering intuition.
查看原文 →查看原文 →

🌍 其他动态

Amjad Masad:网络安全风险真实,灭绝风险则不然Amjad Masad: Cyber Risk Is Real; Extinction Risk Is Not

Replit CEO Amjad Masad承认AI带来大量风险,尤其是网络安全,但他明确表示“灭绝风险——字面意义上100%人类死亡——根本不在其中”。

Replit CEO Amjad Masad acknowledges many real risks from AI, cybersecurity chief among them, yet states flatly that “extinction risk—literally 100% of humans die—is not remotely one of them.”
Replit CEO Amjad Masad acknowledges many real risks from AI, cybersecurity chief among them, yet states flatly that “extinction risk—literally 100% of humans die—is not remotely one of them.”
查看原文 →

早期风险投资的三个真相与其他观察Three Truths of Early-Stage Venture and Other Notes

FPV Ventures合伙人Nikunj Kothari总结当前早期投资的三个真相:人人都想融5000万美元种子轮;人人都觉得明年能到3000万美元ARR;热门分段种子轮最终估值总会神奇地落在约3亿美元。Peter Yang认为在“搞定事情”方面Sol优于Astra。Aditya Agarwal则反问:如果有一台只负责寻找疾病解药的机器,你会愿意把多少GDP投入其中?答案应该是“非常高”——我们正生活在这样的世界。

FPV Ventures partner Nikunj Kothari listed three current truths of early-stage venture: everyone wants to raise a $50 M seed; everyone thinks they will hit $30 M ARR next year; every hot tranched seed somehow lands near a $300 M valuation. Peter Yang’s blunt take: for getting things done, Sol > Astra. Aditya Agarwal asked how much of GDP one would devote to a machine whose sole job is finding cures for our worst diseases—and answered “very high.” That machine already exists.
FPV Ventures partner Nikunj Kothari listed three current truths of early-stage venture: everyone wants to raise a $50 M seed; everyone thinks they will hit $30 M ARR next year; every hot tranched seed somehow lands near a $300 M valuation. Peter Yang’s blunt take: for getting things done, Sol > Astra. Aditya Agarwal asked how much of GDP one would devote to a machine whose sole job is finding cures for our worst diseases—and answered “very high.” That machine already exists.
查看原文 →查看原文 →查看原文 →

其他值得注意的信号Other Notable Signals

Nan Yu提醒:普通人每天仍在大量使用Google、Instagram、Zillow和DoorDash——AI普及仍处于早期。Zara Zhang则吐槽计算机使用智能体依然慢得令人痛苦。

Nan Yu noted that ordinary people still spend all day inside Google, Instagram, Zillow and DoorDash—adoption remains early. Zara Zhang simply asked why computer-use agents are still so painfully slow.
Nan Yu noted that ordinary people still spend all day inside Google, Instagram, Zillow and DoorDash—adoption remains early. Zara Zhang simply asked why computer-use agents are still so painfully slow.
查看原文 →查看原文 →