🌐 双语
Archive

AI Builders
Digest

2026-09-19 13 builders · 28 tweets · 1 podcasts · 0 blogs

🔥 热点话题

扩散模型将在AI推理中获胜:Inception CEO Stefano Ermon详解Why Diffusion Will Win AI Inference: Inception CEO Stefano Ermon

核心要点:扩散模型因为天生并行、更契合GPU推理工作负载,将在AI推理上击败自回归模型,带来更快、更高效的LLM。

斯坦福教授、扩散模型之父、Inception联合创始人兼CEO Stefano Ermon,从2014年开始研究生成模型。当时领域冷门,模型只能勉强生成模糊的MNIST数字。他始终相信生成模型是从无标签数据中学习结构的正确方式,最初更关注世界模型,而非今天LLM的能力。

他实验室在2019年与学生Yang Song提出基于去噪的score-based模型,最终演变为现代扩散模型。如今图像、视频、音乐甚至蛋白质生成的最佳模型多基于扩散。2024年,他们首次证明扩散语言模型在GPT-2规模上能匹配自回归模型的困惑度,同时生成速度快约10倍。

Inception约两年、50人左右,正在把这项技术规模化成商用扩散LLM(Mercury系列)。这些模型在基准上已与OpenAI等前沿实验室的flash/mini模型质量相当,但显著更快,并已在生产中服务真实客户。他们自建了服务引擎,因为现有LLM引擎无法直接运行扩散LLM。

Ermon强调推理时扩展才是关键:自回归模型生成时仍是串行的,严重内存受限;扩散模型在推理时能并行处理大量token,工作负载更像训练,完美映射GPU优势。他引用“苦涩的教训”:更并行的方案最终会赢。扩散模型还更易控制(可从粗到细引导),可能更数据高效。客户如语音代理OpenCall已从专用芯片转向他们的模型,在NVIDIA GPU上达到同等速度、更低成本。

“我们赌的是扩散基LLM,因为它天生更并行。苦涩的教训是,更并行的方案最终会赢。”
The Takeaway: Diffusion models will win AI inference because they are inherently parallel and map far better to GPU workloads than sequential autoregressive models, delivering higher intelligence per watt and per dollar.

Stefano Ermon, longtime Stanford professor and one of the fathers of diffusion models, is now co-founder and CEO of Inception. He has worked on generative models since 2014, when the field was unfashionable and success meant blurry MNIST digits. He always viewed generative modeling as the right way to capture structure in unlabeled data, originally thinking in terms of world models for decision-making rather than today’s LLM capabilities.

In 2019 his lab (with PhD student Yang Song) introduced score-based generative models that train a network to denoise images—the foundation of modern diffusion. Diffusion now powers the best image, video, music and many protein models. In 2024 they showed for the first time that a diffusion language model could match autoregressive perplexity at GPT-2 scale while generating text roughly 10× faster because it outputs many tokens in parallel.

Inception (≈2 years old, ~50 people) is scaling that technology into commercial diffusion LLMs. Their Mercury models already match the quality of frontier labs’ speed-optimized flash/mini models on benchmarks while being significantly faster, and they are serving real production traffic. They had to build their own serving engine because existing systems cannot run diffusion LLMs.

Ermon’s core bet is inference-time scaling. Autoregressive generation remains sequential and memory-bound; diffusion’s parallel workload at inference closely resembles training and therefore exploits GPUs efficiently. He invokes the bitter lesson: the more parallel solution eventually wins. Diffusion models are also easier to steer (coarse-to-fine control) and may prove more data-efficient. Customers such as voice-agent company OpenCall have switched from custom silicon to Mercury, achieving comparable latency on ordinary NVIDIA GPUs at lower cost and higher availability.

“We bet on a diffusion-based LLM because it’s inherently more parallel. And the bitter lesson is that the more parallel solution is the one that is eventually going to win.”
查看原文 →

Claude Code正式支持AGENTS.mdClaude Code Adds Native AGENTS.md Support

Anthropic Claude Code团队成员Thariq宣布,从版本2.1.277起,Claude Code正式支持AGENTS.md。如果文件夹中没有CLAUDE.md,Claude会自动检查并使用AGENTS.md。该功能基于即将推出的Claude Code mods(自定义harness的方式),这是一个内置mod,用户未来也能构建自己的项目指令版本。可在/config中开关此行为。源码已公开。
Thariq of Anthropic’s Claude Code team announced native AGENTS.md support starting in version 2.1.277. If no CLAUDE.md exists in a folder, Claude will check for and use AGENTS.md. The feature is built on upcoming Claude Code mods—the team’s way to customize the harness. This is a built-in mod; users will be able to create their own project-instruction variants. The behavior can be toggled in /config, and the source is public.
查看原文 →查看原文 →查看原文 →

Vercel AI Gateway开放模型token占比创纪录Record Open-Model Token Share on Vercel AI Gateway

Vercel CEO Guillermo Rauch指出,当天Vercel AI Gateway上开放模型token体积占比达到创纪录的78.4%(封闭模型仅21.6%)。虽然花费通常呈现不同故事,但当天第3、第4名是Moonshot AI和DeepSeek;加上Z.ai,三者合计花费已超过OpenAI(第2名)。他同时提到Jev的采用数据“令人震惊”,所有人都在用,这既源于产品本身优秀,也反映了当前“AI太贵/太慢”的时代情绪,人们急于优化并把AI放到更多地方。
Vercel CEO Guillermo Rauch reported a record day for open-model token volume on Vercel AI Gateway: 78.4% open vs 21.6% closed. While spend usually tells a different story, #3 and #4 were Moonshot AI and DeepSeek; adding Z.ai their combined spend surpassed OpenAI (#2). He also highlighted “shocking” adoption data for Jev—everyone is using it—driven both by product quality and the broader “AI is too expensive/slow” zeitgeist that makes people eager to optimize and deploy AI more widely.
查看原文 →查看原文 →

💰 创业成功案例

Meta Muse个人代理一年帮省800+美元Meta’s Muse Agent Saves $800+ a Year on Bills

实用AI教程作者Peter Yang称Meta的Muse是他用过最好的个人代理。它实际帮他在有线电视和手机账单上一年节省超过800美元,作为免费AI代理价值惊人。他发布了新视频,演示10个最爱用例,包括个性化早间新闻、习惯追踪,以及让Muse直接打电话给客服谈判账单。他甚至分享了谈判Comcast账单节省288美元的通话记录,认为大多数公司客服线还没准备好应对代理。
Practical AI educator Peter Yang calls Meta’s Muse the best personal agent he has tried. It actually saved him $800+ a year on cable and phone bills—insane value for a free AI agent. He released a new video walking through 10 favorite use cases, including a personalized morning news feed, habit tracker, and having Muse call customer support to negotiate bills. He also shared the phone transcript of Muse negotiating $288 in annual Comcast savings and noted that most companies’ support lines are not ready for agents.
查看原文 →查看原文 →

Box CEO展示Jev在企业工作流中的瞬间决策价值Box CEO Aaron Levie Demos Jev for Instant Enterprise Decisions

Box CEO Aaron Levie认为Jev对代理在工作流中做分秒级决策、数据分类、判断调用等数百种企业用例极有帮助。他演示了Box与Jev结合:从Box拉取事故报告,判断是否面向客户及严重程度,将文件移入升级/监控/审查文件夹,并写入元数据模板——几乎瞬时完成且成本几乎为零。可想象用于保险理赔、合同管理、贷款处理、安全审查、客户日志分析等。这是全新一类AI用例。
Box CEO Aaron Levie says Jev will be super helpful for agents making split-second decisions in workflows, data classification, judgment calls and hundreds of other enterprise use-cases. He demoed Box + Jev: pull an incident report from Box, decide whether it is customer-facing and how severe, move the file into escalate/monitor/review folders, and set a metadata template—all nearly instantly and at almost no cost. Think insurance claims, contract management, loan processing, security reviews, customer log analysis. A great new class of AI use-case.
查看原文 →

Jev周末项目与极致性价比演示Jev Weekend Projects and Extreme Cost-Performance Demos

FPV Ventures合伙人Nikunj Kothari搭建了Jevable网站,汇总X上所有有趣的Jev演示,可按类别筛选,并允许用户添加自己的项目。他同时展示Jev在28秒内以0.11美元的成本对3000种儿童零食按多标准打分的结果,令人惊叹。Vercel也宣布免费提供Jev。
FPV Ventures partner Nikunj Kothari built Jevable, a site that showcases all the fun Jev demos on X, filterable by category, with a button for anyone to add their own project. He also showed Jev scoring 3000 kid snacks against multiple criteria in 28 seconds for $0.11. Vercel meanwhile made Jev free for the weekend.
查看原文 →查看原文 →查看原文 →

🛠️ 开发者工具与技巧

OpenClaw团队用roboclaw实时管理会话OpenClaw’s roboclaw Runs Live Team Server and Session Management

Peter Steinberger(OpenClaw)分享:roboclaw在团队服务器上运行,实时在Discord上线,与gpt-live对话,并知道自己同时处理的所有会话。会议中可以直接询问当前和过去会话的上下文。他还推荐打开主页侧边栏让claw重新整理会话,以及在PR落地前用协作工具“去slop”。CUA在所有这些场景也能工作,让代理比单纯截图更高效。
Peter Steinberger of OpenClaw notes that roboclaw runs their team server, is live on Discord, talks with gpt-live and tracks all the sessions it is juggling. The team can ask it about context for current and past sessions right during meetings. He also tips opening the home sidebar to have the claw reorganize sessions, and using collaborators to “deslop” work before PRs land. CUA works across these surfaces too, making the agent more efficient than screenshots alone.
查看原文 →查看原文 →查看原文 →

OpenAI Codex与ChatGPT团队筹备主题演讲OpenAI Codex & ChatGPT Team Prepping Keynote

OpenAI Codex与ChatGPT的Thibault Sottiaux透露,当天与Romain Huet和Sam Altman一起准备主题演讲,最大乐趣是琢磨如何向大家解释,因为里面好东西太多、接连不断。下周就会有一些内容放出,不必久等,未来几个月会逐步展示所有新东西以及它们如何整合。
Thibault Sottiaux of OpenAI’s Codex & ChatGPT team shared that they were working on the keynote with Romain Huet and Sam Altman. Most of the fun was figuring out how to explain everything because there is so much good stuff coming in quick succession. Some things will ship next week so people don’t have to wait, with the full set of new capabilities rolling out over the coming months.
查看原文 →

🌍 其他动态

Zara Zhang:要修好输出,先修好输入Zara Zhang on Fixing Output by Fixing Input

Builder Zara Zhang指出:当你消费的大多是slop时,很难不产出slop。要修好输出,先修好输入。
Builder Zara Zhang observed: it’s hard not to create slop when most things you consume are slop. To fix output, first fix input.
查看原文 →

Dan Shipper为真实AI演示辩护Dan Shipper Defends Real AI Craft Against “Fake” Claims

Every CEO Dan Shipper对有人把Jack Cheng的作品称为“假”感到遗憾。他认为Cheng是最聪明、最 intellectually honest、最注重工艺的人之一,其演示真实且有趣地窥见了未来。这种指控并不好看。
Every CEO Dan Shipper pushed back against claims that Jack Cheng’s work is “fake.” He calls Cheng one of the smartest, most intellectually honest and craft-focused people he has worked with; the demo is real and an interesting peak into the future. Not a good look.
查看原文 →

Aditya Agarwal:AI是一个完整循环Aditya Agarwal: AI Is a Full Circle

SPC普通合伙人、Bevel Health联合创始人Aditya Agarwal写道:AI是一个完整循环。旧的一切又变新了。
SPC General Partner and Bevel Health co-founder Aditya Agarwal noted: AI is a full circle. Everything old is new again.
查看原文 →