🌐 双语
Archive

AI Builders
Digest

2026-09-12 15 builders · 30 tweets · 1 podcasts · 0 blogs

🔥 热点话题

Coinbase CEO Brian Armstrong:Agentic Finance 与给 AI 开户Coinbase CEO Brian Armstrong on Agentic Finance and Banking the AIs

核心结论:AI agent 必须拥有自己的金融账户,因为传统支付轨道无法支撑大量不足 1 美元的交易,而 crypto 轨道天然适配即将到来的 agentic 经济。

Brian Armstrong(Coinbase 联合创始人兼 CEO,同时是长寿公司 New Limit 联合创始人)详细阐述了 Coinbase 的三大方向:Everything Exchange(把股票、商品、crypto、衍生品、预测市场全部上链交易)、稳定币支付,以及 Agentic Finance(AIFi)。数据显示约 76% 的 agent 电商交易金额低于 30 美分,信用卡最低 30 美分固定费用让传统轨道完全不可用。"We don't want the AIs to be unbanked." Coinbase 已提供一键 prompt 工具,让任何 AI agent 立即获得自托管钱包,使用他们孵化并捐给 Linux Foundation 的 X402 协议。agent 可以自主支付数据、工具调用或服务,而不必每次遇到付费墙都找人类刷卡。

公司内部他们在推进递归自我改进:每次 agent 修改服务时都会注入该服务的“大脑”(历史事故、财务控制、A/B 测试、PR 记录),人工修正后必须写回大脑,让一次性通过率持续上升。Armstrong 本人已能用并行 agent 直接交付功能。New Limit 方面,他们用 AI 搜索转录因子组合做表观遗传重编程,已在人源化小鼠模型中成功重编程至少一种人类细胞类型,明年启动一期临床(先做酒精性肝病肝细胞),最终目标是让健康人 40 岁时拥有 20 岁水平的肝、血管、免疫甚至皮肤细胞。

“Crypto was really, really good for humans, and it's going to be essential for AI.”
The Takeaway: AI agents will need their own bank accounts because traditional payment rails cannot handle the flood of sub-dollar transactions in the agentic economy, and crypto is the natural fit.

Brian Armstrong, cofounder and CEO of Coinbase and cofounder of longevity company New Limit, lays out Coinbase’s three big bets: the Everything Exchange (stocks, commodities, crypto, derivatives, prediction markets all on one platform), stablecoin payments, and Agentic Finance (AIFi). Roughly 76% of agent ecommerce transactions are under 30 cents; credit-card minimum fees make traditional rails unusable for them. “We don’t want the AIs to be unbanked.” Coinbase now offers a single-prompt tool that gives any AI agent a self-custodial wallet using the X402 protocol they incubated and contributed to the Linux Foundation. Agents can pay for data, tool calls, or services without a human intervening at every paywall.

Internally Coinbase is building recursive self-improvement: every agent change to a service ingests that service’s “brain” (past incidents, financial controls, A/B tests, PR history); human corrections are written back so one-shot PR acceptance rates climb over time. Armstrong himself is already shipping features by spinning up parallel agents. At New Limit they use AI to search transcription-factor combinations for epigenetic reprogramming, have successfully reprogrammed at least one human cell type in humanized mouse models, and will start Phase 1 trials next year (first indication: alcoholic liver disease hepatocytes), with the long-term goal of restoring youthful function across liver, vasculature, immune cells and skin.

“Crypto was really, really good for humans, and it’s going to be essential for AI.”
查看原文 →

Meta AI 高管 Madhu Guru:企业 AI 为什么大多失败Meta AI Director Madhu Guru: Why Most Enterprise AI Efforts Fail

Meta AI 高级总监 Madhu Guru(前 Google Gemini/Veo 负责人)指出企业 AI 失败的三大常见原因:继续用旧产品方法论(中央 AI 团队由信任的副手组建,增量思维不适合需要实验与发明的 AI)、严重低估 evals 的重要性、以及从外部给业务部门“建设 AI 平台”导致工具与真实工作流脱节。

正确做法是:招聘真正做过 AI 产品的领导者(技术深度 + 能赢得业务信任)、把 evals 当成一等公民、把最强 AI 建设者直接嵌入财务、销售、客服等被改造的部门,与他们一起构建。
Madhu Guru, Sr Director of AI at Meta (previously led Gemini and Veo at Google), diagnoses the three most common reasons enterprise AI programs fail: applying old-school product playbooks (central AI teams staffed by trusted lieutenants, incremental mindset instead of experimentation and invention), under-investing in evals, and building AI platforms “for” the business from the outside so tools never match real workflows and judgment.

What works instead: hire leaders who have actually shipped AI products (technical depth plus the ability to earn trust), make evals a first-class citizen, and embed your strongest AI builders inside the functions you want to transform—build with finance, with sales, with support.
查看原文 →

Peter Yang 对“软件工厂”持怀疑态度Peter Yang Is Skeptical of “Software Factories”

实用 AI 教程作者 Peter Yang 明确表示:除了验证与测试,当前 AI 还远未达到能端到端自我改进产品或从零构建新功能且无需人类在环的程度。一旦过夜循环中出现一个错误假设,整晚的 token 就浪费了。他问:有哪些产品或功能是完全由软件工厂在没有人类定义需求或检查结果的情况下端到端交付的?
Practical AI educator Peter Yang is bluntly skeptical of “software factories.” Outside of verification and testing, he does not believe AI is yet capable of self-improving a product or building a new feature end-to-end without a human in the loop. One wrong assumption overnight turns the entire run into wasted tokens. He asks for concrete examples of products or features that have been built this way with no human defining requirements or reviewing the work.
查看原文 →

💰 创业成功案例

Replit 收购完全基于 Replit 构建的业务Replit Acquires a Business Built Entirely on Replit

Replit CEO Amjad Masad 宣布公司收购了一家完全基于 Replit 构建的业务,并预期这只是第一家,未来会有更多。与此同时 Replit 也在持续推出新能力(Routines with budgets)并看到大量用户用 Astra 做 3D 实验。
Replit CEO Amjad Masad announced that Replit has acquired a business built entirely on Replit and expects it to be the first of many. The company continues shipping new capabilities (Routines with budgets) and is seeing a wave of Astra-powered 3D experiments from the community.
查看原文 →查看原文 →

Zara Zhang:一人公司被高估了Zara Zhang: The “One-Person Company” Is Overrated

Builder Zara Zhang 认为“一人公司”概念被严重高估。AI 确实让个人能做更多事,但创造新事物本身是极度孤独的体验。你需要有人一起头脑风暴、一起受苦、一起庆祝。没有共同绑定的伙伴,极容易失去动力。
Builder Zara Zhang argues the “one-person company” idea is overrated. AI lets one person do far more, yet building something new is a profoundly lonely experience. You need someone to brainstorm with, suffer with, and celebrate with. Without tying yourself to the mast with another person it is extremely easy to lose motivation.
查看原文 →

🛠️ 开发者工具与技巧

OpenAI Thibault Sottiaux:Astra 本周密集上线 + 质量修复OpenAI’s Thibault Sottiaux: Astra Ships Heavily + Quality Reset

OpenAI Codex & ChatGPT 负责人 Thibault Sottiaux 公布 Astra 本周已上线 Images 2.5、GPT-Live-1、Agents API、Data Agent、ChatGPT for Financial Services,而 DevDay 还没到。同时针对用户反馈的质量问题做了全面修复:部分旧技能触发过频、一个导致提前停止的上下文实验已关闭(约影响 4-5k 用户)、移除了配置不当的引擎。社区反馈直接帮助了快速定位。他还欢迎 Git AI 团队加入,该开源工具帮助开发者看清 coding agent 对代码库的真实贡献,并承诺保持开源。
Thibault Sottiaux (Codex & ChatGPT at OpenAI) reported a heavy Astra shipping week—Images 2.5, GPT-Live-1, Agents API, Data Agent, ChatGPT for Financial Services—and DevDay is still ahead. He also detailed a quality reset after user reports: some legacy skills were over-triggering, an opt-in context experiment that caused early stops was disabled (affecting roughly 4-5k users), and misconfigured engines were removed. Community examples accelerated the fixes. He also welcomed the Git AI team; their open-source tool helps developers see exactly where coding agents contribute, and OpenAI will keep it open source.
查看原文 →查看原文 →查看原文 →

Anthropic Thariq:Claude 插件 evals 上线Anthropic’s Thariq: Claude Plugin Evals Are Here

Claude Code 负责人 Thariq 宣布 plugin evals 已可用(claude plugin eval init),解决新模型发布后技能是否仍然有效的问题。他同时指出:如今只看 pass/fail 分数几乎无法解读 evals,很多失败其实来自过于严格的隐藏测试,有时模型的答案比预期结果更合理。
Thariq (Claude Code at Anthropic) launched plugin evals (claude plugin eval init) so developers can check whether their skills still work after new model releases. He also notes that pass/fail scores alone are nearly useless these days—many benchmark failures come from overly strict hidden tests, and sometimes the model’s answer is more sensible than the expected result.
查看原文 →查看原文 →

Vercel、Box、Cursor 的 agent 基础设施进展Agent Infrastructure Progress at Vercel, Box, and Cursor

Vercel CEO Guillermo Rauch 指出 Tailscale 的模型路由器底层由 Vercel AI Gateway 支撑,并宣称“AI Gateways 就是新的 CDN”——直连源站脆弱,自己搭建又痛苦昂贵。Box CEO Aaron Levie 宣布现在可以把 Box 挂载到 agent sandbox,让 agent 像人一样在自己的“电脑”上读写文件。Cursor 设计师 Ryo Lu 展示了 long-lived agents,用于支撑长期大想法。
Vercel CEO Guillermo Rauch highlighted that Tailscale’s model router runs on Vercel AI Gateway and declared “AI Gateways are the new CDNs”—going direct-to-origin is brittle and DIY is painful. Box CEO Aaron Levie announced the ability to mount Box into agent sandboxes so agents can read and write files on their own “computer” the same way humans do. Cursor designer Ryo Lu showcased long-lived agents now available for big, persistent ideas.
查看原文 →查看原文 →查看原文 →

Dan Shipper:从 vibe check 到个人真实工作基准Dan Shipper: From Vibe Checks to Personal Real-Work Benchmarks

Every CEO Dan Shipper 表示公开基准分数几乎无法预测模型在真实工作中的表现。过去三年他们一直用长文 vibe check 做评估,现在内部搭建了平台,让团队基于日常真实任务创建个人基准,把定性 vibe 与定量检查结合起来。
Every CEO Dan Shipper notes that public benchmark scores tell you almost nothing about how a model performs on real work. For three years the team has run long-form vibe checks; they have now built an internal platform so everyone can create personal benchmarks from their actual day-to-day tasks, combining qualitative feel with quantitative measurement.
查看原文 →

🌍 其他动态

Matt Turck 与 Aditya Agarwal 的 9/11 回忆Matt Turck and Aditya Agarwal on 9/11

FirstMark 合伙人 Matt Turck 详细回忆了自己作为位于世贸中心北塔 53 层的年轻创业者在 9/11 当天的经历:团队几乎无人在办公室,唯一在场的同事安全撤离,公司当天就恢复了备份并继续运转。Aditya Agarwal(SPC 合伙人)则回忆自己在 CMU 上课时得知消息,并思考自己是否会有联合 93 航班乘客那样的勇气。
FirstMark partner Matt Turck shared his full 9/11 story as a young founder whose office was on the 53rd floor of the North Tower: almost no one was in yet, the one colleague present escaped, and the team restored from backups and was operational by end of day. SPC partner Aditya Agarwal recalled being in a CMU lecture hall and reflecting on whether he would have had the courage shown by the passengers of United 93.
查看原文 →查看原文 →

其他值得注意的观点Other Notable Takes

Y Combinator CEO Garry Tan 认为考到 1600 分后应该解锁更难的第二考试来识别更高卓越,而不是取消 SAT 让一切变成随机抽签。FPV Ventures 合伙人 Nikunj Kothari 观察到成功有很多父亲,失败却是孤儿——热门交易上 VC 争抢署名的现象越来越明显,新兴 GP 尤其需要靠创始人背书来证明自己的真实贡献。Peter Steinberger 展示了 Astra 在云端用 CUA 玩 Doom 的实验,并顺手给 trycua 提交了 Linux 按键可靠性补丁。
Y Combinator CEO Garry Tan argued that scoring 1600 should unlock a second, harder test to measure higher excellence rather than banning the SAT and turning everything into a lottery. FPV Ventures partner Nikunj Kothari observed that “success has many fathers, but failure is an orphan”—VCs increasingly fight for attribution on the rare hot deals, and emerging GPs especially need founder references to claim their real contributions. Peter Steinberger showed Astra playing Doom via CUA in a cloud session and submitted a Linux key-reliability patch to trycua.
查看原文 →查看原文 →查看原文 →