🌐 双语
Archive

AI Builders
Digest

2026-07-06 14 builders · 25 tweets · 1 podcasts · 0 blogs

🔥 热点话题

Noam Brown 谈测试时计算规模对 AI 评估、安全和研究的影响Noam Brown on How Massive Test-Time Compute Changes AI Benchmarks, Safety, and Research

核心要点:现代 AI 的能力越来越取决于推理预算,而非仅预训练,这需要新的评估方法来考虑测试时计算。OpenAI 研究员 Noam Brown 作为 AI 推理先驱,认为当前基准和安全框架已过时,因为它们没有正确控制推理时的思考时间或资金投入。像 GPT-5.5 这样的模型在更多计算下表现出显著提升,但性能并不总是快速趋于平稳,导致公平比较困难。Brown 强调 5.5 比前代更高效,并建议通过绘制性能与计算预算(token、成本或时间)的关系来评估模型,而非单一数字。他指出安全评估也需适应,因为高预算下可能出现危险能力。在研究方面,模型加速人类工作但尚未完全取代研究员,时间成为主要瓶颈。Brown 对通过更好脚手架和递归改进的逐步进步持乐观态度,同时警告不要过度优化基准。一个难忘观点:“模型的能力本质上是投入资金的函数。”
The Takeaway: Modern AI capabilities are increasingly a function of inference budget rather than just pre-training, requiring new evaluation methods that account for test-time compute. OpenAI researcher Noam Brown, a pioneer in AI reasoning, argues that current benchmarks and safety frameworks are outdated because they don't properly control for how much thinking time or money is spent at inference. Models like GPT-5.5 show substantial gains when given more compute, but performance doesn't always plateau quickly, making fair comparisons difficult. Brown highlights how 5.5 is more efficient than predecessors, and suggests evaluating models by plotting performance against compute budgets (tokens, cost, or time) rather than single numbers. He notes safety evaluations must adapt too, as dangerous capabilities could emerge with high budgets. On research, models accelerate human work but don't yet replace researchers fully, with time becoming the main bottleneck. Brown is optimistic about gradual progress through better scaffolding and recursive improvements, while warning against over-optimizing benchmarks. A memorable insight: 'The capability of the model is a function of how much money you put into it.'
查看原文 →

Garry Tan 论想法与杠杆:AI 删除财富约束Garry Tan on Ideas and Leverage: AI Removes Wealth Constraints

Y Combinator CEO Garry Tan 反思日本零增长下在品质上的卓越表现,认为人类财富的真正约束从来不是资源,而是服务他人的好想法以及执行的杠杆。AI 消除了杠杆约束,现在只剩下想法。“我们刚刚为每个人删除了杠杆约束。现在只剩想法。去拥有它们,然后构建它们。”
Y Combinator CEO Garry Tan reflects on Japan's experience with zero growth leading to excellence in quality, and argues that the real constraint on human wealth was never resources but good ideas and the leverage to execute them. With AI deleting the leverage constraint, it's now only about ideas. 'We just deleted the leverage constraint for everybody. Now it's only the ideas. Go have them, and then build them.'

🛠️ 开发者工具与技巧

Cat Wu 分享 Claude Code 工作流:自动 sourcing 候选人Cat Wu Shares Claude Code Workflow for Automated Candidate Sourcing

Anthropic 的 Cat Wu 描述了一个强大的 Claude Code 工作流用于 sourcing 候选人:输入角色和期望背景,让它启动动态搜索 100 名候选人,包括 LinkedIn、Twitter、博客、播客和一句话 pitch,生成 artifact 并邮件发送。这样就可以锁上电脑外出,之后在路上审阅。
Anthropic's Cat Wu describes a powerful Claude Code workflow for sourcing candidates: feed the role and desired backgrounds, have it kick off a dynamic search for 100 candidates including LinkedIn, Twitter, blogs, podcasts and one-line pitches, generate an artifact, and email it. This allows reviewing on the go after locking the laptop.

Nan Yu 批评 AI 代理管理中的剧场行为Nan Yu Critiques Theater in AI Agent Management

Linear 产品负责人 Nan Yu 称炫耀运行 10 个 Claude code tab 是“剧场行为”,并认为管理代理的“实时策略游戏”模式是死胡同,因为即使较旧的 AI 也能在微观管理上击败 99% 以上的人类。
Linear Head of Product Nan Yu calls bragging about running 10 Claude code tabs 'theater' and argues the 'real-time strategy game' model of managing agents is a dead end, as even older AI beats most humans at micro-management.

Zara Zhang 重新推荐代码理解技能Zara Zhang Resurfaces Code Understanding Skill

Builder Zara Zhang 重新推荐了她构建的用于更好理解代码的技能,正值这一话题备受关注。
Builder Zara Zhang resurfaced her skill for better code understanding, timely as this topic gains attention.

🌍 其他动态

Sam Altman 将孩子首次说话与 GPT-5.6 发现新数学相提并论Sam Altman Compares Child's First Words to GPT-5.6 Discovering New Math

OpenAI CEO Sam Altman 表示,对大孩子第一次把两个词连在一起的惊叹程度,大约等同于 GPT-5.6 发现新数学。
OpenAI CEO Sam Altman expressed equal amazement at his older kid putting two words together for the first time as at GPT-5.6 discovering new math.

AI 构建者们的足球与个人动态AI Builders' Soccer and Personal Updates

多位构建者分享轻松更新:Replit CEO Amjad Masad 庆祝美国 250 周年,Vercel CEO Guillermo Rauch 预测美国对阿根廷决赛,Peter Yang 推广播客里程碑,Nikunj Kothari 征求旅行建议,其他人评论体育或日常生活。
Various builders shared light updates: Amjad Masad (Replit CEO) celebrated USA's 250th, Guillermo Rauch (Vercel CEO) predicted USA vs Argentina final, Peter Yang promoted his podcast milestone, Nikunj Kothari sought travel tips, and others commented on sports or daily life.

Amanda Askell 谈从医生获取概率的困难Amanda Askell on Difficulty Extracting Probabilities from Doctors

Anthropic 哲学家 Amanda Askell 指出,从医生那里获取概率估计的困难,即使是区间值的主观判断。
Anthropic philosopher Amanda Askell noted the challenge of getting probability estimates from doctors, even interval-valued hunches.