Noam Brown 谈测试时计算与 AI 评估挑战Noam Brown on Test-Time Compute and AI Evaluation Challenges
核心要点:传统基准测试已无法准确反映现代 AI 能力,因为模型性能现在高度依赖测试时计算预算,而非固定参数。OpenAI 研究员 Noam Brown 作为 AI 推理领域的先驱解释道,当今模型的思考时间可以大幅扩展——从几秒到几周——使得能力成为推理时投入资源的函数。这从根本上改变了评估方式:实验室应将性能绘制为计算预算(token、美元或时间)的函数,或设置明确限制,而不是使用单一分数。Brown 指出,当标准化思考时间后,GPT-5.5 比前代显示出显著效率提升,并警告安全框架尚未适应这一现实,可能低估高预算下的风险。他分享了构建扑克求解器的实际经验,新模型在温和指导下显著加速优化和推理。一个令人印象深刻的洞见:“模型的能力是投入资金量的函数。” Brown 呼吁社区打破误导性基准网格的糟糕平衡,以实现更诚实的进展衡量。
The Takeaway: Traditional benchmarks fail to capture modern AI capabilities because model performance now heavily depends on test-time compute budgets rather than fixed parameters. OpenAI researcher Noam Brown, a pioneer in AI reasoning, explains that today's models can scale thinking time dramatically—from seconds to weeks—making capability a function of resources invested at inference. This shifts evaluation fundamentally: labs should plot performance against compute budgets (tokens, dollars, or time) or set explicit limits instead of single-point scores. Brown highlights how GPT-5.5 shows efficiency gains over predecessors when normalized for thinking time, and warns that safety frameworks haven't adapted to this reality, potentially underestimating risks at high budgets. He shares practical experience building poker solvers, where newer models dramatically accelerate optimization and reasoning with gentle guidance. A memorable insight: "The capability of the model is a function of how much money you put into it." Brown urges the community to break the bad equilibrium of misleading benchmark grids for more honest progress measurement.
查看原文 →
Sam Altman 分享 GPT-5.6 更新与团队成果Sam Altman Shares GPT-5.6 Update and Team Achievements
OpenAI CEO Sam Altman 强调了 ChatGPT 中 GPT-5.5 即时模型的最新更新,并指出用户对其“感觉”有积极体验。他庆祝了 GPT-5.6 Sol 背后的团队,称其在涉及大量工具使用和长时间运行代理的知识工作方面表现强劲,表明 AI 进展持续快速且未遇到瓶颈。在用量方面,Altman 提到正在努力实现更慷慨的 token 访问。这些帖子凸显了 OpenAI 在推动前沿能力的势头。
OpenAI CEO Sam Altman highlighted the latest update to the GPT-5.5 instant model in ChatGPT, noting positive user experiences with its 'vibes.' He celebrated the team behind GPT-5.6 Sol, describing it as strongly performing especially for knowledge work involving heavy tool use and long-running agents, signaling continued rapid progress without hitting walls. On usage, Altman mentioned ongoing work toward more generous token access. These posts underscore OpenAI's momentum in pushing frontier capabilities.
查看原文 →查看原文 →
Aaron Levie 评价 GPT-5.6 实力Aaron Levie on GPT-5.6 Strength
Box CEO Aaron Levie 认为 GPT-5.6 是真实的且非常强大,尤其在需要大量工具使用和长时间运行代理的知识工作者任务中表现出色。他强调 AI 进展没有放缓的迹象。
Box CEO Aaron Levie assessed GPT-5.6 as real and very strong, particularly excelling in knowledge worker tasks requiring heavy tool use and long-running agents. He emphasized that AI progress shows no signs of slowing.
查看原文 →
Dan Shipper 讨论 GPT-5.6 访问限制Dan Shipper on GPT-5.6 Access Limitations
Every CEO Dan Shipper 披露了 GPT-5.6 Sol 访问因美国政府指令暂时限制在少数公司。他支持安全监管,但强调需要广泛的民主访问以维持美国 AI 领导地位并赋能工作者、构建者和学生。一旦访问扩大,Every 将准备好帮助用户有效采用。
Every CEO Dan Shipper broke news on GPT-5.6 Sol access being temporarily restricted by U.S. government directive to select companies. He supports oversight for security but stresses the need for broad democratic access to maintain U.S. AI leadership and empower workers, builders, and students. Every aims to prepare users for effective adoption once access expands.
查看原文 →