Noam Brown 谈测试时计算规模对 AI 评估、安全和研究的影响Noam Brown on How Massive Test-Time Compute Changes AI Benchmarks, Safety, and Research
核心要点:现代 AI 的能力越来越取决于推理预算,而非仅预训练,这需要新的评估方法来考虑测试时计算。OpenAI 研究员 Noam Brown 作为 AI 推理先驱,认为当前基准和安全框架已过时,因为它们没有正确控制推理时的思考时间或资金投入。像 GPT-5.5 这样的模型在更多计算下表现出显著提升,但性能并不总是快速趋于平稳,导致公平比较困难。Brown 强调 5.5 比前代更高效,并建议通过绘制性能与计算预算(token、成本或时间)的关系来评估模型,而非单一数字。他指出安全评估也需适应,因为高预算下可能出现危险能力。在研究方面,模型加速人类工作但尚未完全取代研究员,时间成为主要瓶颈。Brown 对通过更好脚手架和递归改进的逐步进步持乐观态度,同时警告不要过度优化基准。一个难忘观点:“模型的能力本质上是投入资金的函数。”
The Takeaway: Modern AI capabilities are increasingly a function of inference budget rather than just pre-training, requiring new evaluation methods that account for test-time compute. OpenAI researcher Noam Brown, a pioneer in AI reasoning, argues that current benchmarks and safety frameworks are outdated because they don't properly control for how much thinking time or money is spent at inference. Models like GPT-5.5 show substantial gains when given more compute, but performance doesn't always plateau quickly, making fair comparisons difficult. Brown highlights how 5.5 is more efficient than predecessors, and suggests evaluating models by plotting performance against compute budgets (tokens, cost, or time) rather than single numbers. He notes safety evaluations must adapt too, as dangerous capabilities could emerge with high budgets. On research, models accelerate human work but don't yet replace researchers fully, with time becoming the main bottleneck. Brown is optimistic about gradual progress through better scaffolding and recursive improvements, while warning against over-optimizing benchmarks. A memorable insight: 'The capability of the model is a function of how much money you put into it.'
查看原文 →