OpenAI 的 Dan Roberts 探讨 AI 如何通过强化学习实现科学发现OpenAI's Dan Roberts on How AI Makes Scientific Discoveries via RL
The Takeaway: AI 系统正从执行指令转向自主进行深度科学发现,尤其在数学领域,通过强化学习(RL)和测试时计算取得突破。
OpenAI 基础强化学习团队负责人 Dan Roberts 拥有理论物理背景,他解释了 RL 如何让模型在长时程任务中通过与环境互动、获得反馈来学习。OpenAI 在 Erdős 问题上取得进展,使用非正式推理方法假设一个长期被认为正确的猜想为假,并通过持久探索推翻了它。这与 DeepMind 的形式化 Lean 证明方法形成对比。
Roberts 强调 RL 现在是 "蛋糕" 而非 "樱桃",语言作为智能的 grounding 层至关重要。模型在测试时生成思考过程,利用大量计算来解决问题。物理学教给我们从大系统反推小模型以理解 emergent 现象。
"I feel really excited that we will get to really answer a lot of fundamental questions in the field of science that we care about with the aid or the models being the driving force."
OpenAI 基础强化学习团队负责人 Dan Roberts 拥有理论物理背景,他解释了 RL 如何让模型在长时程任务中通过与环境互动、获得反馈来学习。OpenAI 在 Erdős 问题上取得进展,使用非正式推理方法假设一个长期被认为正确的猜想为假,并通过持久探索推翻了它。这与 DeepMind 的形式化 Lean 证明方法形成对比。
Roberts 强调 RL 现在是 "蛋糕" 而非 "樱桃",语言作为智能的 grounding 层至关重要。模型在测试时生成思考过程,利用大量计算来解决问题。物理学教给我们从大系统反推小模型以理解 emergent 现象。
"I feel really excited that we will get to really answer a lot of fundamental questions in the field of science that we care about with the aid or the models being the driving force."
The Takeaway: AI systems are shifting from following instructions to autonomously making deep scientific discoveries, particularly in mathematics, powered by reinforcement learning (RL) and test-time compute.
Dan Roberts, lead of OpenAI's Foundations of Reinforcement Learning team with a theoretical physics background, explains how RL enables models to learn by interacting with environments and receiving feedback over long horizons. OpenAI's progress on Erdős problems involved assuming a long-held conjecture false and pursuing a contrarian path with persistence, contrasting DeepMind's formal Lean proofs.
Roberts notes RL is now "the cake, not the cherry," with language as the key grounding for intelligence. Models generate thought processes at test time, leveraging massive compute. Physics teaches scaling from big systems back to understandable small models.
"I feel really excited that we will get to really answer a lot of fundamental questions in the field of science that we care about with the aid or the models being the driving force."
查看原文 →
Dan Roberts, lead of OpenAI's Foundations of Reinforcement Learning team with a theoretical physics background, explains how RL enables models to learn by interacting with environments and receiving feedback over long horizons. OpenAI's progress on Erdős problems involved assuming a long-held conjecture false and pursuing a contrarian path with persistence, contrasting DeepMind's formal Lean proofs.
Roberts notes RL is now "the cake, not the cherry," with language as the key grounding for intelligence. Models generate thought processes at test time, leveraging massive compute. Physics teaches scaling from big systems back to understandable small models.
"I feel really excited that we will get to really answer a lot of fundamental questions in the field of science that we care about with the aid or the models being the driving force."