OpenAI 研究员 Dan Roberts 谈 RL 与 AI 科学发现OpenAI Researcher Dan Roberts on RL and AI Scientific Discovery
核心要点:基于强化学习的 AI 系统正通过大规模测试时计算、反直觉探索和可验证奖励,推动真正的科学发现。OpenAI 基础强化学习团队负责人、拥有理论物理背景的 Dan Roberts 解释了 RL 如何从后训练润色(RLHF)演变为核心引擎,让模型能在复杂数学问题上推理数小时,例如推翻长期存在的 Erdős 猜想。模型现在会假设假设为假,在漫长计算路径中坚持,并连接不同领域——这感觉像是自主发现。令人难忘的引用:“ChatGPT 能做的一件事就是假设它是假的。当你逆势而行……你真的需要有强烈的信念。”Roberts 认为通往 AI 作为研究伙伴的道路是平滑的,而非突然飞跃,物理启发的扩展洞见有助于理解涌现行为。预训练的语言先验与 RL 的交互反馈的相互作用使这一切成为可能,将计算转化为解决开放科学问题的智能。
The Takeaway: AI systems powered by reinforcement learning are poised to drive genuine scientific discoveries by combining massive test-time compute, contrarian exploration, and verifiable rewards. OpenAI's Dan Roberts, leading the Foundations of Reinforcement Learning team with a background in theoretical physics, explains how RL has evolved from a post-training polish (RLHF) to the core engine enabling models to reason for hours on complex math problems like disproving long-standing Erdős conjectures. Models now assume hypotheses false, persist through long calculation paths, and connect disparate fields—capabilities that feel like autonomous discovery. A memorable quote: "One of the things that ChatGPT was able to do was assume it was false. When you go against the grain... you really have to have strong conviction." Roberts sees a smooth progression toward AI as research partners, not a sudden leap, with physics-inspired scaling insights helping understand emergent behaviors. The interplay of pretraining's language prior and RL's interactive feedback makes this possible, turning compute into intelligence for open scientific questions.
查看原文 →