OpenAI 的 Jan Dubois 解析 AI 真实进步时刻OpenAI's Jan Dubois on Why AI Progress Suddenly Feels Real
核心要点:AI 能力增长是连续的,但跨越可靠性门槛后会让人感觉像突然的阶跃函数,从而在编码和智能体任务中实现真实世界可用性。OpenAI 后训练前沿团队联合负责人 Jan Dubois(曾参与斯坦福 Alpaca 项目)解释了其团队如何将针对可验证奖励优化的推理模型(数学/编码竞赛)转化为处理现实世界复杂知识工作的工具。GPT-5.5 在效率(许多任务快 2 倍)、智能体能力和公司协同上取得重大进展。强化学习现在已从竞赛扩展到用户实用性。Dubois 强调,虽然预训练通过更大模型和合成数据扩展,但后训练和持续学习仍是主要未解决问题。'不同垂直领域总会为最后一公里留下大量空间。'
The Takeaway: AI capability growth is continuous, but crossing reliability thresholds makes it feel like sudden step functions, enabling real-world usefulness especially in coding and agentic tasks. Jan Dubois, who co-leads the Post-Training Frontiers team at OpenAI and previously co-authored Stanford Alpaca, explains how his team turned reasoning models optimized for verifiable rewards (like math/coding competitions) into tools for messy real-world knowledge work. GPT-5.5 represents major gains in efficiency (2x faster on many tasks), agent capabilities, and company-wide alignment. Progress in reinforcement learning now generalizes beyond competitions to user utility. Dubois emphasizes that while pre-training scales with larger models and synthetic data, post-training and continual learning remain key unsolved challenges. 'There will always be a lot of space left for this last mile in different verticals.'
查看原文 →