Rich Sutton:持续学习才是真正的学习,LLM 冻结权重是歧途Rich Sutton: Continual Learning Is Just Learning, Frozen LLM Weights Miss the Point
The Takeaway:真正的智能必须持续从经验中学习并更新权重,而不是一次性预训练后冻结。
强化学习先驱、The Bitter Lesson 作者 Rich Sutton 与前学生、Oak Lab 联合创始人 Khurram Javed 指出,当前 LLM 范式在部署后权重完全不动,只靠上下文,这根本不是学习。Sutton 强调:“我不是怪人,整个领域才怪。他们非要把‘持续学习’单独叫出来,其实那就是学习本身。”他们提出“大世界假设”:世界无限复杂,合成数据永远受限于人类专家的瓶颈,无法替代真实经验。Oak Lab 正推进 Alberta Plan,核心是 continual deep learning(每个权重独立步长 + 持续注入新单元的 continual backprop),让智能体能自我形成抽象、规划并保持知识一致性。Sutton 认为 LLM 是语言能力的重大突破,但只占智能的大约 20-25%,远非全部。
“The Takeaway: True intelligence requires continuous weight updates from experience rather than one-shot pretraining followed by frozen models.
Reinforcement learning pioneer and author of The Bitter Lesson Rich Sutton, together with former student and Oak Lab co-founder Khurram Javed, argue that today’s LLMs stop learning the moment they ship—their weights never change, relying only on context. Sutton insists: “I’m not weird. The field is weird. They need to call it continual learning. It’s just learning.” They advance the big-world hypothesis: the world is massively more complex than any agent, so synthetic data remains bottlenecked by human expertise and cannot replace real experience. Oak Lab is executing the Alberta Plan, centered on continual deep learning (per-weight step-size optimization plus continual backprop that keeps injecting new units) so agents can form their own abstractions, plan, and maintain coherent knowledge. Sutton credits LLMs as a breakthrough in language but estimates they cover only about 20-25% of intelligence.
强化学习先驱、The Bitter Lesson 作者 Rich Sutton 与前学生、Oak Lab 联合创始人 Khurram Javed 指出,当前 LLM 范式在部署后权重完全不动,只靠上下文,这根本不是学习。Sutton 强调:“我不是怪人,整个领域才怪。他们非要把‘持续学习’单独叫出来,其实那就是学习本身。”他们提出“大世界假设”:世界无限复杂,合成数据永远受限于人类专家的瓶颈,无法替代真实经验。Oak Lab 正推进 Alberta Plan,核心是 continual deep learning(每个权重独立步长 + 持续注入新单元的 continual backprop),让智能体能自我形成抽象、规划并保持知识一致性。Sutton 认为 LLM 是语言能力的重大突破,但只占智能的大约 20-25%,远非全部。
“The Takeaway: True intelligence requires continuous weight updates from experience rather than one-shot pretraining followed by frozen models.
Reinforcement learning pioneer and author of The Bitter Lesson Rich Sutton, together with former student and Oak Lab co-founder Khurram Javed, argue that today’s LLMs stop learning the moment they ship—their weights never change, relying only on context. Sutton insists: “I’m not weird. The field is weird. They need to call it continual learning. It’s just learning.” They advance the big-world hypothesis: the world is massively more complex than any agent, so synthetic data remains bottlenecked by human expertise and cannot replace real experience. Oak Lab is executing the Alberta Plan, centered on continual deep learning (per-weight step-size optimization plus continual backprop that keeps injecting new units) so agents can form their own abstractions, plan, and maintain coherent knowledge. Sutton credits LLMs as a breakthrough in language but estimates they cover only about 20-25% of intelligence.
The Takeaway: True intelligence requires continuous weight updates from experience rather than one-shot pretraining followed by frozen models.
Reinforcement learning pioneer and author of The Bitter Lesson Rich Sutton, together with former student and Oak Lab co-founder Khurram Javed, argue that today’s LLMs stop learning the moment they ship—their weights never change, relying only on context. Sutton insists: “I’m not weird. The field is weird. They need to call it continual learning. It’s just learning.” They advance the big-world hypothesis: the world is massively more complex than any agent, so synthetic data remains bottlenecked by human expertise and cannot replace real experience. Oak Lab is executing the Alberta Plan, centered on continual deep learning (per-weight step-size optimization plus continual backprop that keeps injecting new units) so agents can form their own abstractions, plan, and maintain coherent knowledge. Sutton credits LLMs as a breakthrough in language but estimates they cover only about 20-25% of intelligence.
查看原文 →
Reinforcement learning pioneer and author of The Bitter Lesson Rich Sutton, together with former student and Oak Lab co-founder Khurram Javed, argue that today’s LLMs stop learning the moment they ship—their weights never change, relying only on context. Sutton insists: “I’m not weird. The field is weird. They need to call it continual learning. It’s just learning.” They advance the big-world hypothesis: the world is massively more complex than any agent, so synthetic data remains bottlenecked by human expertise and cannot replace real experience. Oak Lab is executing the Alberta Plan, centered on continual deep learning (per-weight step-size optimization plus continual backprop that keeps injecting new units) so agents can form their own abstractions, plan, and maintain coherent knowledge. Sutton credits LLMs as a breakthrough in language but estimates they cover only about 20-25% of intelligence.