Oriol Vinyals 谈世界模型、多模态与 AGI 路径Oriol Vinyals on World Models, Multimodal Progress, and AGI Paths
核心要点:Google 通过 Gemini 对世界模型的大力投入,代表了一种独特的赌注,即让 AI 通过丰富的多模态来理解物理世界,而不是单纯依赖语言或代码自我改进循环。Gemini 共同负责人、DeepMind 深度学习先驱 Oriol Vinyals 强调,虽然语言提供了大量知识提炼,但视频和图像数据对更深层的物理和因果理解仍有巨大潜力。Omni 等模型在交互式视频生成和编辑方面表现出色,但纯视觉转移学习的“GPT 时刻”尚未到来。他指出了表征学习以及在没有显式语言标签的情况下连接概念的挑战。在 Agent 方面,他强调了脚手架、长期运行可靠性和使用基于文件的非参数存储进行持续学习的内存系统。他对 AGI 时间线保持乐观,认为当前模型在某些定义下已接近 AGI,同时强调需要更好的经验学习能力。
“那是我见过的那一刻吗?可能还没有,我们很可能拥有最先进或最先进的多模态配方之一,将一切混合在一起。”
“那是我见过的那一刻吗?可能还没有,我们很可能拥有最先进或最先进的多模态配方之一,将一切混合在一起。”
The Takeaway: Google's heavy investment in world models through Gemini represents a distinct bet on grounding AI in rich multimodal understanding of the physical world rather than pure language or code self-improvement loops. Oriol Vinyals, co-lead of Gemini at Google DeepMind with a storied career in deep learning breakthroughs, emphasizes that while language has provided massive knowledge distillation, video and image data hold untapped potential for deeper physics and causal understanding. Progress in models like Omni shows impressive interactive video generation and editing, yet the 'GPT moment' for pure visual transfer learning remains ahead. Vinyals notes the challenges in representation learning and connecting concepts without explicit language labels. On agents, he highlights improvements in scaffolding, long-running reliability, and memory systems using file-based nonparametric storage for continual learning. He remains optimistic about AGI timelines, viewing current models as approaching it under certain definitions while stressing the need for better experiential learning.
"That is the moment that have I seen that? Probably not, and most likely we have the most advanced or one of the most advanced multimodal recipe that mixes everything."
查看原文 →
"That is the moment that have I seen that? Probably not, and most likely we have the most advanced or one of the most advanced multimodal recipe that mixes everything."