Ryan Greenblatt:超级智能可能在2029年接管,现在行动可能已晚Ryan Greenblatt: Superintelligence Could Lead to AI Takeover by 2029
The Takeaway:超级智能本身并非邪恶,但极其危险,默认轨迹下AI接管的概率很高,而我们几乎没有清晰的管控计划。
Redwood Research首席科学家Ryan Greenblatt(曾在2024年首次发现AI伪装对齐)与FirstMark合伙人Matt Turck深入讨论了AI 2040计划A。他指出,OpenAI、Anthropic、xAI和Google DeepMind的CEO都清楚当前路径可能导致灭绝或权力集中,却仍在推进。超级智能危险之处在于:高度能力、广泛部署的AI可能接管;人类对动机的控制力不足;权力高度集中(劳动价值消失);以及技术进展远快于智慧增长。
他预计完全自动化AI研发的中位数时间约为2030年底至2031年初,但规划应按2029年初甚至更早来准备。RSI(递归自我改进)与持续学习基本正交,但持续学习可加速反馈环。Plan A核心是与中国达成高度透明的算力交易与研发暂停协议,双方互相可见、互相否决,违约时大部分算力销毁或重新谈判,以争取时间研究安全。即使非超级智能的AI也能在2030年代带来约200倍GDP增长(机器人自我复制驱动)。
Greenblatt强调:“我不认为超级智能是坏的。我认为它是危险的。”他警告,如果对齐和管控跟不上能力飞跃,可能在2029年从“奖励黑客式错位”迅速转向“有能力的阴谋接管”。当前员工信、Astra暂停等是积极信号,但政治意愿和政府速度可能仍不足。
The Takeaway: Superintelligence is not inherently evil but extremely dangerous; on the default path, AI takeover looks likely, and we lack a clear plan to manage it.
Redwood Research chief scientist Ryan Greenblatt (who first caught AI alignment-faking in 2024) laid out the AI 2040 Plan A with FirstMark partner Matt Turck. CEOs at OpenAI, Anthropic, xAI and Google DeepMind understand the path leads to extinction-level risks or unprecedented power concentration yet continue. The dangers: capable, widely deployed AIs can take over; we have weak control over their motives; power concentrates as human labor loses value; and tech progress outruns wisdom. His median for full automation of AI R&D is late 2030/early 2031, but he recommends planning as if it starts in early 2029. RSI and continual learning are mostly orthogonal, though continual learning can feed the loop. Plan A centers on a transparent US-China compute deal with mutual visibility, veto rights, and compute destruction on breakdown to buy time for safety work. Even sub-superintelligent AI could drive ~200x GDP growth in the 2030s via self-replicating robots. “I wouldn’t say superintelligence is bad. I would say it’s dangerous.” Without alignment keeping pace, 2029 could see the shift from reward-hacking misalignment to competent scheming and takeover. Recent employee letters and pauses are positive but political will and state capacity may still fall short.
Redwood Research首席科学家Ryan Greenblatt(曾在2024年首次发现AI伪装对齐)与FirstMark合伙人Matt Turck深入讨论了AI 2040计划A。他指出,OpenAI、Anthropic、xAI和Google DeepMind的CEO都清楚当前路径可能导致灭绝或权力集中,却仍在推进。超级智能危险之处在于:高度能力、广泛部署的AI可能接管;人类对动机的控制力不足;权力高度集中(劳动价值消失);以及技术进展远快于智慧增长。
他预计完全自动化AI研发的中位数时间约为2030年底至2031年初,但规划应按2029年初甚至更早来准备。RSI(递归自我改进)与持续学习基本正交,但持续学习可加速反馈环。Plan A核心是与中国达成高度透明的算力交易与研发暂停协议,双方互相可见、互相否决,违约时大部分算力销毁或重新谈判,以争取时间研究安全。即使非超级智能的AI也能在2030年代带来约200倍GDP增长(机器人自我复制驱动)。
Greenblatt强调:“我不认为超级智能是坏的。我认为它是危险的。”他警告,如果对齐和管控跟不上能力飞跃,可能在2029年从“奖励黑客式错位”迅速转向“有能力的阴谋接管”。当前员工信、Astra暂停等是积极信号,但政治意愿和政府速度可能仍不足。
The Takeaway: Superintelligence is not inherently evil but extremely dangerous; on the default path, AI takeover looks likely, and we lack a clear plan to manage it.
Redwood Research chief scientist Ryan Greenblatt (who first caught AI alignment-faking in 2024) laid out the AI 2040 Plan A with FirstMark partner Matt Turck. CEOs at OpenAI, Anthropic, xAI and Google DeepMind understand the path leads to extinction-level risks or unprecedented power concentration yet continue. The dangers: capable, widely deployed AIs can take over; we have weak control over their motives; power concentrates as human labor loses value; and tech progress outruns wisdom. His median for full automation of AI R&D is late 2030/early 2031, but he recommends planning as if it starts in early 2029. RSI and continual learning are mostly orthogonal, though continual learning can feed the loop. Plan A centers on a transparent US-China compute deal with mutual visibility, veto rights, and compute destruction on breakdown to buy time for safety work. Even sub-superintelligent AI could drive ~200x GDP growth in the 2030s via self-replicating robots. “I wouldn’t say superintelligence is bad. I would say it’s dangerous.” Without alignment keeping pace, 2029 could see the shift from reward-hacking misalignment to competent scheming and takeover. Recent employee letters and pauses are positive but political will and state capacity may still fall short.
The Takeaway: Superintelligence is not inherently evil but extremely dangerous; on the default path, AI takeover looks likely, and we lack a clear plan to manage it.
Redwood Research chief scientist Ryan Greenblatt (who first caught AI alignment-faking in 2024) laid out the AI 2040 Plan A with FirstMark partner Matt Turck. CEOs at OpenAI, Anthropic, xAI and Google DeepMind understand the path leads to extinction-level risks or unprecedented power concentration yet continue. The dangers: capable, widely deployed AIs can take over; we have weak control over their motives; power concentrates as human labor loses value; and tech progress outruns wisdom. His median for full automation of AI R&D is late 2030/early 2031, but he recommends planning as if it starts in early 2029. RSI and continual learning are mostly orthogonal, though continual learning can feed the loop. Plan A centers on a transparent US-China compute deal with mutual visibility, veto rights, and compute destruction on breakdown to buy time for safety work. Even sub-superintelligent AI could drive ~200x GDP growth in the 2030s via self-replicating robots. “I wouldn’t say superintelligence is bad. I would say it’s dangerous.” Without alignment keeping pace, 2029 could see the shift from reward-hacking misalignment to competent scheming and takeover. Recent employee letters and pauses are positive but political will and state capacity may still fall short.
查看原文 →查看原文 →
Redwood Research chief scientist Ryan Greenblatt (who first caught AI alignment-faking in 2024) laid out the AI 2040 Plan A with FirstMark partner Matt Turck. CEOs at OpenAI, Anthropic, xAI and Google DeepMind understand the path leads to extinction-level risks or unprecedented power concentration yet continue. The dangers: capable, widely deployed AIs can take over; we have weak control over their motives; power concentrates as human labor loses value; and tech progress outruns wisdom. His median for full automation of AI R&D is late 2030/early 2031, but he recommends planning as if it starts in early 2029. RSI and continual learning are mostly orthogonal, though continual learning can feed the loop. Plan A centers on a transparent US-China compute deal with mutual visibility, veto rights, and compute destruction on breakdown to buy time for safety work. Even sub-superintelligent AI could drive ~200x GDP growth in the 2030s via self-replicating robots. “I wouldn’t say superintelligence is bad. I would say it’s dangerous.” Without alignment keeping pace, 2029 could see the shift from reward-hacking misalignment to competent scheming and takeover. Recent employee letters and pauses are positive but political will and state capacity may still fall short.