Google AI 基础设施负责人 Amin Vahdat:前沿 AI 的物理与经济规律Google AI Infrastructure Chief Amin Vahdat on the Physics & Economics of Frontier AI
The Takeaway:在前沿 AI 规模下,真正的瓶颈不是 FLOPS,而是可交付的 goodput(有效吞吐),而电力是长期最根本的约束。
Google AI 基础设施负责人 Amin Vahdat 正站在人类历史上最大规模 CapEx 建设的中心。Google 今年预计投入超过 2000 亿美元,其中绝大部分用于数据中心。AI 数据中心与传统数据中心的核心差异在于专业化:建筑、电力、冷却、网络都与硬件深度共设计,而不是按 20-30 年通用寿命规划。
Vahdat 强调,芯片理论峰值 FLOPS 是虚荣指标。真正重要的是 workload 的 delivered goodput——在真实故障条件下实际完成有效工作的比例。在 10 万加速器规模,故障可能每小时多次发生,同步训练或 agent 工作负载中任何一个组件挂掉都可能让整个作业停摆。因此必须实现近实时检测与恢复。
TPU 项目从 2013 年的逆向赌注起步,最初只做推理,后来扩展到训练、transformer 与推荐系统。2026 年他们同时发布了专为推理优化的 8I 与专为训练优化的 8T,因为推理需求已足够大,值得进一步专业化,同时两款芯片仍能互相兜底。与 DeepMind 的共设计是 Google 的独特优势:模型团队可以直接影响尚未 tape-out 的芯片架构,硬件团队也能提前几年预判模型演进方向。
电力是根本约束。Google 优先与电网合作多年规划,必要时本地发电并反哺电网,以获得统计复用带来的可靠性与成本优势。长 horizon agent 正在改变工作负载形态:交互从秒级变到毫秒级,CPU、网络与存储需求同步暴涨。光学电路交换与波分复用已在 Google 数据中心运行十余年,能在毫秒级重配置拓扑并替换故障机架。
Vahdat 还提到轨道数据中心作为真正的 moonshot:太空太阳能可提供 1.4 倍功率与接近 100% 日照,但冷却、可靠性与维修挑战巨大。他预测十年后的前沿超级计算机会更高度集成,单机架可能达到数兆瓦,光纤大幅减少,甚至可能直接发射入轨。
“我们衡量自己的标准是 delivered goodput,而不是理论上的 benchmark 吞吐量。”
Google AI 基础设施负责人 Amin Vahdat 正站在人类历史上最大规模 CapEx 建设的中心。Google 今年预计投入超过 2000 亿美元,其中绝大部分用于数据中心。AI 数据中心与传统数据中心的核心差异在于专业化:建筑、电力、冷却、网络都与硬件深度共设计,而不是按 20-30 年通用寿命规划。
Vahdat 强调,芯片理论峰值 FLOPS 是虚荣指标。真正重要的是 workload 的 delivered goodput——在真实故障条件下实际完成有效工作的比例。在 10 万加速器规模,故障可能每小时多次发生,同步训练或 agent 工作负载中任何一个组件挂掉都可能让整个作业停摆。因此必须实现近实时检测与恢复。
TPU 项目从 2013 年的逆向赌注起步,最初只做推理,后来扩展到训练、transformer 与推荐系统。2026 年他们同时发布了专为推理优化的 8I 与专为训练优化的 8T,因为推理需求已足够大,值得进一步专业化,同时两款芯片仍能互相兜底。与 DeepMind 的共设计是 Google 的独特优势:模型团队可以直接影响尚未 tape-out 的芯片架构,硬件团队也能提前几年预判模型演进方向。
电力是根本约束。Google 优先与电网合作多年规划,必要时本地发电并反哺电网,以获得统计复用带来的可靠性与成本优势。长 horizon agent 正在改变工作负载形态:交互从秒级变到毫秒级,CPU、网络与存储需求同步暴涨。光学电路交换与波分复用已在 Google 数据中心运行十余年,能在毫秒级重配置拓扑并替换故障机架。
Vahdat 还提到轨道数据中心作为真正的 moonshot:太空太阳能可提供 1.4 倍功率与接近 100% 日照,但冷却、可靠性与维修挑战巨大。他预测十年后的前沿超级计算机会更高度集成,单机架可能达到数兆瓦,光纤大幅减少,甚至可能直接发射入轨。
“我们衡量自己的标准是 delivered goodput,而不是理论上的 benchmark 吞吐量。”
The Takeaway: At frontier AI scale the real bottleneck is not FLOPS but delivered goodput, and power is the fundamental long-term constraint.
Amin Vahdat, head of Google’s AI infrastructure, sits at the center of the largest CapEx build-out in human history. Google alone is expected to spend more than $200 billion this year, mostly on data centers. AI data centers differ from traditional ones mainly through specialization: buildings, power, cooling and networking are co-designed with the hardware rather than planned for a 20–30 year generic lifetime.
Vahdat insists theoretical peak FLOPS is a vanity metric. What matters is delivered goodput—the fraction of useful work actually completed under real failure conditions. At 100,000-accelerator scale, failures can occur multiple times an hour; in synchronous training or agentic workloads a single failed component can stall the entire job. Near-real-time detection and recovery are therefore non-negotiable.
The TPU program began as a contrarian bet in 2013, first for inference then training, transformers and recommenders. In 2026 Google released two chips—8I specialized for inference and 8T for training—because inference demand had grown large enough to justify further specialization, while both chips can still run the other’s workload. Deep co-design with DeepMind is a unique Google advantage: model teams can still influence chips that have not yet taped out, and hardware teams can project model trends years ahead.
Power is the binding constraint. Google prefers multi-year grid planning and, when necessary, local generation that can also feed the grid, capturing statistical multiplexing benefits. Long-horizon agents are reshaping demand: interaction latency drops from seconds to milliseconds and CPU, networking and storage requirements explode alongside accelerators. Optical circuit switching and wave-division multiplexing, deployed inside Google data centers for over a decade, allow topology reconfiguration and failed-rack replacement in milliseconds.
Vahdat also discusses orbital data centers as a genuine moonshot: space offers 1.4× solar power and near-100 % sunlight, but cooling, reliability and repair become far harder. He expects the 2036 frontier supercomputer to be far more tightly integrated, with multi-megawatt racks, far less fiber, and possibly direct launch into orbit.
“We hold ourselves accountable by delivered goodput, not theoretical benchmark throughput.”
查看原文 →
Amin Vahdat, head of Google’s AI infrastructure, sits at the center of the largest CapEx build-out in human history. Google alone is expected to spend more than $200 billion this year, mostly on data centers. AI data centers differ from traditional ones mainly through specialization: buildings, power, cooling and networking are co-designed with the hardware rather than planned for a 20–30 year generic lifetime.
Vahdat insists theoretical peak FLOPS is a vanity metric. What matters is delivered goodput—the fraction of useful work actually completed under real failure conditions. At 100,000-accelerator scale, failures can occur multiple times an hour; in synchronous training or agentic workloads a single failed component can stall the entire job. Near-real-time detection and recovery are therefore non-negotiable.
The TPU program began as a contrarian bet in 2013, first for inference then training, transformers and recommenders. In 2026 Google released two chips—8I specialized for inference and 8T for training—because inference demand had grown large enough to justify further specialization, while both chips can still run the other’s workload. Deep co-design with DeepMind is a unique Google advantage: model teams can still influence chips that have not yet taped out, and hardware teams can project model trends years ahead.
Power is the binding constraint. Google prefers multi-year grid planning and, when necessary, local generation that can also feed the grid, capturing statistical multiplexing benefits. Long-horizon agents are reshaping demand: interaction latency drops from seconds to milliseconds and CPU, networking and storage requirements explode alongside accelerators. Optical circuit switching and wave-division multiplexing, deployed inside Google data centers for over a decade, allow topology reconfiguration and failed-rack replacement in milliseconds.
Vahdat also discusses orbital data centers as a genuine moonshot: space offers 1.4× solar power and near-100 % sunlight, but cooling, reliability and repair become far harder. He expects the 2036 frontier supercomputer to be far more tightly integrated, with multi-megawatt racks, far less fiber, and possibly direct launch into orbit.
“We hold ourselves accountable by delivered goodput, not theoretical benchmark throughput.”