Fractile CEO Walter Goodwin:全栈推理芯片与内存带宽的极限追求Fractile CEO Walter Goodwin: Full-Stack Inference Chips and the Memory Bandwidth Frontier
The Takeaway:在AI芯片领域,谁能结构性拿下3到6个月的领先优势,谁就能拿走所有前沿部署。
Fractile创始人兼CEO Walter Goodwin正在打造全栈AI推理芯片,目标是让超大模型以远超今天GPU的速度运行,同时能扩展到更长上下文和更大规模。他强调公司从2022年夏天起步时就押注推理时代,并坚持端到端自研——从前端架构、物理设计到先进封装全部自己做,团队只有约150人,却能形成极短的闭环反馈。
最初他们探索基于SRAM的方案(类似Grok或Cerberus),能实现每秒数千token的速度,但很快意识到上下文长度爆炸会让容量成为瓶颈。过去两年他们转向与内存厂商深度合作,追求对低成本DRAM的极高带宽访问,希望把GPU的可扩展性与SRAM芯片的速度优势结合起来。平台计划在2027年下半年量产。
Goodwin指出,真正拉开差距的不是“更灵敏的聊天机器人”,而是让多万亿参数模型在数千token/s下舒适运行,从而让长程Agent能力发生质变。他引用了Henry Ford的“更快的马”比喻:用户要的是汽车,而不是更快的马。芯片必须同时具备极高带宽和可负担的内存容量,因为数据中心规模推理的经济学最终塌缩到每GB内存成本。
在市场结构上,他观察到所有大规模部署方都在同时押注多个平台,但真正有新能力的芯片会成为刚需——前沿实验室需要最快的部署速度来守住溢价智能窗口。第三方芯片玩家之所以能长期存在,是因为任何实验室如果All-in专有硅,一旦对手发现只在对方芯片上才能跑的计算突破,自己可能在9个月内就出局。因此实验室反而需要共享相同的平台,而差异化下注则由芯片公司来承担。
“如果你能结构性切出3到6个月的优势,你就会拿走所有那些部署。”
Fractile创始人兼CEO Walter Goodwin正在打造全栈AI推理芯片,目标是让超大模型以远超今天GPU的速度运行,同时能扩展到更长上下文和更大规模。他强调公司从2022年夏天起步时就押注推理时代,并坚持端到端自研——从前端架构、物理设计到先进封装全部自己做,团队只有约150人,却能形成极短的闭环反馈。
最初他们探索基于SRAM的方案(类似Grok或Cerberus),能实现每秒数千token的速度,但很快意识到上下文长度爆炸会让容量成为瓶颈。过去两年他们转向与内存厂商深度合作,追求对低成本DRAM的极高带宽访问,希望把GPU的可扩展性与SRAM芯片的速度优势结合起来。平台计划在2027年下半年量产。
Goodwin指出,真正拉开差距的不是“更灵敏的聊天机器人”,而是让多万亿参数模型在数千token/s下舒适运行,从而让长程Agent能力发生质变。他引用了Henry Ford的“更快的马”比喻:用户要的是汽车,而不是更快的马。芯片必须同时具备极高带宽和可负担的内存容量,因为数据中心规模推理的经济学最终塌缩到每GB内存成本。
在市场结构上,他观察到所有大规模部署方都在同时押注多个平台,但真正有新能力的芯片会成为刚需——前沿实验室需要最快的部署速度来守住溢价智能窗口。第三方芯片玩家之所以能长期存在,是因为任何实验室如果All-in专有硅,一旦对手发现只在对方芯片上才能跑的计算突破,自己可能在9个月内就出局。因此实验室反而需要共享相同的平台,而差异化下注则由芯片公司来承担。
“如果你能结构性切出3到6个月的优势,你就会拿走所有那些部署。”
The Takeaway: In the AI chip space, whoever can structurally carve out a three-to-six-month advantage will win all the frontier deployments.
Walter Goodwin, founder and CEO of Fractile, is building full-stack AI inference chips designed to run the world's largest models far faster than today's hardware while scaling to longer contexts and bigger models. The company started in summer 2022 with a clear bet on the inference era and insists on owning the entire stack—from architecture and front-end design through physical design and advanced packaging. With roughly 150 people they keep a tight, agile loop that lets them place and refine bets ahead of the workload curve.
Early on they pursued an SRAM-based approach (in the spirit of Grok or Cerberus) that could deliver thousands of tokens per second. By late 2023 they grew concerned about scalability as context lengths exploded. For the past two years they have worked with memory vendors to unlock extremely high bandwidth to higher-capacity, lower-cost DRAM, aiming to combine GPU-like scalability with the raw speed of SRAM chips. The resulting platform is scheduled to ramp in the second half of next year.
Goodwin argues the real payoff is not a snappier chatbot—the “faster horse”—but the ability to run multi-trillion-parameter models at many thousands of tokens per second, which becomes a fundamental capability elevator for long-running agents. The technical requirement is an “ineffable” combination: extremely high bandwidth so weights and state can be loaded thousands of times per second, plus economical memory, because data-center-scale inference economics collapse to cost per gigabyte.
On market structure he notes that every large deployer is diversifying across platforms, yet chips that unlock genuinely new capabilities become must-haves. Frontier labs need the fastest possible deployment to defend their premium intelligence window against open-source pressure. Third-party chip players remain viable precisely because any lab that goes all-in on proprietary silicon risks being blindsided if a competitor discovers a breakthrough that only runs on the other chip; the nine-month lag to catch up could be fatal. Differentiated hardware bets are therefore safer left to specialists while labs share common platforms.
“If you can just find a way to structurally carve out a three to six months advantage, you will be winning all of those deployments.”
查看原文 →
Walter Goodwin, founder and CEO of Fractile, is building full-stack AI inference chips designed to run the world's largest models far faster than today's hardware while scaling to longer contexts and bigger models. The company started in summer 2022 with a clear bet on the inference era and insists on owning the entire stack—from architecture and front-end design through physical design and advanced packaging. With roughly 150 people they keep a tight, agile loop that lets them place and refine bets ahead of the workload curve.
Early on they pursued an SRAM-based approach (in the spirit of Grok or Cerberus) that could deliver thousands of tokens per second. By late 2023 they grew concerned about scalability as context lengths exploded. For the past two years they have worked with memory vendors to unlock extremely high bandwidth to higher-capacity, lower-cost DRAM, aiming to combine GPU-like scalability with the raw speed of SRAM chips. The resulting platform is scheduled to ramp in the second half of next year.
Goodwin argues the real payoff is not a snappier chatbot—the “faster horse”—but the ability to run multi-trillion-parameter models at many thousands of tokens per second, which becomes a fundamental capability elevator for long-running agents. The technical requirement is an “ineffable” combination: extremely high bandwidth so weights and state can be loaded thousands of times per second, plus economical memory, because data-center-scale inference economics collapse to cost per gigabyte.
On market structure he notes that every large deployer is diversifying across platforms, yet chips that unlock genuinely new capabilities become must-haves. Frontier labs need the fastest possible deployment to defend their premium intelligence window against open-source pressure. Third-party chip players remain viable precisely because any lab that goes all-in on proprietary silicon risks being blindsided if a competitor discovers a breakthrough that only runs on the other chip; the nine-month lag to catch up could be fatal. Differentiated hardware bets are therefore safer left to specialists while labs share common platforms.
“If you can just find a way to structurally carve out a three to six months advantage, you will be winning all of those deployments.”