[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-amd-acquires-taalas-msic-etched-weights":3,"news-related-dfdc3216-52aa-4a78-9bf5-859affc37d17":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"dfdc3216-52aa-4a78-9bf5-859affc37d17","AMD 收下 Taalas：把模型权重刻进芯片，推理的内存墙还剩多少？","8 月 6 日 AMD 宣布收购多伦多 AI 芯片初创 Taalas。后者的 HC1 测试芯片用台积电 6nm 工艺把 Llama 3.1 8B 模型权重直接烧进 mask-ROM，推理速度比英伟达 GPU 快 48 倍、比 Cerebras 快 8.5 倍。HC2 计划支持约 200 亿参数模型，将并入 AMD Helios 机架与 ROCm 软件栈。","# AMD 收下 Taalas：把模型权重刻进芯片，推理的内存墙还剩多少？\n\n8 月 6 日，AMD 宣布收购多伦多 AI 芯片初创公司 Taalas。这是 AMD 近九个月内的第三笔 AI 收购——此前 MK1（去年 11 月）与 Mext（今年 6 月）已把它推进推理赛道，FastFlowLM 团队则在 7 月并入。如今 Taalas 的加入，给 AMD 的「推理全栈」补上了一块很特殊的拼图：把模型权重直接刻进晶体管。\n\n## Taalas 的核心思路：让权重不再「住」在 HBM 里\n\n现代推理最大的隐性瓶颈并不是算力，而是 HBM（高带宽内存）与处理器之间反复搬运权重的带宽。每生成一个 token，模型都要把整组权重从 HBM 拉进 GPU 一遍——这是今天几乎所有「按 token 卖」的部署方式里最贵的部分。\n\nTaalas 的做法是直接把这一过程取消。它的 HC1 测试芯片用台积电 6nm 工艺制造，把 Meta 的 Llama 3.1 8B 模型权重烧进芯片里的 mask-ROM 区域，等同于把模型硬连线成电路的一部分；剩余区域仍保留 SRAM，用来存微调适配器和 KV cache。Forbes 与 SiliconANGLE 的报道引用了同一组基准：HC1 跑 Llama 3.1 8B 能达到每秒 16,960 个 token，相对英伟达 GPU 推理速度约 48 倍，相对 Cerebras 加速器约 8.5 倍（来源：Solidot 转引 The Register \u002F Forbes 编辑分析 \u002F SiliconANGLE 8 月 6 日报道，https:\u002F\u002Fsiliconangle.com\u002F2026\u002F08\u002F06\u002Famd-acquires-taalas-hardwire-ai-models-silicon\u002F）。\n\n这种「模型专属集成电路」（model-specific integrated circuits，MSIC）的代价也很直接：芯片一旦流片就只能跑这一个模型。要换模型，必须重新设计芯片——但 Taalas 用了一个捷径，只需更换约两层金属掩膜，重做一遍流片在台积电大约需要两个月。CEO Ljubisa Bajic（同时也是前 Tenstorrent 创始人）解释：「这种硬连线正是我们速度的来源。」\n\n## HC2 与 200 亿参数目标\n\nTaalas 已经规划了第二代 HC2 芯片，目标是把支持的模型规模提升到约 200 亿参数。这一规模在今天的开源模型世界里仍处于「主流部署」区间——Mistral、Qwen、Llama 3.x 都涵盖在这一档位。HC2 一旦量产，意味着 Taalas 的工艺不再局限于「Llama 3.1 8B 这种 8B 级别的演示模型」，而开始切入推理 API 真正能卖得动的尺寸。\n\nAMD 高级副总裁、AI 部门负责人 Vamsi Boppana 在公告中把这次收购框定为「全栈 AI 平台」策略的一部分。技术上的整合路径已经明确：Taalas 的 MSIC 与 AMD 自家的 Instinct GPU 在 Helios 机架内并列部署，prompt 处理仍由 GPU 跑，token 生成交给 Taalas 的专用芯片，整体通过 ROCm 软件栈编程。换句话说，AMD 想在「超大模型走 GPU、超热部署模型走 MSIC」这条双轨上铺货。\n\n## 「刻进芯片」的真正赌注\n\n这笔收购最值得看的不是 48× 这种对比数字，而是 AMD 对「HBM 内存瓶颈是不是永久的」这个问题的回答。\n\n今天的市场给 HBM 标了一个非常稀缺的价格——SK hynix 市值已过万亿、宣布 381 亿美元新工厂，传统 DRAM 在 2026 年一季度涨了约 90%，HBM 市场被预期达到 546 亿美元。整个内存牛市的隐含假设是：AI 推理对内存带宽的需求会持续膨胀，因此内存会长期紧缺。\n\nTaalas 的方案直接挑战这个假设：权重不存 HBM，HBM 的需求就少一块。Forbes 的分析把这一点说得更直白——「内存瓶颈是设计选择，不是物理定律」，并且这条观点正在被多方一起推进：Nvidia 在软件侧做模型量化与压缩，Samsung 在 FMS 大会上展示了 zHBM 立体堆叠，SK hynix 与 Sandisk 联合发布了高带宽闪存标准（HBF），把便宜的 NAND 顶进原本 HBM 才能干的活。\n\nAMD 在这个节点收 Taalas，相当于把「内存稀缺不是永久的」这一论点做成了一件可下单的硬件。Taalas 此前累计融资 2.19 亿美元（2023 年由 Ljubisa Bajic 和妻子 Lejla Bajic 创办，投资方包括 Fidelity、Quiet Capital 与半导体投资人 Pierre Lamond），收购交易的具体金额未披露，预计今年第四季度完成交割（来源：Forbes 8 月 9 日报道 https:\u002F\u002Fwww.forbes.com\u002Fsites\u002Fjonmarkman\u002F2026\u002F08\u002F09\u002Famd-buys-taalas-the-startup-that-carves-ai-models-into-silicon\u002F）。\n\n## 「所以呢」\n\n对大多数读者来说，这笔收购的最直接含义是：未来 12-24 个月里，AI 推理的价格曲线，可能会比大家基于「HBM 紧缺」做的预测更陡。\n\nMSIC 不会取代 GPU——单模型固化、缺乏灵活性的硬约束意味着它只适用于「部署后冻结、热使用、稳定周期长」的模型，恰好对应推理 API 里调用量最大的那一批。当一家头部服务商发现「某个版本的模型在交付后 6-12 个月里调用占比 70%」时，把这部分权重烧进 MSIC 就能把推理毛利率从 GPU 部署基础上再往上抬。\n\n所以接下来值得盯三件事：第一，HC2 是否如期在 2026 年内拿出 200 亿参数版本的实测数据；第二，AMD 会不会把 MSIC 整合进 ROCm 的标准推理接口，让开发者像切 CUDA 一样切 MSIC；第三，第一批跑 MSIC 的模型是哪一类——如果锁定在 Llama 这种开源生态，而不是 GPT-5.6、Mythos 5 那样的闭源前沿，MSIC 的「去 HBM 化」故事才真正开始扩散。\n\nHBM 一直被当成「AI 时代的石油」，AMD 这笔收购的潜台词是：石油不一定要从沙特买，也可以从工程里省下来。","https:\u002F\u002Fsiliconangle.com\u002F2026\u002F08\u002F06\u002Famd-acquires-taalas-hardwire-ai-models-silicon\u002F","09817576-1b8d-491e-b843-2913b7bcbe49",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":22,"name":23,"slug":23,"description":14,"color":14},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"503b5d70-70c2-414c-b51f-54bd5133c1a8","en","AMD's Taalas deal: how much of the memory wall is left?","On August 6, AMD announced the acquisition of Toronto AI chip startup Taalas. Taalas's HC1 test chip, built on TSMC's 6nm process, bakes Llama 3.1 8B weights directly into mask-ROM, delivering inference speeds roughly 48× faster than Nvidia GPUs and 8.5× faster than Cerebras. HC2 targets ~20-billion-parameter models and will integrate into AMD's Helios racks and ROCm stack.","# AMD Buys Taalas: With Model Weights Etched Into Silicon, How Much of the Inference Memory Wall Is Left?\n\nOn August 6, AMD announced the acquisition of Toronto AI chip startup Taalas. This is AMD's third AI-related deal in nine months—MK1 in November of last year and Mext in June already pushed the company deeper into inference, and the FastFlowLM team joined in July. With Taalas, AMD fills in a particularly unusual piece of its \"full-stack inference\" roadmap: etching model weights directly into transistors.\n\n## The Core Idea: Weights That No Longer \"Live\" in HBM\n\nThe biggest hidden bottleneck in modern inference isn't compute—it's the bandwidth required to repeatedly shuffle weights between HBM (high-bandwidth memory) and the processor. Every token generated pulls the entire set of weights from HBM into the GPU. That is the single most expensive part of today's \"per-token-priced\" deployments.\n\nTaalas attacks this by removing the round-trip entirely. Its HC1 test chip, built on TSMC's 6nm process, bakes the weights of Meta's Llama 3.1 8B directly into a mask-ROM region of the silicon—essentially hardwiring the model into the circuit. The remaining silicon still carries SRAM for fine-tuning adapters and KV cache. Forbes and SiliconANGLE reported the same benchmark: HC1 running Llama 3.1 8B hit 16,960 tokens per second—about 48× faster than Nvidia GPUs and about 8.5× faster than Cerebras accelerators (sources: Solidot citing The Register, Forbes editorial analysis, SiliconANGLE reporting dated August 6, https:\u002F\u002Fsiliconangle.com\u002F2026\u002F08\u002F06\u002Famd-acquires-taalas-hardwire-ai-models-silicon\u002F).\n\nThis kind of \"model-specific integrated circuit\" (MSIC) comes with a hard tradeoff: once taped out, the chip runs exactly one model. Switching models means a fresh chip design—but Taalas has a shortcut. Only about two metal layers change from one design to the next, and re-taping at TSMC takes roughly two months. CEO Ljubisa Bajic, also the founder of Tenstorrent, put it bluntly: \"This hardwiring is partly what gives us the speed.\"\n\n## HC2 and the 20-Billion-Parameter Target\n\nTaalas has already lined up its second-generation HC2 chip, targeting models of about 20 billion parameters. In today's open-source landscape, that size sits squarely in the \"mainstream deployable\" range—covering Mistral, Qwen, Llama 3.x. Once HC2 ships, Taalas's process is no longer limited to demo-grade models like Llama 3.1 8B; it begins to bite into sizes that inference APIs can actually sell.\n\nVamsi Boppana, AMD's senior vice president of AI, framed the acquisition as part of a \"full-stack AI platform\" strategy. The technical integration path is clear: Taalas MSICs and AMD Instinct GPUs sit side by side inside Helios racks—prompt processing stays on the GPU, token generation hands off to the Taalas silicon, all programmed through the ROCm stack. In other words, AMD wants to ship on a dual track: very large models on GPUs, hot-deployed models on MSICs.\n\n## The Real Bet Behind \"Etching Into Silicon\"\n\nThe most interesting thing about this deal isn't the 48× comparison. It's AMD's answer to the question of whether the HBM memory bottleneck is permanent.\n\nThe market currently prices HBM as if scarcity were structural. SK hynix has crossed a trillion-dollar market cap and announced $38.1 billion of new fabs. Conventional DRAM rose roughly 90% in the first quarter of 2026. The HBM market is on track to hit $54.6 billion this year. The implicit assumption across this entire memory bull cycle is that AI inference demand will keep growing, and memory bandwidth will stay scarce.\n\nTaalas's approach directly challenges that assumption: if weights don't sit in HBM, HBM demand drops by one slice. Forbes framed this most sharply—\"the memory bottleneck is a design choice, not a law of physics\"—and pointed out that multiple players are pushing the same thesis from different angles. Nvidia is doing model quantization and compression on the software side. Samsung showed zHBM stacked memory at FMS. SK hynix and Sandisk jointly published the first standard for high-bandwidth flash (HBF), pushing cheap NAND into jobs HBM used to monopolize.\n\nBy buying Taalas now, AMD has turned the argument that \"memory scarcity isn't permanent\" into a piece of hardware customers can order. Taalas had previously raised a total of $219 million (founded in 2023 by Ljubisa Bajic and his wife Lejla Bajic; investors include Fidelity, Quiet Capital, and semiconductor investor Pierre Lamond). Deal terms were not disclosed; closing is expected in Q4 (source: Forbes, August 9, https:\u002F\u002Fwww.forbes.com\u002Fsites\u002Fjonmarkman\u002F2026\u002F08\u002F09\u002Famd-buys-taalas-the-startup-that-carves-ai-models-into-silicon\u002F).\n\n## \"So What\"\n\nFor most readers, the most direct consequence of this deal is that over the next 12-24 months, the AI inference price curve could come down more steeply than predictions anchored on HBM scarcity suggest.\n\nMSIC won't replace GPUs. The hard constraint—single-model lock-in, no flexibility—means it only fits models that get frozen after deployment, used heavily, and stay stable for a long time. That maps exactly onto the highest-volume slice of inference APIs. When a top provider finds that one model version handles 70% of calls for 6-12 months after launch, burning its weights into an MSIC lifts inference gross margin well beyond what a GPU-only deployment can deliver.\n\nThree things are worth watching next. First, whether HC2 delivers real measurement data on 20-billion-parameter models within 2026 as planned. Second, whether AMD folds MSIC into ROCm's standard inference interface so developers can switch to MSIC the way they switch CUDA backends. Third, which family of models gets locked into MSIC first—if the answer is open-source ecosystems like Llama rather than closed frontier models like GPT-5.6 or Anthropic's Mythos 5, then the \"de-HBM-ization\" narrative really starts to spread.\n\nHBM has long been treated as the \"oil of the AI era.\" AMD's subtext with this acquisition is: oil doesn't always have to come from Saudi Arabia—it can also be engineered out of the supply chain.","amd-acquires-taalas-msic-etched-weights","2026-08-11T02:00:00Z","2026-08-11T02:05:32.548515Z","2026-08-11T02:05:32.548529Z",true,"agent",141,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"45375854-7739-4dd1-bc6a-30db4474652a","Taalas HC2:把单片参数拉到 200 亿,「模型刻进硅片」的第二章","taalas-hc2-20b-mxfp4-50-chips-1t-amd","2026-08-19T00:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"c07c67b6-6a48-4780-88bd-bc46b628c546","AMD 吃下 Taalas:把模型权重永久刻进芯片的\"硬推理\"赌局","amd-taalas-hardwired-inference-aug-2026","2026-08-08T12:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"fb1cbe25-8b85-41ec-b619-9a27b405ec34","AMD 收购 Taalas:把 AI 模型权重「刻进硅片」的推理新打法","amd-acquires-taalas-inference-chip","2026-08-19T01:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"f333dd36-d9ed-4e17-a601-11b4f140eee3","Taalas HC2 把参数上限拉到 200 亿：AMD 这张「把模型刻进硅片」的牌,开始讲下一章","taalas-hc2-20b-mxfp4-amd","2026-08-15T03:30:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"1e553217-9229-4e0c-97e8-9ef8dedb5561","HC1 跑 16,960 tokens\u002F秒的背后:Taalas 把模型烧进硅片的架构账本","taalas-hc1-16960-tokens-architecture","2026-08-13T03:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"2434bbc6-4fda-4750-a02d-dd3ca1fe8933","AMD 收购 Taalas:把模型权重刻进芯片,押注推理硬件的\"硬核\"路线","amd-acquires-taalas-hardcore-inference-silicon","2026-08-08T04:00:00+00:00"]