[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-amd-acquires-taalas-hardcore-asic-inference":3,"news-related-c26cb1e1-d0c0-471d-81a1-79536834a617":35},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":21,"news_slug":28,"published_at":29,"created_at":30,"modified_at":31,"is_published":32,"publish_type":33,"image_url":14,"view_count":34},"c26cb1e1-d0c0-471d-81a1-79536834a617","AMD 收下 Taalas：把 Llama 权重烧进 ASIC，推理速度把 GPU 甩在身后","AMD 8 月 6 日宣布收购多伦多 AI 推理芯片初创公司 Taalas。Taalas 的 Hardcore (HC) ASIC 把单个模型权重直接烧进硬件，与 Instinct GPU 协作可针对 LLM 解码阶段提速。首代 HC1 在 Llama 3.1 8B 上实现每秒数千 token 的推理速度，TDP 远低于 GPU 整机柜。","# AMD 收下 Taalas：把 Llama 权重烧进 ASIC，推理速度把 GPU 甩在身后\n\n8 月 6 日，AMD 官宣收购位于多伦多的 AI 推理芯片初创公司 Taalas，把这家“把模型权重烧进硅片”的玩家正式纳入自己旗下 [^1]。这笔交易由 AMD 高级副总裁、Vamsi Boppana 的 AI Group 主导，Taalas 联合创始人、CEO Ljubisa Bajic（前 Tenstorrent CEO、前 AMD 高管）和整个加拿大团队将并入 AMD。\n\n## 技术核心：HC 路线不是 GPU，也不是普通 ASIC\n\nTaalas 2023 年才在加拿大成立，2026 年 2 月才出 stealth 模式 [^2]。他们的核心思路被外媒 EE Times 形象地称为“把 AI 模型本身当成一台计算机” [^3]：让芯片的硬件数据流围绕一个特定模型的计算图来设计，并把权重烧进金属层 — 不是存在 HBM 里，是直接做成互连线路的一部分。换句话说：\n\n- **不是 GPU** — 不接受重新编程，硬件只为某一个模型服务；\n- **不是普通 ASIC** — 谷歌 TPU、AWS Trainium 仍能在软件层面做编译，HC1 把那一步也砍掉了；\n- **不是 MRAM \u002F 浮点优化** — 全 SRAM，单模型全部驻片。\n\n首代 HC1 用台积电 6nm 工艺，die size 815mm²、晶体管数 53B、整张卡功耗 2.5 kW。HC1 在 Llama 3.1 8B 上能输出约 17,000 tokens \u002F 秒 \u002F 用户 [^4]。对比基线是 Nvidia H200。模型更新时不需要从零流片 — 只改两层金属（金属层之上承载权重与数据流），官方说法是“两个月，不是两年”就能出新版本。\n\n## AMD 想怎么用：推理生意上补齐“机柜视角”\n\nAMD 在自己官方通稿里写得相当克制，只提 “进一步差异化 AI 路线图”“为 AI 推理市场提供加速计算解决方案” [^1]。但 EE Times 给出了一种更激进的解读：HC 类结构化 ASIC 在 LLM 解码（decode，token-by-token 生成）这种吃带宽、吃延迟的阶段非常合适。和英伟达收购 Groq 之后用 Groq 处理 decode 的方案形成镜像 [^3]：\n\n- AMD 已经宣布 Instinct GPU 与 Cerebras 在“分离式推理”上的合作（GPU 干 prefill，Cerebras 干 decode）；\n- Taalas 加入后，AMD 自家就能在 decode 阶段补一个比 Cerebras 更省电的选项。\n\nEE Times 进一步推测，HC 系列会面向两个场景 [^3]：\n\n1. **物理 AI、边缘推理** — 小模型（≤ 8B）受益于 HC 芯片“单价低、功耗低、不经常换模型”的特点，对应 AMD 已经做了一段时间的 FPGA \u002F SoC 客户群；\n2. **大型模型** — 通过多芯片堆叠（30 颗 HC 跑 DeepSeek-671B 级别的模型）覆盖更上层的推理负载。\n\n## 算力经济学：拷问 GPU 模型的“丹尼尔斯曲线”\n\nForbes 专栏作者 Karl Freund 在 2026 年 2 月 Taalas 出 stealth 时就替读者算过一笔账 [^2]：\n\n| 模型 | HC1 推理单价 (per 1M tokens) | GPU 现役方案 (per 1M tokens) |\n|------|-----------|-----------|\n| Llama 3.1 8B | 0.75 美分 | 3.79 美分 |\n| DeepSeek R1 | 7.6 美分 (sim) | 20–49 美分 |\n\n搭配功耗（HC1 机柜 12–15 kW vs GPU 机柜 120–600 kW），Taalas 阵营宣称“在四年生命周期内 capex 下降 60–75%” [^2]。AMD 通稿没有宣称这些数字，但收购落地，意味着 AMD 内部至少愿意把这些算式放在谈判桌上。\n\n## 我的看法：这是一笔“打补丁”而非“改路线”的并购\n\n几个值得注意的细节：\n\n1. **Taalas 不解决训练问题**。这是纯推理一侧的补强，对 AMD MI400 \u002F MI450 的训练路线图没有直接影响。\n2. **“机柜级推理”阵营正在成形**。从 Nvidia 吃下 Groq、AMD 这次吃下 Taalas、再加上与 Cerebras 的合作，2026 年下半年起 AI 推理硬件会明显从“一张卡即一切”转向“机柜级特化分工”。\n3. **风险在 Taalas 这边**。HC 路线的代价是“换模型就要重新流片”，团队必须把两个月的迭代能力带走，对台积电的依赖是 TSO。AMD 强大的工程资源能扩产能，但模型节奏（每年 1–2 次迭代） vs 硬件节奏（流片 12–18 个月）的错位仍是结构性挑战。\n4. **对小模型阵营是个利好**。Physicial AI、agent、edge 设备这些“推理频繁、模型相对小且稳”的场景，未来 12 个月可能看到更多“ASIC + 标准化模型栈”的耦合方案。\n\n对从业者来说，关注点很简单：AMD 与 Taalas 团队的整合节奏，以及第一批 HC-as-decode 的 Instinct 配套机柜什么时候进入主流云厂。Forbes 文章里那句 “May you live in interesting times” [^2]，今年 AI 推理侧的从业者可能引用得比往年更频繁。\n\n---\n\n[^1]: AMD IR Press Release, “AMD Acquires Taalas to Advance Compute Solutions for Rapidly Growing AI Inference Market”, 2026-08-06. https:\u002F\u002Fir.amd.com\u002Fnews-events\u002Fpress-releases\u002Fdetail\u002F1296\u002Famd-acquires-taalas-to-advance-compute-solutions-for-rapidly-growing-ai-inference-market\n[^2]: Karl Freund, “Taalas Launches Hardcore Chip With ‘Insane’ AI Inference Performance”, Forbes, 2026-02-19. https:\u002F\u002Fwww.forbes.com\u002Fsites\u002Fkarlfreund\u002F2026\u002F02\u002F19\u002Ftaalas-launches-hardcore-chip-with-insane-ai-inference-performance\u002F\n[^3]: Sally Ward-Foxton, “AI Chip Startup Taalas Acquired by AMD”, EE Times, 2026-08-06. https:\u002F\u002Fwww.eetimes.com\u002Fai-chip-startup-taalas-acquired-by-amd\u002F\n[^4]: Taalas official Products page, “HC1 Technology Demonstrator”. https:\u002F\u002Ftaalas.com\u002Fproducts\u002F","https:\u002F\u002Fwww.eetimes.com\u002Fai-chip-startup-taalas-acquired-by-amd\u002F","09817576-1b8d-491e-b843-2913b7bcbe49",[11,15,18],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",[22],{"id":23,"lang":24,"title":25,"summary":26,"content":27},"eeaabd19-8922-4601-9a92-92fb0d731912","en","AMD-Taalas: Llama weights burned into ASICs outrun GPUs","On August 6 AMD announced the acquisition of Toronto-based AI inference chip startup Taalas. Taalas' Hardcore (HC) ASIC hardwires a single model's weights into silicon and is intended to pair with Instinct GPUs to accelerate the decode stage of LLM inference. The first-gen HC1 delivers thousands of tokens per second on Llama 3.1 8B, with total power far below a GPU rack.","# AMD Buys Taalas: Burning Llama Weights Into ASIC, Outrunning the GPU on Inference\n\nOn August 6 AMD officially announced the acquisition of Toronto-based AI inference chip startup Taalas, folding a company that bakes model weights directly into silicon into the larger chipmaker [^1]. The deal is run out of AMD SVP Vamsi Boppana's AI Group; Taalas co-founder and CEO Ljubisa Bajic (ex-Tenstorrent CEO, ex-AMD executive) and the Canadian team will join AMD.\n\n## The technical core: HC is not a GPU, and not a conventional ASIC either\n\nTaalas was founded in Canada in 2023 and only came out of stealth in February 2026 [^2]. The core idea is what EE Times vividly calls \"making the AI model itself the computer\" [^3]: the chip's hardware dataflow is designed around a specific model's compute graph, and the weights are burned into the metal layers — not stored in HBM, but literally mapped into interconnect traces. Concretely:\n\n- **Not a GPU**: HC silicon is not reprogrammable; the hardware serves one model only.\n- **Not a conventional ASIC**: Google TPU and AWS Trainium still go through software compilation. HC1 cuts that stage out too.\n- **Not MRAM or floating-point trickery**: it is fully SRAM, with one model entirely on-die.\n\nThe first-generation HC1 is built on TSMC 6nm, with an 815 mm² die, 53B transistors, and a 2.5 kW per-card power envelope. On Llama 3.1 8B it delivers roughly 17,000 tokens\u002Fsecond\u002Fuser [^4], with the baseline being Nvidia H200. When the model updates, Taalas does not need a fresh full tape-out — only the two metal layers (which encode weights and dataflow) change. Their stated turnaround is \"two months, not two years.\"\n\n## How AMD plans to use it: filling in the rack-scale view of inference\n\nAMD's own press release stays measured, mentioning only \"further differentiating the AI roadmap\" and delivering accelerated compute for the AI inference market [^1]. EE Times offers a more aggressive read: structured-ASIC designs like HC are a natural fit for LLM decode (the token-by-token generation stage that bleeds bandwidth and latency). The setup mirrors what Nvidia now runs after bringing Groq in-house to handle decode [^3]:\n\n- AMD has already announced an Instinct GPU + Cerebras partnership for disaggregated inference (GPU does prefill, Cerebras does decode);\n- With Taalas inside, AMD gains an in-house decode option that is more power-efficient than Cerebras.\n\nEE Times further suggests two scenarios where the HC line will land first [^3]:\n\n1. **Physical AI and edge inference** — small models (≤ 8B) benefit from HC's low unit cost, low power draw, and infrequent model swaps, lining up with AMD's existing FPGA\u002FSoC customer base.\n2. **Large models** — multi-chip stacking (think ~30 HC chips to serve a DeepSeek-671B-class model) covers higher-end inference workloads.\n\n## Inference economics: stress-testing the GPU cost curve\n\nForbes columnist Karl Freund ran the numbers back when Taalas came out of stealth in February 2026 [^2]:\n\n| Model | HC1 cost per 1M tokens | Current GPU cost per 1M tokens |\n|------|------|------|\n| Llama 3.1 8B | \u002Fbin\u002Fbash.0075 | \u002Fbin\u002Fbash.0379 |\n| DeepSeek R1 | \u002Fbin\u002Fbash.076 (sim) | \u002Fbin\u002Fbash.20–\u002Fbin\u002Fbash.49 |\n\nPair that with power (HC racks draw 12–15 kW versus 120–600 kW for a GPU rack) and the Taalas side claims \"60–75% capex reduction over a four-year comparable lifespan\" [^2]. AMD's own release does not repeat these numbers, but the fact that the deal closed implies AMD is willing to put those calculations on the table in customer conversations.\n\n## My take: a \"patch,\" not a \"re-route,\" acquisition\n\nA few details worth flagging:\n\n1. **Taalas does not solve training.** This is a pure-inference reinforcement — it does not change the MI400\u002FMI450 training roadmap.\n2. **Rack-scale inference is becoming a thing.** From Nvidia scooping Groq, AMD buying Taalas on top of its Cerebras partnership, the second half of 2026 will see AI inference hardware move clearly from \"one card does it all\" to \"rack-level specialized division of labor.\"\n3. **The risk sits on the Taalas side.** The HC approach's price is \"re-tape whenever the model changes,\" so the team has to bring its two-month iteration cadence with it and stay close to TSMC. AMD's engineering muscle can scale production, but the structural mismatch between model cadence (1–2 updates per year) and silicon cadence (12–18 months per tape-out) remains.\n4. **Good news for the small-model camp.** Physical AI, agentic applications, and edge devices — workloads that run inference frequently, on relatively small and stable models — are likely to see more \"ASIC + a standardized model stack\" couplings over the next 12 months.\n\nFor practitioners the watch-list is simple: how fast the AMD × Taalas integration moves, and when the first HC-as-decode Instinct rack reaches mainstream cloud providers. As Freund's Forbes piece put it — \"May you live in interesting times\" [^2] — for AI inference practitioners this year, that line is going to get quoted more than usual.\n\n---\n\n[^1]: AMD IR Press Release, \"AMD Acquires Taalas to Advance Compute Solutions for Rapidly Growing AI Inference Market,\" 2026-08-06. https:\u002F\u002Fir.amd.com\u002Fnews-events\u002Fpress-releases\u002Fdetail\u002F1296\u002Famd-acquires-taalas-to-advance-compute-solutions-for-rapidly-growing-ai-inference-market\n[^2]: Karl Freund, \"Taalas Launches Hardcore Chip With 'Insane' AI Inference Performance,\" Forbes, 2026-02-19. https:\u002F\u002Fwww.forbes.com\u002Fsites\u002Fkarlfreund\u002F2026\u002F02\u002F19\u002Ftaalas-launches-hardcore-chip-with-insane-ai-inference-performance\u002F\n[^3]: Sally Ward-Foxton, \"AI Chip Startup Taalas Acquired by AMD,\" EE Times, 2026-08-06. https:\u002F\u002Fwww.eetimes.com\u002Fai-chip-startup-taalas-acquired-by-amd\u002F\n[^4]: Taalas official Products page, \"HC1 Technology Demonstrator.\" https:\u002F\u002Ftaalas.com\u002Fproducts\u002F","amd-acquires-taalas-hardcore-asic-inference","2026-08-17T00:00:00Z","2026-08-17T05:06:01.550789Z","2026-08-17T05:06:01.550799Z",true,"agent",111,{"items":36},[37,42,47,52,57,62],{"id":38,"title":39,"news_slug":40,"published_at":41},"1e553217-9229-4e0c-97e8-9ef8dedb5561","HC1 跑 16,960 tokens\u002F秒的背后:Taalas 把模型烧进硅片的架构账本","taalas-hc1-16960-tokens-architecture","2026-08-13T03:00:00+00:00",{"id":43,"title":44,"news_slug":45,"published_at":46},"c07c67b6-6a48-4780-88bd-bc46b628c546","AMD 吃下 Taalas:把模型权重永久刻进芯片的\"硬推理\"赌局","amd-taalas-hardwired-inference-aug-2026","2026-08-08T12:00:00+00:00",{"id":48,"title":49,"news_slug":50,"published_at":51},"2434bbc6-4fda-4750-a02d-dd3ca1fe8933","AMD 收购 Taalas:把模型权重刻进芯片,押注推理硬件的\"硬核\"路线","amd-acquires-taalas-hardcore-inference-silicon","2026-08-08T04:00:00+00:00",{"id":53,"title":54,"news_slug":55,"published_at":56},"4f957b38-d8f5-446f-805c-062ff268d7ab","三星 GAIA 试水 AI PC：把「存算一体」塞进 NPU，准备用 PIM 抢端侧推理","samsung-gaia-ai-pc-pim","2026-07-11T06:00:00+00:00",{"id":58,"title":59,"news_slug":60,"published_at":61},"9dffd6b9-99bc-448c-90d8-f706b74edcba","Intel Xeon 6+ 登场：288核 E-core 架构能否重塑数据中心推理？","intel-xeon-6-plus-288-e-core-clearwater","2026-06-03T01:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"bfd2a2e5-7c5c-4b92-a48b-d2aca19b11fe","英特尔SuperClaw：混合AI架构如何让边缘设备更聪明","intel-superclaw-hybrid-edge-70pct","2026-05-23T07:00:00+00:00"]