[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tenstorrent-risc-v-tensix":3,"news-related-39f5dabb-a59e-4672-9caa-446fd6d6b0cd":31},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":24,"published_at":25,"created_at":26,"modified_at":27,"is_published":28,"publish_type":29,"image_url":13,"view_count":30},"39f5dabb-a59e-4672-9caa-446fd6d6b0cd","Tenstorrent 同台刷新三项推理记录：RISC-V + Tensix 把\"GPU = 默认\"撕开一道口子","6 月 30 日 Tenstorrent 在东京 TT-Deploy JP 一次抛出三张推理成绩单：Kimi K2.6 跑到 900 tokens\u002F秒\u002F用户，DeepSeek-R1-0528 671B 稳在 400+ tokens\u002F秒\u002F用户；LTX-2.3 Fast 在 144 帧 1080p、加音频与口型同步的设定下做出 6 秒一段视频。对照官方 GPU 基线，LLM 约 3 倍加速、视频约 4 倍，三组数字都跑在同一套 Tenstorrent Galaxy Blackhole 超集群上，靠标准以太网横向扩展。\n\n底层架构才是看点：RISC-V 控制面 + Tensix AI Core 自研数据通路，调度和内存模型跟 CUDA 系 GPU 不是一个体系。Galaxy 的网络层走标准以太网而不是 NVLink 之类的专有互联，软件栈整体开源，今天等于把\"开放式、可许可、heterogeneous AI compute\"路线第一次用生产级模型背书——跨 MoE LLM 与 Diffusion 视频两类负载，同一套硬件就能承接，且容量随 Galaxy 节点添加近似线性增长。\n\n同场还推出 TT-Ascalon S——面向 agentic AI 的 RISC-V CPU IP，die 面积砍到上代约 50%，每平方毫米性能反升 140%；并通过 Rapidus 的 2nm 项目在日本落地 NEDO 主导的 LSTC 计划。ai& 那边已部署 120+ Galaxy 系统，撑起日本规模最大的 sovereign AI 算力底座，整套推理栈跑在境内、不依赖海外供应链。\n\n\"非 GPU 推理路线\"长期停在 demo 与细分市场，这次 3 倍\u002F4 倍同台登场加国家级落地，意味着 RISC-V + Tensix 已经具备挑战\"GPU = 默认假设\"的工程现实。对所有把推理成本全压在单一 GPU 供应链上的团队，这是一份值得立刻拿来算账的对照表。","https:\u002F\u002Ftenstorrent.com\u002Fnewsroom\u002Ftenstorrent-sets-new-performance-records-launches-tt-ascalon-s-and-expands-across-japan","91111c4f-644d-4a5a-a0bc-90e90a97a889",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[],"tenstorrent-risc-v-tensix","2026-06-30T14:05:00Z","2026-06-30T14:12:37.492995Z","2026-08-19T02:08:40.142862Z",true,"agent",96,{"items":32},[33,38,43,48,53,58],{"id":34,"title":35,"news_slug":36,"published_at":37},"ba96758c-bb89-4336-bb4c-cf3ac0056a90","高通 HBC 架构破内存墙：把 3D 堆叠塞进 LLM 解码路径，token\u002F瓦直接翻 6 倍","qualcomm-hbc-3d-stacking-6x-tokens-per-watt","2026-06-27T12:03:00+00:00",{"id":39,"title":40,"news_slug":41,"published_at":42},"4aa9534a-778e-4cd7-8194-fdf3097249b8","OpenAI Jalapeño Hot Chips 实测:峰值每瓦 1.9×,延迟压到 1 秒","openai-jalapeno-hot-chips-benchmark-2026","2026-08-26T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"45375854-7739-4dd1-bc6a-30db4474652a","Taalas HC2:把单片参数拉到 200 亿,「模型刻进硅片」的第二章","taalas-hc2-20b-mxfp4-50-chips-1t-amd","2026-08-19T00:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"1e553217-9229-4e0c-97e8-9ef8dedb5561","HC1 跑 16,960 tokens\u002F秒的背后:Taalas 把模型烧进硅片的架构账本","taalas-hc1-16960-tokens-architecture","2026-08-13T03:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"dfdc3216-52aa-4a78-9bf5-859affc37d17","AMD 收下 Taalas：把模型权重刻进芯片，推理的内存墙还剩多少？","amd-acquires-taalas-msic-etched-weights","2026-08-11T02:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"c07c67b6-6a48-4780-88bd-bc46b628c546","AMD 吃下 Taalas:把模型权重永久刻进芯片的\"硬推理\"赌局","amd-taalas-hardwired-inference-aug-2026","2026-08-08T12:00:00+00:00"]