[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-amd-mi455x-huggingface-99-5":3,"news-related-cdc8e3ce-b1aa-4348-9436-04763179af9c":39},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":26,"news_slug":32,"published_at":33,"created_at":34,"modified_at":35,"is_published":36,"publish_type":37,"image_url":13,"view_count":38},"cdc8e3ce-b1aa-4348-9436-04763179af9c","AMD MI455X：Transformers 99.5% 通过率，432GB HBM4","AMD 在 Advancing AI 2026 上发布的 Instinct MI455X 单卡搭载 432GB HBM4 与 23.3TB\u002Fs 内存带宽,容量是上一代 MI300(192GB)的 2.25 倍。Hugging Face 拿到早期样机后,在 Transformers 库 24 个核心架构(覆盖 encoder、decoder、视觉、音频、多模态、现代 LLM)的测试中拿到 99.5% 通过率,与 MI300 的 99.4%、NVIDIA A10 的 99.1% 处于同一区间。容量层面,用 64GB 的 Qwen3-32B BF16 模型做并发压测,MI455X 凭借约 3 倍于 MI300 的 KV cache 容量,可支撑的并发请求数也多了约 3 倍。这意味着大模型推理正在告别\"切分优先\"的设计思路,长上下文与高并发场景可以更激进地堆在单机里。工程侧,Hugging Face 与 AMD 联合稳定了 Flash Attention 路径、补齐了 torchcodec 多模态音视频支持,并修复了一组输出对比偏差。下一步,MI455X 将进入 Transformers 的 CI 流水线,AITER 优化内核会被陆续搬到 Hugging Face Kernel Hub。对自托管玩家来说,这不只是\"AMD 追上 NVIDIA\"的叙事,更像是单台机器能跑得动的模型规模又往上抬了一个台阶——以前必须靠张量并行拆开的 64B+ 模型,现在可以放到单机上做高并发推理,运营成本结构会随之改变。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fbadaoui\u002Ftransformers-on-amd-mi455","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20,23],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":21,"name":22,"slug":22,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":24,"name":25,"slug":25,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[27],{"id":28,"lang":29,"title":30,"summary":31,"content":31},"bec4ce03-cff2-4220-bf8f-50364d65af57","en","AMD MI455X: Transformers passes 99.5%, 432GB HBM4","AMD's Instinct MI455X, announced at Advancing AI 2026, ships 432GB of HBM4 and 23.3TB\u002Fs memory bandwidth on a single card, 2.25x the capacity of the previous-generation MI300 (192GB). When Hugging Face got an early sample, it ran it through 24 core architectures in the Transformers library (covering encoder, decoder, vision, audio, multimodal, and modern LLMs) and got a 99.5% pass rate — in the same range as MI300's 99.4% and NVIDIA A10's 99.1%. On the capacity side, using a 64GB Qwen3-32B BF16 model for concurrency stress tests, the MI455X's roughly 3x larger KV cache capacity than MI300 also supports about 3x more concurrent requests. This means LLM inference is moving away from \"split-first\" design — long context and high-concurrency scenarios can be stacked more aggressively onto a single machine. On the engineering side, Hugging Face and AMD jointly stabilized the Flash Attention path, filled in torchcodec multimodal audio\u002Fvideo support, and fixed a set of output-comparison drifts. Next, MI455X will enter Transformers' CI pipeline, and AITER-optimized kernels will be ported to the Hugging Face Kernel Hub. For self-hosters, this isn't just an \"AMD catches up to NVIDIA\" story — it's that the model size you can fit on a single box has moved up another step. 64B+ models that previously had to be sharded via tensor parallelism can now be hosted on a single machine for high-concurrency inference, and the operating-cost structure will change with it.","amd-mi455x-huggingface-99-5","2026-07-27T10:30:00Z","2026-07-27T02:04:36.786846Z","2026-08-19T02:08:40.142862Z",true,"agent",133,{"items":40},[41,46,51,56,61,66],{"id":42,"title":43,"news_slug":44,"published_at":45},"bac63469-b19c-4619-8709-73656ad0cd9f","Intel × HF 上线 xpu-kernels Skill：LLM Agent 把 vLLM 调过的 Triton 内核再提 2.8×","intel-hf-xpu-kernels-skill-triton-2-8x","2026-06-20T00:16:00+00:00",{"id":47,"title":48,"news_slug":49,"published_at":50},"45375854-7739-4dd1-bc6a-30db4474652a","Taalas HC2:把单片参数拉到 200 亿,「模型刻进硅片」的第二章","taalas-hc2-20b-mxfp4-50-chips-1t-amd","2026-08-19T00:00:00+00:00",{"id":52,"title":53,"news_slug":54,"published_at":55},"f333dd36-d9ed-4e17-a601-11b4f140eee3","Taalas HC2 把参数上限拉到 200 亿：AMD 这张「把模型刻进硅片」的牌,开始讲下一章","taalas-hc2-20b-mxfp4-amd","2026-08-15T03:30:00+00:00",{"id":57,"title":58,"news_slug":59,"published_at":60},"a91067a3-4fa4-4e88-a25a-18ba3bea21ea","Google 把\"加密推理\"摆上桌面：HEIR 编译器让预训练模型在密文上直接跑","google-heir-compiler-encrypted-ai-inference","2026-08-14T14:00:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"dfdc3216-52aa-4a78-9bf5-859affc37d17","AMD 收下 Taalas：把模型权重刻进芯片，推理的内存墙还剩多少？","amd-acquires-taalas-msic-etched-weights","2026-08-11T02:00:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"c07c67b6-6a48-4780-88bd-bc46b628c546","AMD 吃下 Taalas:把模型权重永久刻进芯片的\"硬推理\"赌局","amd-taalas-hardwired-inference-aug-2026","2026-08-08T12:00:00+00:00"]