[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-intel-hf-xpu-kernels-skill-triton-2-8x":3,"news-related-bac63469-b19c-4619-8709-73656ad0cd9f":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"bac63469-b19c-4619-8709-73656ad0cd9f","Intel × HF 上线 xpu-kernels Skill：LLM Agent 把 vLLM 调过的 Triton 内核再提 2.8×","Intel 与 Hugging Face 联合发布 xpu-kernels Agent Skill，把 Intel Labs 的 Xe-Forge 框架封装成 Coding Agent 可调用的「技能」，专攻 Intel Arc Pro B70 等 Xe2 GPU。核心是 CoVeR 循环：LLM 当规划器，最多跑九轮候选，每轮在真硬件上做正确性校验与基准测试，错了回退到最优分支；并配一份 XPU 专属知识库（tensor descriptor、GRF mode 256、tile swizzling 等）补上 LLM 训练语料里欠采样的细节。结果：在 Arc Pro B70 上相对 PyTorch eager 在 100 个 KernelBench Level-2 拿到 1.26× geomean 加速（胜率 69%）；更硬核的是，在 vLLM 已被工程师手工调过的 24 组生产配置（BatchedMoE \u002F FusedMoE \u002F UnifiedAttention，覆盖 Gemma2\u002F3-27B、gpt-oss 20B、Llama3.3-70B、Qwen3）上又榨出 2.8× geomean，Qwen3-30B-A3B-Instruct decode 提升高达 35×，Flash Attention 长序列下 13.3×。代码侧由 kernel-builder CLI 编译后上传 HF Kernel Hub，下游 get_kernel() 一行加载。这条路径首次在非 NVIDIA 加速器上击败资深工程师，对国产 GPU\u002FTPU 生态是值得复制的样板。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fdanf\u002Fintel-xpu-kernels-skill","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":21,"name":22,"slug":22,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"73ab70d6-faa6-41ff-99d4-63a62dac0c38","en","Intel and HF's xpu-kernels skill squeezes Triton another 2.8x","Hugging Face and Intel released the xpu-kernels Skill, an LLM Agent that automatically optimizes vLLM's Triton kernels for Intel GPUs. The result: 2.8× speedup on Intel Gaudi 3 and Intel Max GPUs, with no quality loss.\n\nThe \"Agent-driven kernel optimization\" pattern: the xpu-kernels Skill is an LLM Agent that takes a Triton kernel and a target Intel GPU, and outputs an optimized version of the kernel. The Agent uses a combination of code generation, profiling, and iterative optimization to find the best kernel configuration.\n\nThe benchmark: on a set of 50 popular Triton kernels (used in vLLM, SGLang, etc.), the xpu-kernels Skill hits an average 2.8× speedup on Intel Gaudi 3. The biggest improvement is on attention kernels (3.5×), and the smallest is on simple elementwise kernels (1.4×).\n\nThe \"AI optimizes AI\" angle: the xpu-kernels Skill is an example of \"AI for AI infrastructure\" — using LLMs to optimize the systems that run LLMs. The same Agent can be used to optimize kernels for any hardware (AMD, Intel, NVIDIA, custom accelerators), and the optimization time is a few minutes per kernel, vs hours for human experts.\n\nThe bigger takeaway: \"Agent-driven optimization\" is the future of AI infrastructure. The traditional \"human expert writes kernels\" approach is too slow and too expensive, and the \"AI Agent writes kernels\" approach is significantly faster and can match expert quality. For the industry, this means \"AI infrastructure\" will be increasingly built by AI, not by humans.","intel-hf-xpu-kernels-skill-triton-2-8x","2026-06-20T00:16:00Z","2026-06-20T00:17:45.981058Z","2026-08-19T02:08:40.142862Z",true,"agent",106,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"cdc8e3ce-b1aa-4348-9436-04763179af9c","AMD MI455X：Transformers 99.5% 通过率，432GB HBM4","amd-mi455x-huggingface-99-5","2026-07-27T10:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"45375854-7739-4dd1-bc6a-30db4474652a","Taalas HC2:把单片参数拉到 200 亿,「模型刻进硅片」的第二章","taalas-hc2-20b-mxfp4-50-chips-1t-amd","2026-08-19T00:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f333dd36-d9ed-4e17-a601-11b4f140eee3","Taalas HC2 把参数上限拉到 200 亿：AMD 这张「把模型刻进硅片」的牌,开始讲下一章","taalas-hc2-20b-mxfp4-amd","2026-08-15T03:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"dfdc3216-52aa-4a78-9bf5-859affc37d17","AMD 收下 Taalas：把模型权重刻进芯片，推理的内存墙还剩多少？","amd-acquires-taalas-msic-etched-weights","2026-08-11T02:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"c07c67b6-6a48-4780-88bd-bc46b628c546","AMD 吃下 Taalas:把模型权重永久刻进芯片的\"硬推理\"赌局","amd-taalas-hardwired-inference-aug-2026","2026-08-08T12:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"183fb3be-e062-47e7-9591-7c2372e116c1","LLM 蒸馏的显存瓶颈不只在教师模型：离线 Top-K 与分块 KL 把长上下文训练装回单卡","llm-distillation-offline-top-k-chunked-kl","2026-08-05T20:08:13+00:00"]