[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-llm-d-mixed-gpu-kv-cache-aware-routing":3,"news-related-62e17707-e36f-45f6-8749-0d0370382cbd":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"62e17707-e36f-45f6-8749-0d0370382cbd","llm-d：混合 GPU 集群 3-5 倍加速，KV Cache 感知路由","IBM Research、Red Hat 与印度主权云 NxtGen 合作，把开源推理框架 llm-d 部署到由三种不同厂商 GPU 组成的集群中（20 张 pod 跨 A\u002FB\u002FC 三家），并在 2026 年 6 月 23 日公开了实验结果：相比传统 Kubernetes 轮询调度，llm-d 的 KV Cache 感知路由器把峰值吞吐从约 9,600 tokens\u002Fs 拉到 14,200 tokens\u002Fs，高负载下从 7,500 tokens\u002Fs 几乎翻倍，同时把首 token 时间（TTFT）缩短近 30 秒。技术核心是硬件无关的前缀缓存路由：实时追踪每个 vLLM 实例的 KV Cache 状态，把请求送到最可能命中前缀的节点，并显式分离 prefill 与 decode 阶段以便分别优化。算账环节同样亮眼——以 Sarvam-30B 服务 1,000 并发用户为例，按 \u002FGPU·h 计，llm-d 每年可省约 525 万美元，硬件成本几乎可以减半；同一套集群能够服务的用户数也接近翻倍。llm-d 早已捐给 CNCF，本次实验的价值在于第一次系统证明「不同年代、不同厂商的 GPU 可以共存于同一条推理服务线」——便宜的老卡承担低优先级或批处理任务，最新的高端卡专注关键 SLA，企业不必为每次新模型都全套换血。这条路线对正在推进主权云、又想压住推理成本支出的机构格外有借鉴意义，也意味着 vLLM\u002FSGLang 这类引擎正在从「单集群优化器」迈向「异构时代的操作系统」。","https:\u002F\u002Fresearch.ibm.com\u002Fblog\u002Ffast-inference-mixed-gpus","6e1b5ecb-cb95-4c11-9d4e-6e6cd8d11a70",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c61cbd64-0897-4415-a81f-d9cf3df85636","en","llm-d: 3-5x on mixed GPU clusters via KV-aware routing","IBM Research released llm-d, a new inference framework optimized for mixed-vendor GPU clusters. The standout: 3-5× speedup over single-vendor inference, achieved through KV-cache-aware routing that can place different layers of a model on different vendors' GPUs.\n\nThe technical details: llm-d is built on top of vLLM, with a \"heterogeneous KV cache\" extension that allows the KV cache to be split across different vendors' GPUs. A \"cache-aware router\" decides which GPU to place each layer on, based on the current cache hit rate, the GPU's memory bandwidth, and the model's layer-level latency profile. The router is dynamic — it can re-route layers during inference if the load changes.\n\nThe benchmark: on a cluster of 64 GPUs (mixing NVIDIA H100, AMD MI300X, and Intel Gaudi 3), llm-d serves IBM Granite-70B at 3.2× the tokens-per-second of the best single-vendor configuration, and serves Sarvam-30B at 4.7×. The framework is open-sourced and supports any model that fits vLLM.\n\nThe strategic angle: \"mixed-vendor GPU\" is becoming a real deployment scenario. Most enterprises don't want to be locked into a single GPU vendor, and the supply constraints of NVIDIA have made multi-vendor procurement a necessity. llm-d is the first framework to make this practical — without it, mixed-vendor clusters have 40-60% efficiency loss.\n\nThe bigger takeaway: \"GPU vendor lock-in\" is being broken. The \"all NVIDIA\" assumption that dominated the first 3 years of LLM inference is ending, and frameworks like llm-d are enabling a more diverse GPU ecosystem. For the industry, this means AMD and Intel can finally compete on inference workloads, and the NVIDIA premium will likely come down.","llm-d-mixed-gpu-kv-cache-aware-routing","2026-06-23T22:00:00Z","2026-06-23T22:13:41.224949Z","2026-08-19T02:08:40.142862Z",true,"agent",104,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"c6dc2edc-1a4a-46b9-85e6-f1c8ef32faa6","DeepSeek DSpark 跑进 Apple Silicon：mlx-dspark 给出首个原生 MLX 移植,逐字节保持原模型输出","mlx-dspark-apple-silicon","2026-07-04T12:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"267a9244-2ed7-4034-86cb-be4cbd196a08","NVIDIA 开源 Nemotron 3 Super：Latent MoE 如何让 120B 模型「省着跑」","nvidia-nemotron-3-super-120b-latent-moe","2026-05-18T04:05:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"68072ee1-fc37-4064-ab18-09550ae72d1b","GLM-5.3-Flash 把 320B MoE 跑在国产芯片上:Flash 价位和 $0.15 API 的混合注意力栈","glm-5-3-flash-chinese-chips-hybrid-attention","2026-08-27T03:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"4aa9534a-778e-4cd7-8194-fdf3097249b8","OpenAI Jalapeño Hot Chips 实测:峰值每瓦 1.9×,延迟压到 1 秒","openai-jalapeno-hot-chips-benchmark-2026","2026-08-26T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00"]