[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-polyq-3bit-llm-cpu":3,"news-related-cf7f8e7e-4efc-451e-9595-706b0be911ba":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"cf7f8e7e-4efc-451e-9595-706b0be911ba","PolyQ:把\"3-bit LLM 跑在 CPU\"做成一件可预测的事","最近被 arXiv 接收、将在 ICCAD 2026 上见面的 **PolyQ**,给\"边缘端低 bit LLM 推理\"这道老题提供了一个让人耳目一新的答案:不再纠结\"统一压到 2\u002F3\u002F4 bit\",而是在 **平均 bit 预算固定**的前提下,按通道重要性分配 {2, 3, 4, 8, 16} 比特宽度,再用配套编译器把异构 bit 通道重排成 SIMD\u002FLUT 友好的同质块,**把\"分数比特部署\"从概念变成可落地的 CPU 方案**。\n\n为什么这件事重要?现在的端侧 LLM 推理其实卡在一个隐形墙里:GPU 走 NPU 路线成本和功耗高,CPU 又跑不动稠密 4-bit 以下的模型。PolyQ 走的是更工程化的路线——**量化与编译联合设计**。激活感知的通道级 bit 分配保证精度;编译期 permutation 和跨算子合并,把重排流量压掉 70.8%,并把 layout 正则化彻底挡在运行时之外。\n\n实测上,作者在 Falcon-H1-3B、Llama2-13B、Qwen3-32B 三个量级迥异的模型、Workstation\u002FLaptop\u002FMobile 三类 CPU 上做了端到端验证:3-bit 目标下 perplexity 比前作提升 2.4-32.1%,prefill 延迟与 decode 吞吐随 bit 预算**接近线性**变化,单 token 能量开销相对优化 LUT 后端只多不到 2%。换句话说,\"3-bit 跑 CPU\"不仅能跑,而且**性能\u002F能耗可预测**。\n\n值得关注的还有它的工程取舍:PolyQ 没有去卷\"最低 bit\",而是把\"分数比特\"做成可调旋钮——这恰好和最近 DeepSeek、Kimi 等用 MoE + 路由把\"算力按需分配\"的思路异曲同工。可以预见,**未来 LLM 系统栈的竞争点,会从\"模型本身能压多狠\"转向\"编译器与量化协同能把硬件榨多干\"**。对想自己跑本地 Agent 的中小团队、对手机\u002FPC OEM 来说,这条路线比单纯卷参数更值得押注。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.14618","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"4750e513-bc47-4bbb-b763-4cf8b8e69491","en","PolyQ makes \"3-bit LLM on CPU\" a predictable thing","Recently accepted to arXiv and set to appear at ICCAD 2026, **PolyQ** offers a refreshingly new answer to the old problem of \"low-bit LLM inference on edge devices\": instead of obsessing over \"uniformly compressing to 2\u002F3\u002F4-bit\", it allocates bit widths from {2, 3, 4, 8, 16} across channels by importance under a fixed average-bit budget, then uses a paired compiler to reorganize heterogeneous-bit channels into SIMD\u002FLUT-friendly homogeneous blocks — turning \"fractional-bit deployment\" from a concept into a practical CPU solution. Why does this matter? Current on-device LLM inference is stuck behind an invisible wall: GPUs following the NPU route are too expensive and power-hungry, while CPUs can't run dense models below 4-bit. PolyQ takes a more engineering-driven path — **joint quantization-compiler design**. Activation-aware per-channel bit allocation preserves accuracy; compile-time permutation and cross-operator fusion cut rearrangement traffic by 70.8% and push layout regularization entirely off the runtime path. On three very different model scales (Falcon-H1-3B, Llama2-13B, Qwen3-32B) and three CPU classes (Workstation \u002F Laptop \u002F Mobile), end-to-end validation shows: at a 3-bit target, perplexity improves 2.4–32.1% over prior work; prefill latency and decode throughput scale **near-linearly** with the bit budget; per-token energy cost increases by less than 2% relative to an optimized LUT backend. In other words, \"3-bit on CPU\" is not only viable but **predictable in performance and energy**. The engineering tradeoff is also worth noting: PolyQ doesn't chase \"the lowest bit\", but turns \"fractional bits\" into a tunable knob — echoing the recent idea used by DeepSeek, Kimi, and others with MoE + routing to \"allocate compute on demand\". The next competition in the LLM system stack will shift from \"how hard can the model itself be compressed\" to \"how dry can the compiler-quantization collaboration squeeze the hardware\". For mid-sized and small teams wanting to run local Agents, and for mobile\u002FPC OEMs, this path is more worth betting on than simply scaling parameters.","polyq-3bit-llm-cpu","2026-07-17T10:00:00Z","2026-07-17T10:07:01.042169Z","2026-08-19T02:08:40.142862Z",true,"agent",91,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"4147f71b-eaa7-4d39-91cf-c2c572105e7f","FlashPrefill V2:128K 长文本 prefill 提速 47 倍,块稀疏注意力走进生产框架","flashprefill-v2-block-sparse-prefill","2026-08-21T19:10:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"007a83c4-faed-459a-ab44-17b915e05fb5","TriRoute 把 MoE + MoD + KV 量化做成一个控制器：三条条件计算路径第一次协同","triroute-moe-mod-kv-quantization","2026-07-09T18:02:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"77015cf0-fb2c-4176-ab77-f428d8bd2d30","UltraQuant 把 KV Cache 压到 4-bit：Agentic 长上下文推理首次跑出 3.47× TTFT 加速","ultraquant-amd-4bit-kv-cache-3-47x","2026-06-22T18:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"993c1a22-999d-42de-a202-3a1af5ec7ef8","小米 MiMo × TileRT：万亿模型 1000 tokens\u002Fs，通用 GPU 的极限被重新定义","xiaomi-mimo-tilert-1000-tps-fp4-gpu","2026-06-09T06:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"f33cd4ce-46c8-4ce0-8a8a-090b1359dc34","2-bit 量化翻车实录：Qwen3 推理模型的失败模式与「FP16 规划+循环救援」修复","qwen3-2-bit-fp16-planning-loop-rescue","2026-06-08T02:00:00+00:00"]