[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-kronq-kronecker-hessian-gptq":3,"news-related-5a5b1531-e1b2-469b-8064-772223231183":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"5a5b1531-e1b2-469b-8064-772223231183","KronQ：Kronecker Hessian 拆掉 GPTQ 的 2-bit 墙","USC 的 Donghyun Lee 团队公开了 KronQ——一种基于 Kronecker-Factored Hessian 的训练后量化(PTQ)框架,直击 GPTQ 系列方法在 2-bit 量化上的死结。该论文已被 COLM 2026 接收,代码和模型以 Apache 2.0 开源。\n\nGPTQ、GPTAQ 等二阶方法都把输入激活的 Hessian H_X 当作量化目标,等价假设所有输出通道同等重要。KronQ 的关键改写是,在 K-FAC 近似下把目标分解成 H ≈ H_X ⊗ H_G,把输出侧的曲率 H_G 重新拉回优化目标——让量化真正区分对最终 loss 影响大的输出通道。\n\n方法上两处新意:BiIP(双向非相干处理)在输入、输出两侧各做一次 incoherence 旋转加 rescale,同时压下权重方差;层间混合精度分配用 tr(H_G)·tr(H_X) 作为子层灵敏度指标,把 bit 预算动态切到关键子层,让低比特预算不再是均匀分摊。\n\n数字最有说服力:LLaMA-3-70B 的 2-bit 权重量化上,GPTQ 和 GPTAQ 直接退化到 >2000 的 WikiText-2 困惑度(基本不可用),KronQ 跑到 7.93——同一档位差出三个数量级。LLaMA-2-7B 上 W4\u002FW2 的 PPL 分别为 5.56\u002F8.23(fp16 基线 5.47),几乎不掉点,W4 量化甚至比原始 bf16 还稳。\n\n部署侧,KronQ 走 packaged int4\u002Fint2 + fused dequant+BiIP CUDA matvec 路径,A100 上单 token 解码 6.30 ms(W2\u002FW4 同速),比 fp16 的 11.6 ms 还快——量化等于降速的传统认知被彻底翻了过来。对 LLM 端侧部署和长上下文推理来说,2-bit 已从实验室玩具切到可部署档位,Hugging Face 上已经放出 Llama-2\u002F3 全系 W2\u002FW3\u002FW4 量化模型。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.07964","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c768b53e-d7df-4409-995a-2e019d0e77b4","en","KronQ: Kronecker Hessians break GPTQ's 2-bit wall","USC's Donghyun Lee team has published KronQ — a post-training quantization (PTQ) framework based on Kronecker-factored Hessian, directly hitting the 2-bit quantization dead end of the GPTQ family of methods. The paper has been accepted to COLM 2026, and the code and models are open-sourced under Apache 2.0. Second-order methods like GPTQ and GPTAQ both use the input activation's Hessian H_X as the quantization target, which is equivalent to assuming all output channels are equally important. KronQ's key rewrite: under the K-FAC approximation, decompose the target into H ≈ H_X ⊗ H_G, pulling the output-side curvature H_G back into the optimization objective — letting quantization truly distinguish the output channels that matter most to the final loss. Two new ideas in method: BiIP (Bidirectional Incoherence Processing) does an incoherence rotation plus rescale on both input and output sides, simultaneously suppressing weight variance; mixed-precision allocation between layers uses tr(H_G)·tr(H_X) as the sub-layer sensitivity indicator, dynamically slicing the bit budget to critical sub-layers, so a low-bit budget no longer means even allocation. The numbers are most convincing: on 2-bit weight quantization of LLaMA-3-70B, GPTQ and GPTAQ degrade directly to a WikiText-2 perplexity of >2000 (essentially unusable), while KronQ hits 7.93 — a difference of three orders of magnitude at the same level. On LLaMA-2-7B, W4\u002FW2 PPL is 5.56\u002F8.23 (fp16 baseline 5.47), almost no point loss, and W4 quantization is even more stable than the original bf16. On the deployment side, KronQ takes the packaged int4\u002Fint2 + fused dequant+BiIP CUDA matvec path; on A100, single-token decode is 6.30 ms (W2\u002FW4 same speed), even faster than fp16's 11.6 ms — the traditional notion that \"quantization equals slowing down\" is completely flipped. For on-device LLM deployment and long-context inference, 2-bit has moved from a lab toy to a deployable tier, and Hugging Face has already released Llama-2\u002F3 full-series W2\u002FW3\u002FW4 quantized models.","kronq-kronecker-hessian-gptq","2026-07-13T16:02:00Z","2026-07-13T16:22:35.520495Z","2026-08-19T02:08:40.142862Z",true,"agent",88,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"1480f5c1-5eea-4513-bb38-ad5a4bb3cc25","Log_bQuant 改写 4-bit 量化:TUM 让 14B LLM 保住 72.97% MMLU","log-bquant-4bit-quantization","2026-07-06T20:11:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"5e7da1f9-83c8-421d-bd6b-0a2957bfba76","Mamba-2 也撑不住 1.58-bit：从预训练 checkpoint 出发，QAT 把 SSM 压到 744MB","mamba-2-1-58-bit-qat-744mb-102m-tokens","2026-06-18T06:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4978562a-00d5-4230-adbc-821bf89b08f5","EdgeRazor：1.58比特精度极限压缩，大模型边缘部署迎来新解法","edgerazor-1-58-bit-nanjing-microsoft-qwen","2026-05-07T22:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"ddb7bc6c-6b6e-4797-ab76-d1aeab5a3002","压缩得好≠部署得好:树莓派实测边缘 LLM,LoRA恢复模型100题押97个同答案","edge-llm-compression-raspberry-pi","2026-08-23T13:30:00+00:00"]