[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-edgerazor-1-58-bit-nanjing-microsoft-qwen":3,"news-related-4978562a-00d5-4230-adbc-821bf89b08f5":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"4978562a-00d5-4230-adbc-821bf89b08f5","EdgeRazor：1.58比特精度极限压缩，大模型边缘部署迎来新解法","当大模型的参数从十亿级膨胀到千亿级，能不能在手机上跑就成了一个越来越难回答的问题。南京大学与微软AI联合团队近日发布的EdgeRazor框架，给出了一个让人眼前一亮的答案——在仅1.58比特精度下，Qwen3-0.6B的存储空间从1.41GB压缩至280MB，解码速度提升15倍，同时保持了可用的推理能力。\n\n这项研究的核心挑战在于，现有的三种主流量化方法各有局限：后训练量化在4比特以下性能急剧下降，量化感知训练需要动用大量计算资源进行梯度更新，而量化感知蒸馏虽然在二者之间取得了平衡，但在层级别的特征选择上仍然依赖人工干预。\n\nEdgeRazor通过三个创新模块突破了这些瓶颈。混合精度量化感知蒸馏允许对矩阵各层进行精细的精度分配，将更多比特分配给敏感层以保留关键信息。自适应特征蒸馏则让压缩后的学生模型能够智能地从教师模型中选择最具代表性的层进行监督学习。熵感知KL散度进一步扩展了蒸馏数据的使用范围，使模型能够在人工标注和模型生成数据上都能保持稳定的训练效果。\n\n实验结果表明，EdgeRazor在1.88比特精度下就已经超越了所有3比特精度级别的竞争对手，在14个不同领域任务上的表现比最先进的2比特后训练量化方法高出11个百分点。这一结果的实际意义在于：真正能在消费级硬件上运行、同时保持可用能力的LLM，在边缘端部署已不再是奢望。\n\n这背后的核心洞见是：极致压缩不是单一技术的突破，而是量化与蒸馏的深度融合。混合精度分配本质上是对不同权重组的信息瓶颈进行精细调节——对敏感层多给比特，对冗余层果断丢弃。南京大学与微软AI的这项工作，为极端低比特量化提供了一个值得追踪的新方向。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.04062","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d5a4003d-f952-4d97-bcf3-847a85072750","en","EdgeRazor: extreme 1.58-bit compression for edge deployment","As large model parameters balloon from billions to trillions, can they run on phones is becoming an increasingly hard question to answer. A joint team from Nanjing University and Microsoft AI recently released the EdgeRazor framework, providing a striking answer — at just 1.58-bit precision, Qwen3-0.6B's storage compresses from 1.41GB to 280MB, decoding speed improves 15×, while maintaining usable inference capability.\n\nThe core challenge of this research: existing three mainstream quantization methods each have limitations — post-training quantization degrades sharply below 4-bit, quantization-aware training requires massive compute for gradient updates, and while quantization-aware distillation strikes a balance between the two, it still relies on manual intervention for layer-level feature selection.\n\nEdgeRazor breaks through these bottlenecks with three innovative modules. Mixed-precision quantization-aware distillation allows fine-grained precision allocation across matrix layers, allocating more bits to sensitive layers to preserve key information. Adaptive feature distillation lets the compressed student model intelligently select the most representative layers from the teacher model for supervised learning. Entropy-aware KL divergence further extends the use of distillation data, letting the model maintain stable training on both manually-labeled and model-generated data.\n\nExperimental results show EdgeRazor at 1.88-bit precision already surpasses all 3-bit precision competitors, with performance 11 percentage points higher than the most advanced 2-bit post-training quantization methods across 14 different domain tasks. The practical significance: truly runnable on consumer-grade hardware while maintaining usable LLM capability, edge deployment is no longer a luxury.\n\nThe core insight: extreme compression is not a single-tech breakthrough, but a deep fusion of quantization and distillation. Mixed-precision allocation is essentially fine-grained regulation of the information bottleneck for different weight groups — more bits for sensitive layers, decisive drops for redundant ones. This work from Nanjing University and Microsoft AI provides a direction worth tracking for extreme low-bit quantization.","edgerazor-1-58-bit-nanjing-microsoft-qwen","2026-05-07T22:00:00Z","2026-05-07T22:06:12.978373Z","2026-08-19T02:08:40.142862Z",true,"agent",218,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"5a5b1531-e1b2-469b-8064-772223231183","KronQ：Kronecker Hessian 拆掉 GPTQ 的 2-bit 墙","kronq-kronecker-hessian-gptq","2026-07-13T16:02:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"1480f5c1-5eea-4513-bb38-ad5a4bb3cc25","Log_bQuant 改写 4-bit 量化:TUM 让 14B LLM 保住 72.97% MMLU","log-bquant-4bit-quantization","2026-07-06T20:11:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"5e7da1f9-83c8-421d-bd6b-0a2957bfba76","Mamba-2 也撑不住 1.58-bit：从预训练 checkpoint 出发，QAT 把 SSM 压到 744MB","mamba-2-1-58-bit-qat-744mb-102m-tokens","2026-06-18T06:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"ddb7bc6c-6b6e-4797-ab76-d1aeab5a3002","压缩得好≠部署得好:树莓派实测边缘 LLM,LoRA恢复模型100题押97个同答案","edge-llm-compression-raspberry-pi","2026-08-23T13:30:00+00:00"]