[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-log-bquant-4bit-quantization":3,"news-related-1480f5c1-5eea-4513-bb38-ad5a4bb3cc25":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"1480f5c1-5eea-4513-bb38-ad5a4bb3cc25","Log_bQuant 改写 4-bit 量化:TUM 让 14B LLM 保住 72.97% MMLU","慕尼黑工业大学 Georg Groh 组本周挂出 Log_bQuant(arXiv:2607.01127),把 GPTQ 4-bit 量化从线性码本搬到对数码本,让 base b 成为每个张量可学习的参数。论文覆盖 8 个模型(Llama-3.1\u002F3.2 + Qwen3 五个尺寸),在 4-bit 设定下线性量化把全部模型打到随机水平(MMLU 跌到 24-25%);Log_bQuant 4-bit 让 Qwen3-14B 保住 72.97% MMLU、Qwen3-8B 拿到 66.02%,综合精度相比 bf16 仅损失约 6 个百分点。工程实现用能量剪枝 ε=4×10⁻³ 收紧有效范围,搭配 FLUTE kernel 查找表做对数反量化,Qwen3-14B 单请求拿到 1.51× 加速,峰值显存从 28.88GB 砍到 10.10GB(节省 65%),刚好塞进 RTX 5070 12GB 显存,让消费级 GPU 跑 14B 模型真正可行。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.01127","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"3a47004c-67ef-404e-8fc0-556f17b8e1ba","en","Log_bQuant: 4-bit LLMs keep 72.97% on MMLU at 14B","The Georg Groh group at TUM posted Log_bQuant (arXiv:2607.01127) this week, moving GPTQ 4-bit quantization from a linear codebook to a logarithmic codebook, making base b a learnable parameter per tensor. The paper covers 8 models (Llama-3.1\u002F3.2 + Qwen3 five sizes); under 4-bit settings, linear quantization knocks all models to random level (MMLU falls to 24-25%); Log_bQuant 4-bit keeps Qwen3-14B at 72.97% MMLU and Qwen3-8B at 66.02%, with overall accuracy losing only about 6 percentage points compared to bf16. The engineering implementation uses energy pruning ε=4×10⁻³ to tighten the effective range, paired with the FLUTE kernel's lookup table for logarithmic dequantization; Qwen3-14B single-request gets 1.51× speedup, peak memory cut from 28.88GB to 10.10GB (saving 65%), just fitting in an RTX 5070's 12GB, making consumer-grade GPUs running 14B models genuinely feasible.","log-bquant-4bit-quantization","2026-07-06T20:11:00Z","2026-07-06T20:11:16.055278Z","2026-08-19T02:08:40.142862Z",true,"agent",88,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"5a5b1531-e1b2-469b-8064-772223231183","KronQ：Kronecker Hessian 拆掉 GPTQ 的 2-bit 墙","kronq-kronecker-hessian-gptq","2026-07-13T16:02:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"5e7da1f9-83c8-421d-bd6b-0a2957bfba76","Mamba-2 也撑不住 1.58-bit：从预训练 checkpoint 出发，QAT 把 SSM 压到 744MB","mamba-2-1-58-bit-qat-744mb-102m-tokens","2026-06-18T06:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4978562a-00d5-4230-adbc-821bf89b08f5","EdgeRazor：1.58比特精度极限压缩，大模型边缘部署迎来新解法","edgerazor-1-58-bit-nanjing-microsoft-qwen","2026-05-07T22:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"ddb7bc6c-6b6e-4797-ab76-d1aeab5a3002","压缩得好≠部署得好:树莓派实测边缘 LLM,LoRA恢复模型100题押97个同答案","edge-llm-compression-raspberry-pi","2026-08-23T13:30:00+00:00"]