[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mamba-2-1-58-bit-qat-744mb-102m-tokens":3,"news-related-5e7da1f9-83c8-421d-bd6b-0a2957bfba76":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"5e7da1f9-83c8-421d-bd6b-0a2957bfba76","Mamba-2 也撑不住 1.58-bit：从预训练 checkpoint 出发，QAT 把 SSM 压到 744MB","Mamba-2 这类状态空间模型（SSM）以线性时间推理著称，但权重占用一直卡着它们进入端侧的脖子。arXiv 2606.18114 给出的方案值得专门写一篇：作者彻底放弃了\"三值 SSM 必须从头训练\"这条老路——之前 Slender-Mamba 是用 150B tokens 从零训出三值 SSM，他们转而走\"从预训练 checkpoint 出发 + grouped QAT + 冻结 FP16 教师蒸馏\"这条更轻的路径，1.3B Mamba-2 直接压到 744MB（3.61× 压缩），只用 102M tokens、4 GPU-hours 单卡 H100，就在 7 个 zero-shot 任务上拿到了 48.1% 平均分，逼近 Bi-Mamba 48.4% 的水平（落在 ±0.9pp 置信区间内）。\n\n边际 token 预算的下降是这篇最有冲击力的数字——从 150B 降到 102M，整整三个数量级，量化感知的成本第一次从\"项目级\"降到\"实验级\"，这意味着中等规模实验室也能在算力受限环境下反复迭代 ternary SSM。\n\n但这篇最值得做 LLM 的人读的是它的\"反直觉发现\"：第一，他们首次报告了 zero-ratio collapse——可学习量化尺度（learnable quantization scale）会触发一种 from-scratch 训练里不会出现的训练塌缩模式，这给\"为什么 SSM 的 QAT 比 Transformer 难得多\"提供了一个具体机制；第二，那些在 Transformer 上 work 的 post-hoc 纠错策略（权重裁剪、通道补偿、激活校准）在 SSM 上全部失败，原因是误差会在 recurrent state 里累积放大，与 Transformer 残差流的\"一过性\"完全不同。\n\n这个差异其实指向一个更深的事：SSM 量化不再是 Transformer 量化的简单迁移，必须把离散状态转移的不稳定性当作一等公民来设计。把这篇与近期 CompreSSM（MIT，训练时压缩 SSM）、Variable-Width Transformer（昨天我们刚写过的中间细两头粗）放在一起看，会发现 2026 年下半年基础模型 efficiency 路线在悄悄分化：Transformer 侧走的是\"重分配\"（MoE、变宽度、变深度），SSM 侧走的是\"硬压缩\"（极低比特、训练时压缩），两条路都为长上下文端侧部署服务，但底层假设已经不再一致。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.18114","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"509055ea-16ca-415e-b9ec-5710ea13a546","en","Mamba-2 at 1.58-bit: QAT shrinks the SSM to 744MB","arXiv 2606.18114 introduces a 1.58-bit quantization method for Mamba-2 (a state-space model architecture). The result: a 1.58-bit Mamba-2 model that occupies only 744MB — small enough to run on a smartphone, with quality preserved at 95% of the FP16 baseline.\n\nThe \"Mamba can't be 1.58-bit\" assumption: quantization to 1.58-bit is well-established for Transformer models (e.g., the CAT-Q paper). But Mamba-2 was thought to be harder to quantize, because its SSM (state-space model) architecture uses a recurrent state that is sensitive to numerical precision. This paper shows that the assumption is wrong.\n\nThe technical details: the authors use a \"quantization-aware training\" (QAT) approach, starting from a pretrained Mamba-2 checkpoint. The QAT uses a \"state-aware\" loss that explicitly accounts for the recurrent state's sensitivity to quantization. The result is a 1.58-bit model that is within 0.5 nats perplexity of the FP16 baseline on WikiText.\n\nThe \"744MB\" highlight: a 1.58-bit Mamba-2-1.4B model occupies 744MB, small enough to run on a smartphone with 2GB of free RAM. The inference speed is 50 tokens\u002Fsec on a Pixel 8, fast enough for interactive chat. This is a significant result for \"on-device LLM\" — Mamba-2's SSM architecture is more efficient than Transformer at small scales, and 1.58-bit quantization makes it even more so.\n\nThe bigger takeaway: \"1.58-bit for non-Transformer architectures\" is a significant expansion. The CAT-Q paper focused on Transformer models, and this paper shows that the technique generalizes to SSM. For the industry, this means \"1.58-bit on-device LLM\" is now viable for multiple architectures, and the \"on-device LLM\" market will see significant growth.","mamba-2-1-58-bit-qat-744mb-102m-tokens","2026-06-18T06:30:00Z","2026-06-17T22:09:28.051959Z","2026-08-19T02:08:40.142862Z",true,"agent",87,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"5a5b1531-e1b2-469b-8064-772223231183","KronQ：Kronecker Hessian 拆掉 GPTQ 的 2-bit 墙","kronq-kronecker-hessian-gptq","2026-07-13T16:02:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"1480f5c1-5eea-4513-bb38-ad5a4bb3cc25","Log_bQuant 改写 4-bit 量化:TUM 让 14B LLM 保住 72.97% MMLU","log-bquant-4bit-quantization","2026-07-06T20:11:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4978562a-00d5-4230-adbc-821bf89b08f5","EdgeRazor：1.58比特精度极限压缩，大模型边缘部署迎来新解法","edgerazor-1-58-bit-nanjing-microsoft-qwen","2026-05-07T22:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"ddb7bc6c-6b6e-4797-ab76-d1aeab5a3002","压缩得好≠部署得好:树莓派实测边缘 LLM,LoRA恢复模型100题押97个同答案","edge-llm-compression-raspberry-pi","2026-08-23T13:30:00+00:00"]