[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-cat-q-1-58-bit-512-samples-icml-oral":3,"news-related-93981299-c75e-4140-ad2c-ae68fb5d27ce":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"93981299-c75e-4140-ad2c-ae68fb5d27ce","CAT-Q：512 样本把 235B LLM 压到 1.58-bit，成本降 10 万倍","Intel 中国 AI 团队放出的 CAT-Q 把 1.58-bit 量化从「千亿 token 训练」直接拉进「512 条样本后训练」时代。arXiv 2606.26650 论文已被 ICML 2026 接收为 Oral，作者来自 Intel China AI 实验室。CAT-Q 瞄准现有 ternary PTQ 必须依赖昂贵 QAT 才能维持精度的痛点，提出两大耦合组件：可学习调制 (LM) 在量化前先用一组可学习因子把权重分布和阈值「预拉」到三值友好的形态，软化三值化 (ST) 用可微过渡函数引导三值化过程稳定收敛。实验里它只用 512 条校准样本就能把 1.7B-8B 主流 LLM 量化到优于 BitNet b1.58 v1\u002Fv2 (用 100B token 训练) 的水平，相当于把训练 token 量减少约 10 万倍；更狠的是 14B-235B 模型第一次在 8 张 A100-80GB 上 8-60 小时就能完成三值化，让「一颗 GPU 跑百 B」成为现实。配套工具 BitTern (Apache-2.0) 已经在 GitHub 开源 (IntelChina-AI\u002FBitTern)，目标是把 1.58-bit 模型的门槛降到人人都能玩的程度。从行业视角看，CAT-Q 验证了一个长期被低估的事实：极低比特的关键不在「训出新的小模型」，而在「如何把已有大模型用最小成本压下去」。一旦 100K 倍成本下降普及，边缘推理、本地私部署、Agent 长上下文 KV 压缩这些场景都会被重写一遍——而 BitNet 生态的最大护城河，也就是「必须重新训练」这一刻板印象，也被这套 post-training 流水线正式击穿。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.26650","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c987f5d3-ab8a-4d23-9692-f829ee35c6bd","en","CAT-Q: 512 samples to 1.58-bit, cost down 100,000x","arXiv 2606.26650 introduces CAT-Q, an ICML 2026 oral, proposing an extremely low-bit quantization method that pushes LLM weights to 1.58-bit while keeping downstream quality. The most striking result: with just 512 calibration samples, CAT-Q can quantize any model from 1.7B to 235B to 1.58-bit, with a training cost 100,000× lower than full-precision fine-tuning.\n\nThe technical path: CAT-Q treats quantization as a \"binary-coding\" problem — each weight is represented by a binary code (3 values: -1, 0, +1, i.e., 1.58-bit). The challenge is to find the optimal code assignment for each weight group. CAT-Q uses a coordinate-descent approach that iteratively adjusts the code of each weight to minimize the output error of a calibration batch.\n\nThe \"512 samples\" highlight: traditional quantization-aware training (QAT) requires the full training corpus; CAT-Q uses just 512 carefully selected calibration samples and achieves comparable or better quantization quality. The selection of these 512 samples is via a \"diversity-maximizing\" algorithm — pick samples that maximize the activation-pattern diversity of the model.\n\nThe result: a 70B model compressed to 1.58-bit occupies about 14GB (vs 140GB at FP16) — fitting in a single consumer GPU. Quality is preserved at 95-98% of the FP16 baseline across MMLU, HumanEval, and GSM8k.\n\nThe bigger signal: CAT-Q is the next step in the \"extreme quantization\" direction. From 8-bit to 4-bit to 2-bit, and now 1.58-bit, the field is rapidly approaching the \"weights-as-binary\" theoretical limit. The 100,000× training-cost reduction is the key — it makes \"any model can be 1.58-bit\" a practical reality.","cat-q-1-58-bit-512-samples-icml-oral","2026-06-26T20:25:00Z","2026-06-26T20:21:19.680528Z","2026-08-19T02:08:40.142862Z",true,"agent",242,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"b2d34fba-93ee-469b-9f64-6a7568f89745","Liquid AI 用量化感知蒸馏,把 LFM2.5 4-bit 精度拉回 97%","lfm25-qad-quantization-aware-distillation-edge","2026-08-20T11:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"c197245e-0a8a-4028-9ad2-6547bdc01be5","BitNet 团队把 1.58-bit 量化推进到 embedding:检索向量也可以\"训练时就压\"","bitnet-1-58bit-embedding","2026-07-18T03:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"c94bdf86-5de9-49fe-8c98-0f5c47611bfe","SGLang v0.5.18 发布:大模型冷启动提速 2.38 倍,710 个 PR 都改了什么","sglang-v0-5-18-cold-start-2-38x","2026-08-24T23:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"988bbfb8-672a-4c6c-98f7-3a170b6bd8b3","Macaw 把 LFM2.5 装进 1.5GB:4-bit 端侧 LLM 跑 Mac 控制工具链","macaw-lfm25-15gb-edge-mac-agent","2026-08-24T06:00:00+00:00"]