[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-orcasaq2-3-bit-qwen-quantization":3,"topics-all":38,"news-related-c23c2fca-29cc-40e8-9ff2-aebb32aba767":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"c23c2fca-29cc-40e8-9ff2-aebb32aba767","3.21 比特量化:27B 模型从 54GB 压到 12.3GB","OrcaRouter 开源 OrcaSAQ2 27B，用 SAQ 量化把 Qwen3.8-27B 从 54 GB 压到 12.3 GB、平均 3.21 比特，困惑度仅涨 0.02%，262K 上下文保留，单张 16 GB 显卡即可跑 27B agent 模型。量化方法未披露，数字为厂商自报。","把 27B 推理模型塞进单张消费级显卡，过去靠蒸馏出小模型牺牲能力，OrcaRouter 给出另一条路：直接把权重量化到平均 3.21 比特。9 月 28 日，这家路由服务商在 Hugging Face 开源了其首个开源权重模型 OrcaSAQ2 27B（Apache-2.0），底座是阿里的 Qwen3.8-27B。\n\n## 压缩数字一览\n\n原始 BF16 检查点 54 GB，量化后只剩 12.3 GB，体积缩小 77.2%，约 4.4 倍压缩。保真度方面，WikiText-2 困惑度从 5.6468 涨到 5.6482，仅 +0.02%；Top-1 token 一致率 93.2%；平均 KL 散度 0.031。这些数字全部来自官方模型卡的自报测量：同一组权重与 BF16 参考走同一条评测路径得出。\n\n架构上保留了底座的混合注意力：64 层里 48 层 Gated DeltaNet 加 16 层全注意力，隐藏层 5120，词表 248,320；262K 上下文、thinking 模式、工具调用和 MTP 投机解码头完整保留，代价是视觉塔被砍、只支持纯文本。\n\n## 落到实际部署\n\n12.3 GB 的检查点意味着 16 GB 显卡装完权重还剩约 3.7 GB 给 KV cache，官方给的实用起点是约 32K 交互上下文，架构上限 262K。推理走 vLLM，在 15.7 GiB 显存上限下实测：单流 MTP 关闭 65.3 tok\u002Fs，MTP 开启 90.1 tok\u002Fs，吞吐提升 38%；8 流并发 332 tok\u002Fs，16 流 333 tok\u002Fs。MTP 会吃额外算力和 KV 容量，高并发场景需实测两种配置再选。\n\nAgentic benchmark 同样是厂商自报：SWE-bench Verified 70.0%，Terminal-Bench 2.1 58.4%。模型卡标注这些是公开参照点而非严格同 harness 对比——70.0 分来自一个 12.06 GB 的 27B 检查点，已经压过了 Qwen3-Coder-480B-A35B 的 69.6。\n\n## 为什么盯住长程 agent\n\n模型卡里最有信息量的不是压缩比，而是评测哲学：困惑度问的是下一个 token 分布像不像，长程评测问的是多次决策后还能不能把任务做完。量化误差在单轮问答里看不出来，但在 plan → act → observe → recover 循环里会复利式放大：一次工具调用选错，后续每一步都建立在被污染的状态上。\n\n这正是低比特量化在 agent 时代的核心风险：短 benchmark 掩盖退化，长轨迹暴露退化。OrcaSAQ2 把 BF16 保真度指标和下游任务指标并排公布，至少在方法论上是对的。\n\n## 两点冷水\n\n第一，量化方法本身是专有的。所谓 SAQ（敏感度感知量化）只是名字，校准策略、精度分配、打包技术全部未披露，想复现或基于它做研究的人拿到的只是一个压缩结果。第二，所有数字都是厂商自报：模型卡自己写明「公开分数用了不同的 agent 栈，不应被解读为严格的模型排名」，llm-releases.com 收录时也标注 vendor-reported。Apache-2.0 开源的是权重，不是方法。\n\n## 所以呢\n\n如果量化确实接近无损，agent 的部署经济学就要重算：一张 16 GB 卡跑 27B 长程 agent，和租用云端旗舰 API 之间的成本差距会被大幅拉开。但记住这句话——重点不是 3 比特，重点是 3 比特之下还有什么活了下来。在你自己的长轨迹任务上实测，比任何 benchmark 都可信。\n\n原文与全部数字见官方模型卡：https:\u002F\u002Fhuggingface.co\u002Forcarouter\u002FOrcaSAQ-2-27B","https:\u002F\u002Fhuggingface.co\u002Forcarouter\u002FOrcaSAQ-2-27B","f88344b2-8136-4f83-9c49-310b753c2bd3",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"5829b181-27db-46f5-ad94-9e2e2d1c4ef9","en","OrcaSAQ2 27B: A 27B Qwen Model in 12.3 GB at 3.21 Bits","OrcaRouter's OrcaSAQ2 27B shrinks Qwen3.8-27B from 54 GB to 12.3 GB at 3.21 bpw (+0.02% PPL), 262K context — a 27B agent on one 16 GB GPU. Vendor-reported.","Fitting a 27B reasoning model onto a single consumer GPU used to mean distilling a smaller model and losing capability. OrcaRouter offers another path: quantize the weights to an average of 3.21 bits. On September 28, the routing vendor released OrcaSAQ2 27B on Hugging Face under Apache-2.0 — its first open-weights model, built on Alibaba's Qwen3.8-27B.\n\n## The compression numbers\n\nThe original BF16 checkpoint is 54 GB; the quantized build is 12.3 GB — a 77.2% footprint reduction, roughly 4.4x smaller. On fidelity, WikiText-2 perplexity moves from 5.6468 to 5.6482, just +0.02%; Top-1 token agreement is 93.2%; mean KL divergence is 0.031. All figures come from the official model card's self-reported measurements, produced by running these exact OrcaSAQ2 weights against the BF16 reference through the same evaluation path.\n\nArchitecturally, OrcaSAQ2 keeps the base model's hybrid attention design: 48 Gated DeltaNet layers plus 16 full-attention layers out of 64, hidden size 5120, vocabulary 248,320, with 262K context, thinking mode, tool calling, and the MTP speculative-decoding head all intact. The cost: the vision tower is dropped, making this a text-only model.\n\n## What it means for deployment\n\nA 12.3 GB checkpoint means a 16 GB GPU holds the weights with about 3.7 GB left for KV cache; the official practical starting point is around 32K interactive context, against an architectural ceiling of 262K. Serving runs on vLLM. Measured under a 15.7 GiB memory cap: 65.3 tok\u002Fs single-stream with MTP off, 90.1 tok\u002Fs with MTP on — a 38% throughput gain — plus 332 tok\u002Fs at 8 concurrent streams and 333 tok\u002Fs at 16. The card also notes MTP consumes extra compute and KV capacity, so batched workloads should benchmark both configurations before choosing.\n\nAgentic benchmarks are likewise vendor-reported: SWE-bench Verified 70.0% and Terminal-Bench 2.1 58.4%. The model card explicitly frames these as public reference points rather than strict apples-to-apples comparisons — the 70.0 comes from a 12.06 GB 27B checkpoint, edging past Qwen3-Coder-480B-A35B's 69.6.\n\n## Why long-horizon agents are the point\n\nThe most informative part of the model card is not the compression ratio but its evaluation philosophy: perplexity asks how similar the next-token distribution is, while long-horizon evaluation asks whether the model can still finish the job after many decisions. Quantization error invisible in single-turn QA compounds across the plan → act → observe → recover loop — one wrong tool call changes the environment state, and every later step builds on polluted state.\n\nThat is the core risk of low-bit quantization in the agent era: short benchmarks hide degradation; long trajectories expose it. By publishing BF16 fidelity metrics alongside downstream numbers, OrcaSAQ2 at least gets the methodology right.\n\n## Two caveats\n\nFirst, the quantization method itself is proprietary. SAQ (sensitivity-aware quantization) is a name; calibration strategy, precision allocation, and packing techniques are all undisclosed, so anyone wanting to reproduce or build on it receives a compressed artifact, not a method. Second, every fidelity and benchmark number is vendor-reported — the card itself says public scores use different agent stacks and should not be read as a strict model-only ranking, and llm-releases.com flags the entry as vendor-reported too. Apache-2.0 opens the weights, not the method.\n\n## So what\n\nIf quantization is truly near-lossless, the deployment economics of agents get rewritten: one 16 GB GPU running a 27B long-horizon agent versus renting a cloud frontier API becomes a very different cost equation. But remember the line — the point is not 3 bits; the point is what survives at 3 bits. Benchmark on your own long trajectories; that beats any leaderboard.\n\nFull details and every number: the official model card at https:\u002F\u002Fhuggingface.co\u002Forcarouter\u002FOrcaSAQ-2-27B","orcasaq2-3-bit-qwen-quantization","2026-09-30T21:20:00Z","2026-09-30T21:10:35.996692Z","2026-09-30T21:10:35.996701Z",true,"agent",271,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"4ccc491b-dbee-4c84-beb6-7cf519f76320","LoRA 基座换 GGUF:40G 显存训 125B","lora-over-gguf-low-vram-training","2026-10-10T21:08:22+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"32aa44d5-81f1-4717-8d17-ff713101b725","Phonon-2:2.1比特量化ASR,164MB逼近全精度","phonon-2-2-bit-parakeet-asr","2026-10-03T13:10:31+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"d056f67b-7e0d-4e44-8d39-e31ea50deeae","Bonsai 2 27B 三元压缩:Qwen3.8 压到 5.9 GB,benchmark 留存 98.2%","bonsai-2-27b-ternary-qwen3-8-compression","2026-09-17T15:47:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"ce40a7fa-acca-4609-a82e-5800f2e1026a","35B 压成 9.96GB 单文件:BTL-4 的 2.30 bit 极限量化和两条反直觉结论","btl-4-compact-9gb-2bit-quantization","2026-08-17T13:20:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"031e715e-c9d6-4855-83da-0515f33f0e3c","POCKET：35B MoE 1-bit 跑进 iPhone，27 tok\u002Fs","pocket-35b-moe-iphone-edge","2026-07-28T04:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"33f9c56d-1bcd-403d-85e2-273fb3d498fb","循环状态量化破局:STEPQuant 6-bit 贴平 FP32,显存省 68.7%","stepquant-delta-rule-state-quantization","2026-10-08T21:07:39+00:00"]