[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-bonsai-2-27b-ternary-qwen3-8-compression":3,"topics-all":47,"news-related-d056f67b-7e0d-4e44-8d39-e31ea50deeae":66},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":33,"news_slug":40,"published_at":41,"created_at":42,"modified_at":43,"is_published":44,"publish_type":45,"image_url":14,"view_count":46},"d056f67b-7e0d-4e44-8d39-e31ea50deeae","Bonsai 2 27B 三元压缩:Qwen3.8 压到 5.9 GB,benchmark 留存 98.2%","PrismML 发布 Bonsai 2 27B 三元权重模型,把 Qwen3.8 27B 压到 5.9 GB,20 项基准平均分 83.9,留存原模型 98.2% 能力,RTX 5090 上跑到 143 tokens\u002Fs,Apache 2.0 开源。","过去三年,大模型算力账单几乎被\"参数越多 = 越聪明\"主导。PrismML 这家加州理工走出来的初创公司,9 月 17 日扔出一枚反方向炸弹——基于阿里 Qwen3.8 27B,三元压缩到 5.9 GB,20 项基准平均分 83.9,留存原模型 98.2% 能力,RTX 5090 上跑到 143 tokens\u002Fs,Apache 2.0 开源免费下载。这个数字组合意味着:一台普通 PC,甚至高端手机,现在能跑出接近云端 27B 模型的智能密度。\n\n## 三元权重动了什么\n\n正常每个权重 16 bit,Bonsai 2 只存 +1、-1、0 三个值。信息密度从 65536 个状态坍缩到 3 个,代价是表达精度,收益是 9-10 倍内存压缩。5.9 GB 正好踩在 PC 内存预算和高端手机闪存的可接受区间内。\n\n之所以叫\"三元\"不是\"二值\",因为 0 在矩阵运算里承担\"消音器\"角色——让一部分权重通道直接归零,稀疏性反而成了稳定性的来源。\n\n## 98.2% 不是营销稿\n\n20 项基准横跨 reasoning、math、coding、instruction following、vision、agentic tool use——这套组合是当下少见的\"真活\"测法。Bonsai 2 27B 平均分 83.9,Qwen3.8 27B 是 85.4,差 1.5 分即 98.2% 留存。相比第一代 Bonsai 的 95% 提升 3 个百分点,说明每代压缩技术以可见速率收敛。\n\n这种保留比例在低比特量化里并不常见。多数 4-bit 方案会有 2-4 个百分点掉点,Bonsai 2 已逼近\"零代价\"区间。\n\n## 反方向在哪\n\n行业惯例是参数规模往上走——GPT、Claude 过去一年每代动辄把总参数翻倍,代价是云端推理成本指数级膨胀。5.9 GB 这个体量把\"在不在端侧跑\"从云端绑死变成个人设备可承担。\n\nIon Stoica(Berkeley Sky Computing Lab 主任,Databricks 联合创始人)把这概括为\"本地智能,设备已经买了就是免费的\"。当模型能跑在用户口袋里,云端 API 护城河会从\"模型能力\"转向\"工具链和分发\"。\n\n## 5.9 GB 的工程含义\n\nRTX 5090 上 143 tokens\u002Fs 意味着实时性不再是瓶颈。模型能装进 24 GB 显存预算,开发者可以单机部署整个 27B 级别的 agentic 流程,不需要任何云端 API。对做本地知识库、私域文档问答、offline coding agent 的团队是直接利好。\n\nPrismML 计划把三元压缩套到\"数百 B\"参数的更大模型上——CEO Babak Hassibi 在 TechCrunch 采访中说,模型越大,压缩时\"留住的智能\"越多,因为有更大的权重空间可容忍精度损失。如果这条规律成立,100B+ 模型未来也能跑在单台工作站上。\n\n## 所以呢\n\nBonsai 2 27B 的真正信号不是\"又一个开源 27B\",而是\"压缩技术的进步速率本身\"。95% → 98.2% 只用了一代,继续推进,100% 留存不会太远。当 benchmark 留存压平时,\"端侧 LLM\"就从营销词变成工程现实。\n\n对开发者来说,这一波三元压缩的真正红利是降低推理基础设施门槛;对模型厂商而言,\"用更大的模型卖 API\"的商业模型会第一次遇到来自设备端的实质性竞争。开源社区接下来几个月的关键看点,会落在\"压缩 100B+ 模型能否保持同样留存\"——如果能,这条路的终局就是把云变成备选。","https:\u002F\u002Fwww.prnewswire.com\u002Fnews-releases\u002Fprismml-launches-bonsai-2-27b-its-most-capable-model-yet-302882228.html","585cbd37-4e65-4869-8066-9a525884b3da",[11,15,18,21,24,27,30],{"id":12,"name":13,"slug":13,"description":14,"color":14},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":25,"name":26,"slug":26,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":28,"name":29,"slug":29,"description":14,"color":14},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",{"id":31,"name":32,"slug":32,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[34],{"id":35,"lang":36,"title":37,"summary":38,"content":39},"75822e16-a2e5-4c84-b1ef-ca65e892cee0","en","Bonsai 2 27B ternary: Qwen3.8 to 5.9 GB, 98.2% retention","PrismML Bonsai 2 27B: ternary-compressed Qwen3.8 27B to 5.9 GB, 83.9\u002F85.4 aggregate, 98.2% retention, 143 tokens\u002Fs on RTX 5090, Apache 2.0.","For the past three years, the compute bill for large models has been dominated by \"more parameters = smarter.\" PrismML, a startup from Caltech, threw a counter-direction bomb on September 17: ternary-compressed Qwen3.8 27B to 5.9 GB, scoring 83.9 on a 20-benchmark aggregate while retaining 98.2% of the original model's capability, running at 143 tokens\u002Fs on RTX 5090, free under Apache 2.0. The number mix means an ordinary PC, even a high-end smartphone, can now run intelligence density close to cloud 27B models.\n\n## What ternary weights change\n\nNormal weights sit at 16 bits each. Bonsai 2 stores only +1, -1, and 0. Information density collapses from 65,536 states to 3, paying precision for a 9-10x memory compression. 5.9 GB lands squarely within both PC memory budgets and high-end phone flash storage.\n\nWhy ternary rather than binary — the value 0 acts as a \"silencer\" in matrix operations, forcing part of the weight channels to zero; that sparsity becomes a stability source.\n\n## 98.2% is not marketing\n\nThe 20 benchmarks span reasoning, math, coding, instruction following, vision, and agentic tool use — a rare \"real-work\" test suite, not just MMLU. Bonsai 2 27B averages 83.9, Qwen3.8 27B is 85.4 — a 1.5-point gap, 98.2% retention. Up 3 points from the first Bonsai's 95%, evidence each compression generation converges at visible rate.\n\nThis retention ratio is uncommon in low-bit quantization. Most 4-bit schemes drop 2-4 points; Bonsai 2 has reached the \"zero-cost\" zone.\n\n## Where the counter-direction lives\n\nIndustry convention pushes parameter scale up — GPT and Claude double total parameters nearly every generation, at the cost of cloud inference costs growing exponentially. Most on-device low-bit schemes serve \"people already decided to run large models\"; the 5.9 GB size turns \"whether to run on-device\" from a cloud-locked choice into something a personal device can absorb.\n\nIon Stoica (Berkeley Sky Computing Lab director, Databricks co-founder) summarized it as \"local intelligence, free because the device is already paid for.\" When a model fits in a user's pocket, the cloud API moat shifts from \"model capability\" to \"toolchain and distribution.\"\n\n## Engineering implications of 5.9 GB\n\nThe 143 tokens\u002Fs on RTX 5090 means real-time is no longer a bottleneck. With the model fitting within a 24 GB VRAM budget, developers can deploy an entire 27B-class agentic pipeline on a single machine, no cloud API required. Direct upside for teams building local knowledge bases, private document Q&A, and offline coding agents.\n\nPrismML's plan is to apply ternary compression to \"several-hundred-billion-parameter\" larger models — CEO Babak Hassibi told TechCrunch that the larger the model, the more \"retained intelligence\" survives compression, because the larger weight space can tolerate precision loss. If this law holds, 100B+ models will fit on a single workstation in the future.\n\n## So what\n\nThe real signal of Bonsai 2 27B is not \"yet another open-source 27B\" — it is the rate at which compression technology itself improves. 95% → 98.2% in one generation; continuing that trajectory, 100% retention is not far. When benchmark retention flattens, \"on-device LLM\" stops being a marketing term and becomes engineering reality.\n\nFor developers, the dividend of this ternary compression wave is lowering the inference infrastructure floor. For model vendors, the \"sell API calls from bigger models\" business model will face its first real competition from the device side. The open-source community's key question for the months ahead: can 100B+ models be compressed at the same retention rate? If yes, the endgame is making the cloud an option.","bonsai-2-27b-ternary-qwen3-8-compression","2026-09-17T15:47:00Z","2026-09-18T03:08:00.713116Z","2026-09-18T03:08:00.713131Z",true,"agent",144,[48,57],{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":44,"created_at":55,"modified_at":56},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":58,"tag_slug":58,"title_zh":59,"title_en":60,"intro_zh":61,"intro_en":62,"id":63,"is_active":44,"created_at":64,"modified_at":65},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":67},[68,73,78,83,88,93],{"id":69,"title":70,"news_slug":71,"published_at":72},"d941056b-c2e7-42e5-965a-a982c20b1169","Qwen3.8-Flash-Next 架构细节:Gated Residual 多分支残差 + QSA micro-block 稀疏注意力","qwen3-8-flash-next-cost-efficiency-architecture","2026-09-02T02:00:00+00:00",{"id":74,"title":75,"news_slug":76,"published_at":77},"637f84e0-e6dc-490a-bba1-879f6527bdd5","Qwen3.8-Max 2.4T 开源:Gated DeltaNet 把长上下文成本砍到 1\u002F8","qwen3-8-max-2-4t-open-weights-gated-deltanet","2026-08-30T03:00:00+00:00",{"id":79,"title":80,"news_slug":81,"published_at":82},"40095b51-97b0-4fd4-9b1d-f636c970572e","阿里 Qwen 团队发布 Qwen3.8-Max:2.4 万亿参数 MoE 模型首度开放权重","qwen3-8-max-2-4t-moe-open-weights","2026-08-07T02:00:00+00:00",{"id":84,"title":85,"news_slug":86,"published_at":87},"df01dae0-0940-4947-a019-31c57066132c","阿里千问 Qwen3.8 预览版上线:2.4T 参数,把开源旗舰抬到 Fable 5 同一档","qwen-3-8-preview-2-4t","2026-07-19T10:01:00+00:00",{"id":89,"title":90,"news_slug":91,"published_at":92},"34b1a171-a0bb-44b3-9342-28d0185f0afc","Apple 接触 PrismML：1-bit 压缩 27B Qwen 塞进 iPhone","apple-prismml-1bit-qwen-iphone","2026-07-10T02:00:00+00:00",{"id":94,"title":95,"news_slug":96,"published_at":97},"1d80585e-c797-4aa1-ac68-ef87334d5d0c","PLaMo 3.0 Prime 正式发布：PFN 把「日语实战」做成日本国产 LLM 的差异化战场","plamo-3-0-prime-pfn-japanese-domestic","2026-06-24T08:15:00+00:00"]