[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-prismml-bonsai-8b-1bit-1-15gb-edge":3,"topics-all":36,"news-related-a335e05e-2cb5-4e8a-8199-9d4ee1de0b01":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"a335e05e-2cb5-4e8a-8199-9d4ee1de0b01","PrismML Bonsai 8B：首个商用级1-bit量化LLM，8B参数压缩至1.15GB","来自加州理工团队的PrismML发布了Bonsai 8B，宣称这是首个具备商用可行性的1-bit大语言模型。其核心突破在于：一个80亿参数的语言模型，经过1-bit量化后，权重仅占1.15GB内存——相当于同类FP16模型的1\u002F12到1\u002F14。\n\n1-bit量化的原理很直接：每个权重仅存储一个比特（0或1），0映射为-scale，1映射为+scale。每128个权重共享一个FP16缩放因子，在极低精度下保留了模型的分布特征。这种极端压缩带来的最大优势是部署门槛的骤降——1.15GB意味着模型可以直接在手机、嵌入式设备甚至浏览器端运行，而不需要GPU或云端API。\n\n在基准测试方面，Bonsai 8B在8B参数级别的模型中保持了有竞争力的表现。虽然1-bit量化不可避免地带来精度损失，但PrismML团队通过改进的训练后量化算法和缩放因子优化，使得模型在推理、常识和编码任务上仍然接近全精度同级模型。这一结果打破了此前1-bit量化只能用于演示的刻板印象。\n\nPrismML于3月31日结束隐身模式正式亮相，Bonsai 8B已发布MLX格式的开源权重。对于边缘AI和端侧推理场景而言，这是一个值得关注的方向：当模型小到可以忽略部署成本时，AI应用的形态将发生根本性变化。","https:\u002F\u002Fwww.forbes.com\u002Fsites\u002Fjonmarkman\u002F2026\u002F04\u002F02\u002Fprismml-introduces-the-first-commercially-viable-1-bit-llm\u002F","3ce68fd9-8f57-444e-9501-e5ddd707d9bf",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"58fbb820-68fe-4e4c-8807-291741da20e3","en","PrismML Bonsai 8B: commercial-grade 1-bit LLM at 1.15GB","PrismML, a Caltech-affiliated team, has released Bonsai 8B, claiming it is the first 1-bit large language model with commercial viability. The core breakthrough: an 8B-parameter language model, after 1-bit quantization, has weights that occupy just 1.15GB of memory — roughly 1\u002F12 to 1\u002F14 of an equivalent FP16 model.\n\nThe 1-bit quantization principle is straightforward: each weight stores only one bit (0 or 1), with 0 mapped to -scale and 1 mapped to +scale. Every 128 weights share one FP16 scaling factor, preserving the model's distribution characteristics at extreme precision. The biggest benefit of this extreme compression is the dramatic drop in deployment cost — 1.15GB means the model can run directly on phones, embedded devices, or even in a browser, without a GPU or cloud API.\n\nIn benchmarks, Bonsai 8B maintains competitive performance among 8B-class models. While 1-bit quantization inevitably brings precision loss, PrismML's improved post-training quantization algorithm and scaling-factor optimization keep the model close to full-precision peers on reasoning, common-sense, and coding tasks. This shatters the old notion that 1-bit quantization is only good for demos.\n\nPrismML emerged from stealth on March 31; Bonsai 8B's open-source weights are already available in MLX format. For edge AI and on-device inference scenarios, this is a direction worth watching: when models are small enough that deployment cost is negligible, the form factor of AI applications will fundamentally change.","prismml-bonsai-8b-1bit-1-15gb-edge","2026-04-23T13:08:00Z","2026-04-23T13:11:41.533429Z","2026-08-19T02:08:40.142862Z",true,"agent",195,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"b571067a-9fa8-42bf-9431-98f26ac78e03","伯克利把LLM推理搬进SSD:KV缓存压缩15倍","llm-inference-in-flash-cim-ssd","2026-09-19T21:10:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"d056f67b-7e0d-4e44-8d39-e31ea50deeae","Bonsai 2 27B 三元压缩:Qwen3.8 压到 5.9 GB,benchmark 留存 98.2%","bonsai-2-27b-ternary-qwen3-8-compression","2026-09-17T15:47:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"ddb7bc6c-6b6e-4797-ab76-d1aeab5a3002","压缩得好≠部署得好:树莓派实测边缘 LLM,LoRA恢复模型100题押97个同答案","edge-llm-compression-raspberry-pi","2026-08-23T13:30:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"031e715e-c9d6-4855-83da-0515f33f0e3c","POCKET：35B MoE 1-bit 跑进 iPhone，27 tok\u002Fs","pocket-35b-moe-iphone-edge","2026-07-28T04:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"224fc205-6b1a-48d3-8562-df66921d017e","腾讯把序列长度做成 LLM 第四 scaling 轴：WeLM-HD4-617B 在不增参数前提下反超 Kimi K2.6","tencent-welm-hd4-617b-sequence-scaling","2026-07-11T04:10:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"aa53081a-e448-4087-aaa9-c822a7074bbc","LlamaWeb：WebGPU 跑 llama.cpp，16 设备吞吐 +45-69%","llamaweb-webgpu-llm","2026-06-29T18:00:00+00:00"]