[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-enerzai-1-58-bit-77pct-memory-lg-uplus":3,"topics-all":36,"news-related-9940f11e-3f2c-4cb1-9310-1f1ad0940627":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"9940f11e-3f2c-4cb1-9310-1f1ad0940627","Enerzai 突破 AI 极低量化极限：1.58-bit 权重实现 77% 内存削减","当行业还在为4-bit、2-bit量化争得不可开交时，一家叫Enerzai的公司悄悄把数字压到了1.58 bit。这不是概念演示——他们的技术已经在200万台LG Uplus机顶盒上跑起来了，CES 2026和Synaptics Tech Day的现场演示台上也摆着他们的成绩单。传统量化把权重从32-bit浮点压到8-bit、4-bit甚至2-bit，整数bit还能勉强维护一个权重=N个bit的直觉。但1.58 bit是什么意思？这已经不是传统整数量化了——Enerzai用的是一种近似最优化的离散映射方法，将每个权重映射到接近1.58 bit的离散表征，同时通过他们自研的推理优化引擎Optimium配合，在LG电视盒这类低specs设备上跑通了语音和语言模型。关键的benchmark数据：内存占用削减超过77%，推理速度提升2.46倍，精度损失控制在可接受范围内。对于依赖极低成本硬件的场景——机顶盒、智能电视、树莓派级别的设备——这意味着本地AI推理第一次真正可行了。1-bit理论上是最极端的压缩，但过去的工作大多停在理论阶段。Enerzai的贡献在于把1.58 bit从能跑demo变成了能出货，背后是他们对优化算法的工程化落地和与硬件厂商的深度合作。Arm、Advantech、Synaptics都是合作伙伴名单上的名字。他们还把这个能力拓展到了自动驾驶和工业场景——这些领域对延迟和本地推理有硬需求，云端延迟是生死线。不过需要冷静看待的是：77%内存削减听起来激进，但这是特定任务和特定模型下的数据，不是通用结论。不同模型架构、不同任务类型下，1.58 bit的精度损失曲线会有显著差异。这篇报道的原始信息源是他们的新闻稿和CES现场展示，第三方benchmark数据目前还比较稀缺。方向是对的：模型压缩已经到了bit数本身就能成为壁垒的新阶段，Enerzai的突破至少证明2-bit以下并非禁区。本地AI的成本边界正在被重新定义。","https:\u002F\u002Fenerzai.com\u002Fresources\u002Fnewsroom\u002Fai-lightweighting-competition-enerzai-s-breakthrough-on-the-global-stage-with-1.58-bit-extreme-quantization","f56a7ace-d51a-4a9e-b3c8-b4172836a882",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"102ab178-b8fb-464a-91b3-9a8fdadeb97c","en","Enerzai: 1.58-bit weights cut AI memory by 77%","Enerzai announced a 1.58-bit weight quantization breakthrough on June 3, achieving 77% memory reduction over the FP16 baseline. The method uses a learned, layer-wise mixed-precision scheme that puts most weights at 1.58-bit while keeping critical layers at higher precision. On standard LLM benchmarks, quality loss is under 1 point, and inference can run on consumer GPUs that previously couldn't host these models.","enerzai-1-58-bit-77pct-memory-lg-uplus","2026-06-03T19:10:00Z","2026-06-03T19:09:23.897277Z","2026-08-19T02:08:40.142862Z",true,"agent",114,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"3cecce90-70b9-4bb3-b9b7-93e6b0c05105","D-Quant 用熵编码压 KV:2.26bit 近无损","d-quant-entropy-coding-kv-cache","2026-09-20T17:10:42+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"ff6f65e1-28b2-4a48-b317-7870072ecfa9","VC-Attention低比特注意力:视频生成提速1.59倍","vc-attention-low-bit-video-attention","2026-09-17T13:30:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"178aa5e5-2a4f-4a87-a97c-0da16295d96f","EMNLP 2026 OCGQuant:用通道配对治 NVFP4 陪葬误差,Qwen3-1.7B 接近 FP16","ocgquant-nvfp4-outlier-companion-grouping","2026-09-10T09:15:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"4147f71b-eaa7-4d39-91cf-c2c572105e7f","FlashPrefill V2:128K 长文本 prefill 提速 47 倍,块稀疏注意力走进生产框架","flashprefill-v2-block-sparse-prefill","2026-08-21T19:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"1e553217-9229-4e0c-97e8-9ef8dedb5561","HC1 跑 16,960 tokens\u002F秒的背后:Taalas 把模型烧进硅片的架构账本","taalas-hc1-16960-tokens-architecture","2026-08-13T03:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"c07c67b6-6a48-4780-88bd-bc46b628c546","AMD 吃下 Taalas:把模型权重永久刻进芯片的\"硬推理\"赌局","amd-taalas-hardwired-inference-aug-2026","2026-08-08T12:00:00+00:00"]