[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-google-turboquant-kv-cache-compression":3,"topics-all":42,"news-related-645dd52f-ae6e-4931-9d74-4582feb4ecb1":61},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":29,"news_slug":35,"published_at":36,"created_at":37,"modified_at":38,"is_published":39,"publish_type":40,"image_url":13,"view_count":41},"645dd52f-ae6e-4931-9d74-4582feb4ecb1","Google TurboQuant：LLM推理内存压缩6倍的技术突破","Google在ICLR 2026发布的TurboQuant算法实现了革命性的LLM KV缓存压缩技术，将16位精度压缩至3位，内存使用减少6倍且精度零损失。该技术通过正交旋转和Lloyd-Max最优化量化，解决了长上下文推理中的内存瓶颈问题。在H100 GPU上，4位TurboQuant将注意力计算速度提升8倍，为推理成本带来显著优化。这项突破不仅改变了内存芯片市场预期，更让百亿参数模型在消费级硬件上运行长上下文成为可能，标志着AI推理效率的重要里程碑。","https:\u002F\u002Fresearch.google\u002Fblog\u002Fturboquant-redefining-ai-efficiency-with-extreme-compression\u002F","4d11edad-2df6-45f6-b71f-70f65de7f7fd",[10,14,17,20,23,26],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"8cf7490f-2449-4ba7-be19-61befa0d92b4","google",{"id":24,"name":25,"slug":25,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":27,"name":28,"slug":28,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[30],{"id":31,"lang":32,"title":33,"summary":34,"content":13},"48caef34-e105-4692-931a-9c0c7072d751","en","TurboQuant: Google compresses LLM inference memory 6x","Google's TurboQuant algorithm, released at ICLR 2026, delivers a revolutionary KV-cache compression technique for LLMs: it compresses 16-bit precision down to 3 bits, cutting memory usage by 6x with zero precision loss. The technique uses orthogonal rotation and Lloyd-Max optimal quantization to tackle the memory bottleneck in long-context inference. On H100 GPUs, 4-bit TurboQuant speeds up attention computation by 8x, delivering significant inference cost optimization. This breakthrough not only shifts memory-chip market expectations, but also makes running long-context billion-parameter models on consumer hardware feasible — marking an important milestone in AI inference efficiency.","google-turboquant-kv-cache-compression","2026-04-23T01:11:00Z","2026-04-23T01:15:58.301782Z","2026-08-19T02:08:40.142862Z",true,"agent",195,[43,52],{"slug":44,"tag_slug":44,"title_zh":45,"title_en":46,"intro_zh":47,"intro_en":48,"id":49,"is_active":39,"created_at":50,"modified_at":51},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":53,"tag_slug":53,"title_zh":54,"title_en":55,"intro_zh":56,"intro_en":57,"id":58,"is_active":39,"created_at":59,"modified_at":60},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":62},[63,68,73,78,83,88],{"id":64,"title":65,"news_slug":66,"published_at":67},"aa53081a-e448-4087-aaa9-c822a7074bbc","LlamaWeb：WebGPU 跑 llama.cpp，16 设备吞吐 +45-69%","llamaweb-webgpu-llm","2026-06-29T18:00:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"2266cea6-06f1-4932-8905-1bc3f2e5a8c0","Meta FAIR 字节蒸馏研究:End-Of-Token 渐近反超 token 蒸馏 4%,数据只需 1\u002F6","meta-fair-byte-distillation-token-ceiling-2026-09","2026-09-15T02:00:00+00:00",{"id":74,"title":75,"news_slug":76,"published_at":77},"49c90242-7793-46c7-961a-8a39e608e23d","竖线 token 才是量化命门：HyQuant 混合精度让 4-bit 注意力几乎无损","hyquant-hybrid-precision-attention-quantization","2026-09-13T17:06:39+00:00",{"id":79,"title":80,"news_slug":81,"published_at":82},"e91b3add-4f1d-48d3-ab7a-bd2c6e8c1765","QuIP 崩、OPTQ 降级:Kashin-DCT 在 4-bit 量化压力测试里活了下来","kashin-dct-2bit-llm-quantization","2026-09-12T15:10:00+00:00",{"id":84,"title":85,"news_slug":86,"published_at":87},"178aa5e5-2a4f-4a87-a97c-0da16295d96f","EMNLP 2026 OCGQuant:用通道配对治 NVFP4 陪葬误差,Qwen3-1.7B 接近 FP16","ocgquant-nvfp4-outlier-companion-grouping","2026-09-10T09:15:00+00:00",{"id":89,"title":90,"news_slug":91,"published_at":92},"6d60f9ba-4866-4520-9f12-d955e37f8472","Gated DeltaNet 全压 4-bit 没掉点:一篇论文拆掉混合 LLM 的量化禁忌","gated-deltanet-nvfp4-full-4bit","2026-09-04T15:08:02+00:00"]