[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-compactifai-llama-3-3-70b-intel-xeon":3,"news-related-c32d3160-4e07-4128-890f-4e135aac2cce":39},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":26,"news_slug":32,"published_at":33,"created_at":34,"modified_at":35,"is_published":36,"publish_type":37,"image_url":13,"view_count":38},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","Multiverse Computing 7 月 23 日宣布,基于其量子软件背景衍生的 CompactifAI 压缩技术,把 Meta 的 Llama 3.3 70B 模型在 Intel Xeon 6 Performance-core 服务器上跑出 3.86 tokens\u002Fs 的输出吞吐,相比未压缩基线提升 93.6%,256 并发场景吞吐提升 107%,延迟下降 51.7%,磁盘占用从 130 GiB 降到 65 GiB,精度保留 97% 以上。这套方案同时支持 Llama 4 Scout、DeepSeek R1、Mistral Small 3.1 等开源旗舰,搭配 vLLM CPU 与 AMX 矩阵扩展,为「不上 GPU 也能跑大模型」的私有化场景补上一块硬通货。技术看点不在于「再压一次 INT4」,而在于压缩后的「healing」再训练——把量化的精度损失通过一轮针对性继续训练消化掉,WinoGrande 等基准反而比原模型高 6.86%,这种「越压越准」的反直觉结果,正是过去一年大模型推理优化从单纯降精度转向「压缩 + 修复」联合优化的代表性样本。对国内做端侧、CPU 推理和私有化部署的团队来说,这条路线比纯 GPU 推理更具工程现实意义:同一台 Xeon 服务器,容量翻倍、并发能力提升,既不用赌新卡供应,也能保住绝大多数问答质量。","https:\u002F\u002Fwww.hpcwire.com\u002Faiwire\u002F2026\u002F07\u002F23\u002Fmultiverse-computing-says-compactifai-nearly-doubles-llama-3-3-performance-on-intel-xeon-6\u002F","5cc8dc91-bc4f-47a5-8f68-c12db79f3b01",[10,14,17,20,23],{"id":11,"name":12,"slug":12,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":24,"name":25,"slug":25,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[27],{"id":28,"lang":29,"title":30,"summary":31,"content":31},"35d0e9d2-d2c1-4164-b83f-5e0d50474625","en","CompactifAI halves Llama 3.3 70B, 1.9x throughput on Xeon 6","On July 23, Multiverse Computing announced that its CompactifAI compression technology — derived from its quantum-software background — runs Meta's Llama 3.3 70B at 3.86 tokens\u002Fs output throughput on Intel Xeon 6 Performance-core servers, a 93.6% improvement over the uncompressed baseline; at 256 concurrent sessions, throughput improves 107%, latency drops 51.7%, disk footprint shrinks from 130 GiB to 65 GiB, and accuracy retention is over 97%. The solution also supports Llama 4 Scout, DeepSeek R1, Mistral Small 3.1 and other open-source flagships, paired with vLLM CPU and AMX matrix extensions, adding a hard-currency option to the \"run big models without a GPU\" private-deployment scenario. The technical highlight isn't \"another INT4 compression\" — it's the \"healing\" retraining after compression, which digests the quantization's accuracy loss through a targeted continuation training pass, and on benchmarks like WinoGrande the compressed model actually scores 6.86% higher than the original. This counter-intuitive \"tighter and more accurate\" result is a representative sample of the past year's shift in LLM inference optimization: from purely dropping precision to \"compression + repair\" joint optimization. For Chinese teams working on edge, CPU inference, and private deployment, this route has more engineering real-world significance than pure-GPU inference: the same Xeon server doubles capacity and concurrency, without betting on new card supply, while keeping the vast majority of Q&A quality.","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00Z","2026-07-25T16:03:54.980484Z","2026-08-19T02:08:40.142862Z",true,"agent",117,{"items":40},[41,46,51,56,61,66],{"id":42,"title":43,"news_slug":44,"published_at":45},"4147f71b-eaa7-4d39-91cf-c2c572105e7f","FlashPrefill V2:128K 长文本 prefill 提速 47 倍,块稀疏注意力走进生产框架","flashprefill-v2-block-sparse-prefill","2026-08-21T19:10:00+00:00",{"id":47,"title":48,"news_slug":49,"published_at":50},"d2c56430-f9a7-4844-bfab-a8651279a70c","ResKV 不再把 KV 缓存压缩等同于删词：给被淘汰的信息留一份残差账本","reskv-residual-kv-cache-compression","2026-08-03T10:43:23+00:00",{"id":52,"title":53,"news_slug":54,"published_at":55},"cf7f8e7e-4efc-451e-9595-706b0be911ba","PolyQ:把\"3-bit LLM 跑在 CPU\"做成一件可预测的事","polyq-3bit-llm-cpu","2026-07-17T10:00:00+00:00",{"id":57,"title":58,"news_slug":59,"published_at":60},"5a5b1531-e1b2-469b-8064-772223231183","KronQ：Kronecker Hessian 拆掉 GPTQ 的 2-bit 墙","kronq-kronecker-hessian-gptq","2026-07-13T16:02:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"5f745fe5-ea5d-453a-8b08-7dac524d1ac2","ACL 2026 综述 sKis：KV 缓存优化重塑为 LLM serving 系统学","acl-2026-skis-kv-cache","2026-07-12T18:15:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"007a83c4-faed-459a-ab44-17b915e05fb5","TriRoute 把 MoE + MoD + KV 量化做成一个控制器：三条条件计算路径第一次协同","triroute-moe-mod-kv-quantization","2026-07-09T18:02:00+00:00"]