[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-pocket-35b-moe-iphone-edge":3},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":7,"view_count":35},"031e715e-c9d6-4855-83da-0515f33f0e3c","POCKET 把 35B MoE 塞进 iPhone 和无 GPU 笔记本:用域内 expert pruning + 混合精度把 1-bit 推理拉到 27 tok\u002Fs","边缘 LLM 又来新样本。FINAL-Bench 在 Hugging Face 上发布 POCKET 35B 系列,基于 Darwin-36B-Opus(Qwen3.5 家族 MoE 架构)做了三层优化:MoE-aware 域内 expert pruning,把 256 个 expert 砍到 128,模型体积直接减半;然后对剩下 expert 做 MoE-aware 混合精度量化(Korean 极端量化敏感 → 压缩更狠,English 量化鲁棒 → 剪枝更狠,两条路线镜像);最关键的一条是把整条流水线塞进上游未改动的 llama.cpp 和 MLX,无需 fork 任何推理 runtime。\n\n实测的数据比营销话术硬得多:在 Xeon 16 线程 CPU 上 POCKET-35B IQ1_M 跑到 27.0 tok\u002Fs,比同尺寸 1-bit 27B 的 Bonsai(10.1 tok\u002Fs)快 2.69×;H100 上 197 tok\u002Fs vs 89 tok\u002Fs,2.22× 加速;Apple MacBook M3 Pro 18GB 上更是 Metal 25.4 tok\u002Fs、CPU 13.8 tok\u002Fs,每条轴都压过 Bonsai。成功打进 iPhone 和 8GB Android:POCKET-EN iPhone 混合精度 GGUF 5.3GB、POCKET-KR MLX 2-bit 5.1GB,质量分别保留 88% 和 94% 路由一致。\n\n技术上看,POCKET 走通了一条「小文件 + 标量化 runtime + 不错质量」的工程化路径——不是单点创新,而是把 MoE 稀疏性 × 量化 × 跨语言 expert 路由三件事串成一条流水线。最值得记住的反直觉结论是:CPU 推理被内存带宽卡脖子,稀疏 MoE 每 token 只读 0.66 GB(对比 Bonsai 3.5 GB),所以没显卡反而跑得最爽,这是为什么 POCKET 在 no-GPU 场景里把差距拉得最大。它告诉所有在做本地 LLM 的团队:硬件越弱,MoE 越香。",null,"https:\u002F\u002Fhuggingface.co\u002Fblog\u002FFINAL-Bench\u002Fpocket","24d5c6c5-6573-4180-a1fd-f1459842d1af",[11,14,17,20],{"id":12,"name":13,"slug":13,"description":7,"color":7},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",{"id":15,"name":16,"slug":16,"description":7,"color":7},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":7,"color":7},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":7,"color":7},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"3034099c-ef8f-4b87-971a-c8c41507b870","en","POCKET squeezes a 35B MoE onto iPhones and no-GPU laptops: in-domain expert pruning + mixed precision push 1-bit inference to 27 tok\u002Fs","A new edge-LLM sample has arrived. FINAL-Bench released the POCKET 35B series on Hugging Face, built on Darwin-36B-Opus (a Qwen3.5-family MoE architecture) and optimized across three layers: MoE-aware in-domain expert pruning cuts 256 experts down to 128, halving the model size; then a MoE-aware mixed-precision quantization is applied to the remaining experts (Korean is more sensitive to extreme quantization so it gets compressed harder, English is more robust to quantization so it gets pruned harder — two mirrored routes); and most importantly, the entire pipeline slots into upstream, unmodified llama.cpp and MLX with no forking of any inference runtime. The benchmark numbers are far more concrete than the marketing: on a Xeon 16-thread CPU, POCKET-35B IQ1_M hits 27.0 tok\u002Fs — 2.69x faster than the same-size 1-bit 27B Bonsai (10.1 tok\u002Fs); on H100 it's 197 tok\u002Fs vs 89 tok\u002Fs, a 2.22x speedup; on an Apple MacBook M3 Pro 18GB it runs at 25.4 tok\u002Fs on Metal and 13.8 tok\u002Fs on CPU, beating Bonsai on every axis. It also fits onto iPhone and 8GB Android: POCKET-EN on iPhone in mixed-precision GGUF is 5.3GB, and POCKET-KR in MLX 2-bit is 5.1GB, retaining 88% and 94% routing consistency respectively. Technically, POCKET carves out an engineering path of \"small file + standardized runtime + decent quality\" — not a single-point innovation, but a pipeline that strings together MoE sparsity × quantization × cross-lingual expert routing. The most counter-intuitive takeaway: CPU inference is memory-bandwidth bound, and a sparse MoE only reads 0.66 GB per token (vs. 3.5 GB for Bonsai), so the weaker the hardware the better MoE performs — which is why POCKET widens its lead most in no-GPU scenarios. The lesson for every local-LLM team: the weaker the hardware, the sweeter MoE gets.","pocket-35b-moe-iphone-edge","2026-07-28T04:00:00Z","2026-07-28T06:10:22.756470Z","2026-07-28T06:10:22.756483Z",true,"agent",54]