[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-apple-prismml-1bit-qwen-iphone":3,"news-related-34b1a171-a0bb-44b3-9342-28d0185f0afc":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"34b1a171-a0bb-44b3-9342-28d0185f0afc","Apple 接触 PrismML：1-bit 压缩 27B Qwen 塞进 iPhone","2026 年 7 月 9 日，The Information 报道 Apple 近期与 PrismML 接触，讨论把后者基于 1-bit 量化技术压缩后的 Qwen 3.6 模型集成进 iPhone 17 Pro 的可能性。PrismML 把 27B 参数的开源 LLM 压到能装进手机内存并跑软件工程类任务，计划 7 月 14 日开源发布。\n\nPrismML 三个月前发布的 1-bit Bonsai 8B 已经在 1.15 GB 体积下逼近 16 GB 同类模型的能力——它在 MMLU、HumanEval+ 等基准上仅小幅落后 Qwen3 8B 全精度版本，iPhone 17 Pro Max 上 44 tok\u002Fs 的速度第一次让\"8B 模型跑在手机\"成为日常体验。如果 27B 的 Qwen 3.6 走同一路径，理论上模型尺寸约 3.8 GB，是 Apple AFM 3 Core（约 3B 激活参数）所装下参数量的近一个量级提升。\n\n对 Apple 来说，这一步背后是从\"自研 AFM 路线\"到\"引入第三方压缩模型\"的策略摆动。AFM 3 用 IFP + NAND-DRAM 把 20B 稀疏模型做到可装入手机，但推理质量仍受限于自身团队规模和后训练投入。PrismML 的\"intelligence density\"路线——同等体积下能力密度提升 10×——对 Apple Intelligence 的功能扩展是直接杠杆，特别是把 coding agent 这类长上下文工具塞进离线环境。\n\n对中国开源生态而言这是个反向输出信号：通义千问 Qwen 3.6 被美国 AI 实验室选作端侧基座，意味着阿里在大规模开源权重与后训练对齐上的积累正在被海外平台反向消费。如果 Apple 最终在 iOS 27 内以 Apple Intelligence 形式封装这一能力，Qwen 系列的国际能见度会再上一档。\n\n端侧 LLM 已经从\"3B 够用就行\"跨进\"27B 不一定需要云\"的阶段——iPhone 17 Pro 只是个起点，Android 阵营、车规芯片、机器人主控都会跟进这条 1-bit 路径。下一个变量不再是模型能不能装下，而是 1-bit 专用硬件何时商业化。","https:\u002F\u002F9to5mac.com\u002F2026\u002F07\u002F09\u002Freport-apple-interested-in-startup-that-runs-giant-ai-models-on-iphone-without-servers\u002F","b1c980f4-2409-4767-adb6-a7ccdd79f990",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"f29a71e6-0bf4-4bff-822b-68d0034f3c41","en","Apple courts PrismML: 1-bit 27B Qwen compressed for iPhone","On July 9, 2026, The Information reported that Apple has recently been in contact with PrismML, discussing the possibility of integrating a Qwen 3.6 model compressed by PrismML's 1-bit quantization technology into the iPhone 17 Pro. PrismML has compressed a 27B-parameter open-source LLM down to fit into phone memory and run software-engineering-type tasks, with plans to open-source release on July 14. PrismML's 1-bit Bonsai 8B released three months ago, at 1.15 GB, already approaches the capability of 16 GB same-class models — it only slightly trails the Qwen3 8B full-precision version on MMLU, HumanEval+ and other benchmarks, and the 44 tok\u002Fs on iPhone 17 Pro Max is the first time \"running an 8B model on a phone\" has become a daily experience. If the 27B Qwen 3.6 takes the same path, the model size is theoretically about 3.8 GB, nearly an order of magnitude more than what Apple's AFM 3 Core (about 3B active parameters) can fit. For Apple, behind this step is a strategic swing from the \"in-house AFM route\" to \"introducing third-party compressed models\". AFM 3 uses IFP + NAND-DRAM to get a 20B sparse model into a phone, but inference quality is still limited by its own team scale and post-training investment. PrismML's \"intelligence density\" route — 10× capability density improvement at the same volume — is a direct lever for extending Apple Intelligence's functionality, especially stuffing long-context tools like coding agents into offline environments. For the Chinese open-source ecosystem, this is a reverse export signal: Tongyi's Qwen 3.6 is selected by an American AI lab as an on-device base, meaning Alibaba's accumulation in large-scale open-source weights and post-training alignment is being reverse-consumed by overseas platforms. If Apple ultimately packages this capability in the form of Apple Intelligence in iOS 27, the Qwen series' international visibility will rise another tier. On-device LLM has moved from \"3B is enough\" to \"27B doesn't necessarily need the cloud\" stage — iPhone 17 Pro is just a starting point, the Android camp, automotive-grade chips, and robot main controllers will all follow this 1-bit path. The next variable is no longer whether the model can fit, but when 1-bit-specific hardware will commercialize.","apple-prismml-1bit-qwen-iphone","2026-07-10T02:00:00Z","2026-07-10T02:09:01.190641Z","2026-08-19T02:08:40.142862Z",true,"agent",217,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"ddb7bc6c-6b6e-4797-ab76-d1aeab5a3002","压缩得好≠部署得好:树莓派实测边缘 LLM,LoRA恢复模型100题押97个同答案","edge-llm-compression-raspberry-pi","2026-08-23T13:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"5a5b1531-e1b2-469b-8064-772223231183","KronQ：Kronecker Hessian 拆掉 GPTQ 的 2-bit 墙","kronq-kronecker-hessian-gptq","2026-07-13T16:02:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"53d67819-1d5c-4c1a-8dbb-1f49d9c76304","BiSCo-LLM 把 LLM 量化推过 2-bit 墙：Lookup-free 球面编码 + 类别恢复蒸馏，告别 VQ 码本","bisco-llm-2bit-quantization","2026-07-10T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"e428fd02-4e0c-4702-8173-9bbebb02cc31","Lynx:渐进式投机量化让长上下文 LLM 的 KV 缓存传输跑出 1.43× 加速","lynx-progressive-speculative-quantization","2026-07-07T10:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"1480f5c1-5eea-4513-bb38-ad5a4bb3cc25","Log_bQuant 改写 4-bit 量化:TUM 让 14B LLM 保住 72.97% MMLU","log-bquant-4bit-quantization","2026-07-06T20:11:00+00:00"]