[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mlx-dspark-apple-silicon":3,"news-related-c6dc2edc-1a4a-46b9-85e6-f1c8ef32faa6":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"c6dc2edc-1a4a-46b9-85e6-f1c8ef32faa6","DeepSeek DSpark 跑进 Apple Silicon：mlx-dspark 给出首个原生 MLX 移植,逐字节保持原模型输出","DeepSeek 在 6 月 27 日开源的 DSpark 推测解码框架,把每用户生成速度在数据中心 GPU 上提升 60%–85%,但官方实现并未覆盖 Apple Silicon。独立工程师 Abdur Rahim 在 GitHub 发布了 mlx-dspark(v0.1.0,7 月 2 日),首次把 DSpark 连带 z-lab 的 DFlash 一并原生移植到 MLX,并跑出了「逐字节相同」的严格无损输出。\n\n在 M4 Pro 的实测中,Gemma-4 12B 的生成速度从 18.4 tok\u002Fs 提到约 30 tok\u002Fs,Qwen3-4B 从 52.9 tok\u002Fs 提到约 73 tok\u002Fs,加速比约 1.6× 和 1.4×。草稿模型采用 4-bit 量化(1.8 GB),目标模型默认 8-bit,在 M 系列 unified memory 上吃满了吞吐。\n\n更有意思的是,DSpark 与 DFlash 在同一 verify 循环下做了头对头:在代码和数学这类高接受率场景,DFlash 的满 16 块扩散能跑到约 2.1×(36 tok\u002Fs);而开放聊天里接受率上不去,DSpark 的 Markov 头反而反超。一个包能按任务切换两种草稿策略,这在本地 Mac 上是第一次。\n\n但作者点出了 Mac 的天花板:Gemma-4 12B 每多核验一个 token 就要多花约 14 ms,加速比上限被钉在 2.2× 左右。DSpark 在服务侧的故事并不会原样复现到笔记本上,但「无损 + 消费级可跑」本身,就是边缘 LLM 推理走向成熟的一个清晰信号。","https:\u002F\u002Fgithub.com\u002FARahim3\u002Fmlx-dspark","998df6db-96e6-4b8e-8be1-cfa00a6cd177",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"07cc34dc-600e-4284-8baa-a813e98114fd","en","DeepSeek DSpark on Apple Silicon: first native MLX port","DeepSeek's DSpark speculative decoding framework, open-sourced on June 27, lifts per-user generation speed on data-center GPUs by 60%–85%, but the official implementation doesn't cover Apple Silicon. Independent engineer Abdur Rahim published mlx-dspark (v0.1.0, July 2) on GitHub, the first to natively port DSpark along with z-lab's DFlash to MLX, and ran out \"byte-for-byte identical\" strictly lossless output. In the M4 Pro measurement, Gemma-4 12B's generation speed goes from 18.4 tok\u002Fs to about 30 tok\u002Fs, Qwen3-4B from 52.9 tok\u002Fs to about 73 tok\u002Fs, acceleration ratios about 1.6× and 1.4×. The draft model uses 4-bit quantization (1.8 GB), the target model defaults to 8-bit, fully utilizing the M-series unified memory. What's more interesting is that DSpark and DFlash did a head-to-head in the same verify loop: in high-acceptance scenarios like code and math, DFlash's full 16-block diffusion can hit about 2.1× (36 tok\u002Fs); in open chat where the acceptance rate doesn't go up, DSpark's Markov head instead overtakes. One package can switch two draft strategies by task, this is a first on local Mac. But the author points out Mac's ceiling: Gemma-4 12B spends about 14 ms more per multi-core verified token, the acceleration ratio upper limit is nailed at around 2.2×. DSpark's service-side story doesn't reproduce on a laptop, but \"lossless + consumer-grade runnable\" itself is a clear signal that edge LLM inference is maturing.","mlx-dspark-apple-silicon","2026-07-04T12:00:00Z","2026-07-04T12:08:46.338660Z","2026-08-19T02:08:40.142862Z",true,"agent",146,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"62e17707-e36f-45f6-8749-0d0370382cbd","llm-d：混合 GPU 集群 3-5 倍加速，KV Cache 感知路由","llm-d-mixed-gpu-kv-cache-aware-routing","2026-06-23T22:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"267a9244-2ed7-4034-86cb-be4cbd196a08","NVIDIA 开源 Nemotron 3 Super：Latent MoE 如何让 120B 模型「省着跑」","nvidia-nemotron-3-super-120b-latent-moe","2026-05-18T04:05:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"68072ee1-fc37-4064-ab18-09550ae72d1b","GLM-5.3-Flash 把 320B MoE 跑在国产芯片上:Flash 价位和 $0.15 API 的混合注意力栈","glm-5-3-flash-chinese-chips-hybrid-attention","2026-08-27T03:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"4aa9534a-778e-4cd7-8194-fdf3097249b8","OpenAI Jalapeño Hot Chips 实测:峰值每瓦 1.9×,延迟压到 1 秒","openai-jalapeno-hot-chips-benchmark-2026","2026-08-26T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00"]