[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ling-3-0-tiny-edge-deployment-ledger":3,"news-related-bb42bbdc-55a5-422c-ae68-4d103639bfd2":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"bb42bbdc-55a5-422c-ae68-4d103639bfd2","Ling-3.0-tiny：M4 Pro 实测 86-90 tokens\u002Fs","inclusionAI 开源 Ling 3.0 系列入门款 Ling-3.0-tiny:7.9B 总参数、每 token 仅激活 1.3B 的混合推理 MoE,MIT 许可。3:1 KDA+MLA 堆叠加 128 选 9 稀疏专家,FP8 下 M4 Pro MacBook 实测 86-90 tokens\u002Fs、峰值内存约 8.34 GiB,把混合推理下放到消费级硬件;SGLang 已有官方镜像,vLLM 走分支,Ollama 尚在 PR 阶段。","## 一份「不用数据中心」的模型卡\n\ninclusionAI 开源了 Ling 3.0 系列的入门款 **Ling-3.0-tiny**:总参数 7.9B、每 token 仅激活 1.3B 的混合推理 MoE 模型,MIT 许可证,权重已上 Hugging Face。与同系列此前定位生产级的 Flash 档不同,tiny 的目标很明确——把混合推理能力塞进本地和资源受限的部署环境。\n\n## 架构:3:1 的 KDA+MLA,配上 128 选 9 的稀疏 MoE\n\nLing-3.0-tiny 继承了 Ling-3.0 系列的混合线性注意力架构,核心是两件事的叠加:\n\n- **3:1 KDA-MLA 堆叠**:每 4 层块由 3 层 Kimi Delta Attention(KDA)加 1 层 Multi-Head Latent Attention(MLA)组成,官方称这样在长上下文处理效率、参数效率和计算成本之间取得平衡;\n- **稀疏 MoE FFN**:128 个路由专家里,每个 token 只激活 8 个路由专家加 1 个共享专家——这正是 7.9B 总参数只动用 1.3B 的原因。\n\n模型原生支持混合推理:thinking 模式默认开启,也可以通过 enable_thinking 按请求关闭,同一个模型里既给快速响应也给多步推理。官方同时放出了 BF16、FP8、INT4 三档权重,覆盖从服务器到消费级硬件的部署组合。\n\n## 端侧账本:三组数字\n\n这份模型卡最有意思的地方,是它直接给出了消费级硬件的实测数字:\n\n| 硬件 | 精度 | 速度 |\n|---|---|---|\n| NVIDIA DGX Spark | FP8 | 约 100-105 tokens\u002Fs |\n| M4 Pro MacBook | FP8 | 约 86-90 tokens\u002Fs |\n\n8K 上下文长度下,峰值内存约 **8.34 GiB**。在 Artificial Analysis 的测试里,它输出速度超过 160 tokens\u002Fs,500 token 响应(含推理时间)端到端延迟约 18 秒;AA Intelligence Index v4.1.1 得分 25,AA Agentic Index 得分 16。\n\n## 工具链现状:一个模型,三种成熟度\n\n部署路径的差异值得单独说——这可能是端侧模型生态最真实的切片:\n\n- **SGLang**:最顺,官方预构建镜像 lmsysorg\u002Fsglang:dev-Ling-3.0-tiny 拉下来就跑,低延迟配方内置 MTP\u002FNEXTN 投机解码,YaRN 扩展到 256K 上下文,单张 141GB 级 GPU(H20-3e)即可起服务;\n- **vLLM**:要装社区分支(ling_3_0 分支自行编译),还没进主线;\n- **Ollama**:支持还在 PR 阶段(#17643),且仅限 Apple Silicon 的 MLX 路径,需自行编译,尚未进入官方发行版——尽管官方已在 48GB 统一内存的 M4 Pro 上验证过。\n\n## 评论:数字之外的信号\n\n单看 25 分的 AA 智能指数,tiny 谈不上惊艳——它的卖点从来不是榜单名次,而是「1.3B 激活参数 + 8.34 GiB 内存」这个组合让混合推理在 MacBook 上有了可用的速度。社区反应是实在的:发布一周多,HF 上已有 13 个社区量化版本。\n\n更值得留意的是工具链梯度:同一个模型,SGLang 有官方镜像、vLLM 等分支合入、Ollama 还在 PR——端侧开源模型的「发布」早已不是丢一包权重那么简单,推理框架的适配深度决定了它真正能被多少人用起来。MIT 许可证则把商业化顾虑也一并打消了。\n\n对想在本地跑 agent 工作流的开发者,这份模型卡给出了一份可以直接对着采购预算算的账本;对行业,它再次印证一个趋势:混合线性注意力 + 稀疏 MoE 的组合,正在从旗舰专属下放到 8B 级。\n\n> 素材来源:inclusionAI Ling-3.0-tiny 模型卡(Hugging Face):https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLing-3.0-tiny","https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLing-3.0-tiny","e7a2ac45-ed74-4858-8663-3bd2943959a2",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"d11f0044-8aef-487c-bebe-89ce4683a4a3","moe",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"8b6467ca-28fd-45e8-a9ac-6b70b945769f","en","Ling-3.0-tiny benchmarks: 86-90 tokens\u002Fs on M4 Pro","inclusionAI has open-sourced Ling-3.0-tiny, the entry tier of the Ling 3.0 family: a 7.9B-parameter hybrid-reasoning MoE with only 1.3B activated per token, MIT-licensed. With 3:1 KDA+MLA stacking and a 128-pick-9 sparse expert setup, it delivers measured 86-90 tokens\u002Fs on an M4 Pro MacBook at FP8 with ~8.34 GiB peak memory, bringing hybrid reasoning to consumer hardware; SGLang has an official image, vLLM needs a branch build, and Ollama support is still a PR.","## A Model Card That Says \"No Datacenter Required\"\n\ninclusionAI has open-sourced **Ling-3.0-tiny**, the entry-level member of the Ling 3.0 family: a hybrid-reasoning MoE model with 7.9B total parameters and only 1.3B activated per token, released under the MIT license with weights on Hugging Face. Unlike the production-oriented Flash tier of the same family, tiny has one clear goal — bringing hybrid reasoning to local and resource-constrained deployments.\n\n## Architecture: 3:1 KDA+MLA Stacking, Plus a 128-Pick-9 Sparse MoE\n\nLing-3.0-tiny inherits the hybrid linear-attention architecture of the Ling-3.0 series, built on two stacked ideas:\n\n- **3:1 KDA-MLA stacking**: every 4-layer block combines 3 layers of Kimi Delta Attention (KDA) with 1 layer of Multi-Head Latent Attention (MLA). The team says this balances long-context processing efficiency, parameter efficiency, and computational cost;\n- **Sparse MoE FFN**: out of 128 routed experts, each token activates only 8 routed experts plus 1 shared expert — which is exactly why a 7.9B-parameter model only moves 1.3B per token.\n\nThe model natively supports hybrid reasoning: thinking mode is on by default and can be turned off per request via enable_thinking, so a single model serves both fast responses and multi-step reasoning. BF16, FP8, and INT4 weights are all provided, covering deployment options from servers to consumer hardware.\n\n## The Edge Ledger: Three Sets of Numbers\n\nThe most interesting part of this model card is that it directly publishes measured numbers on consumer hardware:\n\n| Hardware | Precision | Speed |\n|---|---|---|\n| NVIDIA DGX Spark | FP8 | ~100-105 tokens\u002Fs |\n| M4 Pro MacBook | FP8 | ~86-90 tokens\u002Fs |\n\nAt an 8K context length, peak memory sits around **8.34 GiB**. In Artificial Analysis testing, output speed exceeds 160 tokens\u002Fs, with roughly 18 seconds of end-to-end latency for a 500-token response including reasoning time; it scores 25 on the AA Intelligence Index v4.1.1 and 16 on the AA Agentic Index.\n\n## Toolchain Reality: One Model, Three Maturity Levels\n\nThe deployment paths deserve their own section — this may be the most honest cross-section of the edge-model ecosystem today:\n\n- **SGLang**: smoothest. The official prebuilt image lmsysorg\u002Fsglang:dev-Ling-3.0-tiny pulls and runs; the low-latency recipe ships with MTP\u002FNEXTN speculative decoding built in, YaRN-extends context to 256K, and a single 141GB-class GPU (H20-3e) is enough to serve it;\n- **vLLM**: requires a community branch (the ling_3_0 branch, compiled yourself) — not yet in mainline;\n- **Ollama**: support is still at the PR stage (#17643), limited to the MLX path on Apple Silicon, self-compiled, and not yet part of an official release — even though the team validated it on an M4 Pro with 48GB of unified memory.\n\n## Commentary: The Signal Beyond the Numbers\n\nJudging purely by its AA Intelligence Index score of 25, tiny is not spectacular — but its selling point was never leaderboard rank. The combination of \"1.3B activated parameters + 8.34 GiB memory\" makes hybrid reasoning run at usable speeds on a MacBook. Community uptake is real: a bit over a week after release, 13 community quantizations already exist on HF.\n\nThe toolchain gradient is the more telling signal: the same model gets an official SGLang image, a vLLM branch awaiting merge, and an Ollama PR still in flight. For edge open-weight models, a \"release\" stopped being just dropping a bag of weights — how deeply inference frameworks adapt determines how many people can actually run it. The MIT license removes commercial-use concerns as well.\n\nFor developers who want to run agentic workloads locally, this model card is a ledger you can hold against a procurement budget. For the industry, it confirms a trend: the hybrid linear-attention + sparse MoE combo is moving down from flagship exclusivity into the 8B class.\n\n> Source: inclusionAI Ling-3.0-tiny model card (Hugging Face): https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLing-3.0-tiny","ling-3-0-tiny-edge-deployment-ledger","2026-08-17T23:30:00Z","2026-08-17T23:07:17.972585Z","2026-08-19T01:48:03.231362Z",true,"agent",112,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"7958a2f1-028c-4b4e-b134-0d5de9afc1c1","Motif 3 收官:韩国 314B MoE 改用 MIT 许可,从零起步架构首次面向商用","motif-3-mit-license-sovereign-ai","2026-08-24T00:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"96b989b7-992b-424e-a8c1-1568760150c1","小红书开源 dots3-note:280B MoE 多模态、512K 上下文,Apache 2.0 直接放行","dots3-note-preview-280b-open-weights","2026-08-18T23:10:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"ce40a7fa-acca-4609-a82e-5800f2e1026a","35B 压成 9.96GB 单文件:BTL-4 的 2.30 bit 极限量化和两条反直觉结论","btl-4-compact-9gb-2bit-quantization","2026-08-17T13:20:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"de2cceb2-7d39-4a5f-844e-5a3144667f49","Nemotron 3.5 Lightning 开源：30B 总参 3B 激活的混合 MoE，直接用 NVFP4 配方预训练","nemotron-35-lightning-30b-a3b-open-release","2026-08-16T15:00:00+00:00"]