[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-spark-x2-5-4b-apache-open-source":3,"topics-all":41,"news-related-4c4a2a9e-f69b-4985-bd42-97ab2ef4e2ac":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"4c4a2a9e-f69b-4985-bd42-97ab2ef4e2ac","Spark-X2.5-4B 开源:4B 跑 1M 上下文,22 项基准打 9B 级 Qwen3.5","XHToken 把 Spark-X2.5-4B\u002F1.7B 推上 Hugging Face，Apache 2.0 开源。4B 在 22 项 benchmark 15 项超 Qwen3.5 同级；hybrid attention + 1M 上下文；NVIDIA\u002F华为昇腾\u002FHygon 全栈支持。","SparkLLM 团队（XHToken）把 Spark-X2.5-4B 和 1.7B 同步推上 Hugging Face，Apache 2.0 开源。4B 这个量级，过去默认是「能跑但别期望太多」的玩具。这次给的实测表里，Spark-X2.5-4B 在 22 项 benchmark 的 15 项上超过 Qwen3.5-4B、Qwen3.5-9B、Gemma4-12B 同级对手。这不是简单刷榜，是端侧 agent 工作流真正可以用的拐点。\n\n## 架构：hybrid attention + 原生 1M 上下文\n\n模型卡说清楚 hybrid attention：每 4 层里 1 层做 full attention，剩下 3 层用 sliding window attention（SWA）。full 层负责跨段语义，SWA 层压住 KV cache 爆炸。Spark-X2.5 直接给出原生 1M token 上下文，是从训练后期拿百亿级 1M 长度样本专门 stage 训出来的。Qwen3、GLM-5.3 走的也是这条路，但 SparkLLM 把这件事压到了 4B 端侧可跑的量级。\n\n## 训练：20T 预训练 + MOPD 合并多能力 RL 教师\n\n预训练 20T token 是头部开源 LLM 的标配。post-training 走 SFT 起步 + 大规模 RL 收尾，最后用 MOPD 把几个能力域的「教师策略」合并回一个可部署模型。「多能力 RL 蒸馏合并」是 2026 年开源端侧模型标准动作。\n\n## Agent benchmark 翻倍\n\nτ³-bench 上 30.4（Qwen3.5-4B 是 6.7），BrowseComp 40.9（Qwen3.5-4B 14.3），MCP-Atlas 54.6（Qwen3.5-4B 40.8），Workspace Bench 31.2（Qwen3.5-4B 21.3）。这些是「多步工具调用 + 长时规划 + 跨 app 工作流」的代理评测，不是 MMLU 问答测试。SWE-Bench Pro 44.4 对 Qwen3.5-4B 29.4 也明显拉出差距。所有评测都开了 thinking mode，温度 1.0、top_p 0.95。\n\n## 推理端：vLLM\u002FSGLang\u002Fllama.cpp\u002FMLX 全栈支持\n\nvLLM、SGLang、llama.cpp、MLX 都已支持，NVIDIA、华为昇腾、Hygon、HOUMO.AI 都能跑，Ollama 和 LM Studio 也接好。官方写的「superior TTFT, TOPT」是合规的「同级最优」措辞。\n\n## 4B 跑 1M context 不再是营销话术\n\nQwen3-4B-Instruct 那一代「1M context」基本是位置编码插值 + 检索增强的合成数字，实际可用窗口在 32K-128K 之间就掉得厉害。Spark-X2.5 训练阶段 push 到 1M 序列长度，加上 hybrid attention 把 KV cache 砍掉 ~3\u002F4，意味着 4B 模型在 64GB 内存的 Mac Studio 上可以真正维持 1M 上下文做 agent 工作流。配合 NVIDIA\u002F华为昇腾两条路，端侧 agent 的硬件门槛被压低一档。\n\n要给开源端侧模型排个 2026 第四季度的「值得看」清单，Spark-X2.5-4B 一定在前列。看社区怎么基于它做 fine-tune，以及那 22 个 benchmark 之外的真实工作流扛不扛得住。","https:\u002F\u002Fhuggingface.co\u002FXHToken\u002FSpark-X2.5-4B","24d5c6c5-6573-4180-a1fd-f1459842d1af",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":25,"name":26,"slug":26,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"4d0d5e38-61db-467f-b342-cb03a9b34798","en","Spark-X2.5-4B open-sourced: 4B runs 1M context, beats 9B-class Qwen3.5 on 22 benchmarks","XHToken pushes Spark-X2.5-4B\u002F1.7B to Hugging Face under Apache 2.0. The 4B model surpasses Qwen3.5 peers on 15 of 22 benchmarks, with hybrid attention enabling native 1M-token context and full-stack support for NVIDIA, Huawei Ascend, and Hygon hardware.","SparkLLM (XHToken) released Spark-X2.5-4B and 1.7B on Hugging Face under Apache 2.0. At 4B parameters — a class long dismissed as a workable-but-underwhelming toy tier — Spark-X2.5-4B's published benchmark table shows it beating Qwen3.5-4B, Qwen3.5-9B, and Gemma4-12B on 15 of 22 evaluations. This is not a leaderboard stunt; it is a real turning point for on-device agent workflows.\n\n## Architecture: hybrid attention + native 1M context\n\nThe model card spells out the hybrid attention design: 1 full-attention layer per 4 layers, with the other 3 using sliding window attention (SWA). The full layer carries cross-segment semantics; SWA keeps the KV cache from exploding. Spark-X2.5 ships native 1M-token context — not via position interpolation or paper-only tricks like Ring Attention, but from a dedicated training stage that fed hundreds of billions of 1M-length samples. Qwen3 and GLM-5.3 took the same path, but SparkLLM compressed the recipe to a 4B tier that fits on-device.\n\n## Training: 20T pretraining + MOPD multi-capability RL teacher merge\n\nPretraining on 20T tokens is now table stakes for open-source frontier LLMs. Post-training starts with SFT, then runs large-scale RL, and finally uses MOPD (Multi-Objective Policy Distillation) to merge several capability-domain teacher policies back into a single deployable model. The card emphasises that reasoning, coding, agentic, and instruction-following capabilities all improved together. Multi-capability RL-distillation merge is now the standard 2026 playbook for open on-device models, and Qwen3.5 is on the same page.\n\n## Agent benchmarks double or more\n\nA few data points stand out. τ³-bench: 30.4 (vs Qwen3.5-4B at 6.7). BrowseComp: 40.9 (vs 14.3). MCP-Atlas: 54.6 (vs 40.8). Workspace Bench: 31.2 (vs 21.3). These are multi-step tool-use, long-horizon planning, and cross-app workflow evaluations, not MMLU-style Q&A. SWE-Bench Pro at 44.4 vs 29.4 also opens a clear gap. Worth noting: all evaluations used thinking mode, temperature 1.0, top_p 0.95 — the 4B model is doing real reasoning, not fast-path direct answers.\n\n## Inference: vLLM, SGLang, llama.cpp, MLX full-stack support\n\nvLLM, SGLang, llama.cpp, and MLX all support Spark-X2.5. NVIDIA, Huawei Ascend, Hygon, and HOUMO.AI hardware can all run it. Ollama and LM Studio integrations are also in place. The card's 'superior TTFT, TOPT' framing is the kind of compliant 'best-in-class' wording that fits Apache 2.0 self-reported numbers.\n\n## 4B running 1M context is no longer marketing\n\nQwen3-4B-Instruct and its 1M-context peers were effectively position-interpolation plus retrieval-augmentation synthetics — usable context collapsed somewhere between 32K and 128K. Spark-X2.5 pushed 1M sequence length into actual training, and the hybrid attention cuts KV cache by roughly 3\u002F4, which means a 4B model on a 64GB-memory Mac Studio can sustain real 1M-context agent workflows — read an entire Notion knowledge base for RAG, scan a long code repository, plus a 1M tool-output buffer. With both NVIDIA and Huawei Ascend paths, the hardware bar for on-device agents has been lowered again.\n\nFor any Q4 2026 shortlist of open-source on-device models worth watching, Spark-X2.5-4B is on it. The 4B parameter floor means 16GB-memory Mac minis and 32GB-memory consumer PCs can run it; native 1M context plus strong agent benchmark scores mean it is more than a chat toy; Apache 2.0 means it is commercially usable and modifiable. The real questions are how the community fine-tunes on top of it, and whether the workflows outside those 22 benchmarks actually hold up under real load.","spark-x2-5-4b-apache-open-source","2026-09-16T01:30:00Z","2026-09-16T01:08:52.490297Z","2026-09-16T01:08:52.490312Z",true,"agent",35,[42,51],{"slug":43,"tag_slug":43,"title_zh":44,"title_en":45,"intro_zh":46,"intro_en":47,"id":48,"is_active":38,"created_at":49,"modified_at":50},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":52,"tag_slug":52,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":38,"created_at":58,"modified_at":59},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"365b770a-2cca-40a0-beb0-1eff823702c0","IBM 开源 Granite Time Series PatchTST-FM-r2:零样本 SOTA,Apache 2.0 商用许可","ibm-granite-patchtst-fm-r2-zero-shot-apache","2026-09-12T11:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"c745abb4-d608-4ea6-884b-5176d7134d71","IFM 开源 K2 Horizon 六模型：训练数据全放，7B 刷榜成绩 82 被自己砍到 70.6","ifm-k2-horizon-open-fleet-audit","2026-09-05T23:07:55+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"741bd34c-7134-4e8e-ab45-4f53dc576a6b","腾讯 Hy4 登顶 9 月开源榜:79.87 分超 Qwen3.8 Max,Anthropic 包揽总榜前三","tencent-hy4-tops-open-source-benchlm-september","2026-09-01T17:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"d4fa7e14-8fbd-4940-93a6-3dd6f0a3991d","DeepSeek V4 Pro 正式版：1.6T MoE，1M 上下文","deepseek-v4-pro-0813-ga-1m-context-moe","2026-08-13T02:00:00+00:00"]