[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-dots3-note-preview-280b-open-weights":3,"news-related-96b989b7-992b-424e-a8c1-1568760150c1":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"96b989b7-992b-424e-a8c1-1568760150c1","小红书开源 dots3-note:280B MoE 多模态、512K 上下文,Apache 2.0 直接放行","小红书旗下 dots studio 于 8 月 14 日在 Hugging Face 发布 dots3 家族首个开放权重模型 dots3-note preview:280B 总参数、16B 激活的 MoE 多模态模型,支持 512K token 上下文,可理解文本、图像、视频与音频并输出文本,权重以 Apache 2.0 许可开放,FP8 版本可在单个 8 卡 GPU 节点用 SGLang 或 vLLM 部署,官方自报 SWE-bench Verified 78.4、MMMU Pro 79.1。","8 月 14 日,小红书旗下的 dots studio 把 dots3 家族的第一个开放权重模型 dots3-note preview 挂上了 Hugging Face:280B 总参数、16B 激活的 MoE,上下文直接给到 512K token,许可证是 Apache 2.0。内容平台亲自下场做开源基座模型,这件事本身就值得看一眼。\n\n## 核心规格:一杯\"总参很大、激活很省\"的 MoE\n\n根据官方 model card,dots3-note preview 的架构配置如下:\n\n- **架构**:多模态 MoE,1 层 dense + 45 层 MoE,hidden size 5120;\n- **专家配置**:256 个路由专家 + 1 个共享专家,top-8 路由,280B 总参数、16B 激活参数;\n- **注意力**:13 层 DSA(Top-2048)+ 33 层滑窗注意力(SWA),配比约 1:3;\n- **上下文**:512K token,词表 152K;\n- **多模态输入**:视觉编码器是 MoE ViT(7B 总参、1.2B 激活),音频编码器为 800M dense,支持文本、图像、视频、音频四路输入、文本输出;\n- **MTP**:1 层共享层(1.13B),用于投机解码;\n- **精度**:BF16 \u002F FP8。\n\n官方将其定位为\"家族中最轻量的成员\"——dots3 家族设计上会包含能力、延迟、推理成本不同取舍的多个模型,note preview 只是第一张牌。\n\n## 官方自报评测:编码与多模态都有交代\n\nmodel card 提交到 Hugging Face 的评测数字包括:SWE-bench Verified 78.4、SWE-bench Multilingual 75.7、SWE-bench Pro 61、MMMU Pro(标准 10 选项)79.1、WildClawBench 61.7、Claw-Eval 73.4、HLE 52.6、Apex Agents 30.8。\n\n需要强调:这些均为官方自报数字,第三方复现还要等。但单看分布,它的叙事重心很清楚——编码 agent 与多模态理解两条腿,而不是纯刷通用推理榜。\n\n## 部署:工程协同做得相当前置\n\nFP8 checkpoint 在单个 8 卡 GPU 节点即可跑,vLLM main 分支已原生支持,SGLang(#33829)和 Transformers(#47844)的 PR 还在评审中,官方同步提供了 SGLang Docker 镜像和 vLLM 部署配方。MTP\u002FNEXTN 投机解码是可选项,官方称可将 TPOT 降低超过 50%。ModelScope 同步上架,OpenRouter 上也提供了免费入口。\n\n推理框架在发布日就齐动,说明这不是\"权重一扔就跑\"的开源,工程协同是提前做好的。\n\n## 几点判断\n\n**第一,16B 激活是这类模型真正的卖点。** 280B 总参听起来吓人,但 MoE 的推理成本由激活参数决定,16B 激活把它压进了中杯价位。配上 512K 上下文和多模态输入,model card 明确列出的目标场景是工具调用、多步 agent 工作流和交互式任务——这是照着 agent 时代的接口设计的。\n\n**第二,\"note\" 和 \"preview\" 两个词都值得抠。** 家族最轻量成员 + 预览版,意味着 dots studio 在用小杯探路:先看社区反馈和推理栈适配,再决定后面大杯怎么发。\n\n**第三,内容平台做基座模型,语料是别人没有的牌。** 小红书手握真实、海量的图文与视频多模态数据,dots3-note 的视觉编码器做成了 MoE 结构,大概率是为这种异构输入准备的。Apache 2.0 + OpenRouter 免费试用,把开发者尝试的门槛压到几乎为零。\n\n## 所以呢\n\n对中国开源模型序列来说,这只是又一行加号;但对\"谁有资格做开源多模态基座\"这个问题,小红书交出的答卷是:有真实多模态语料的内容平台同样可以上桌。对开发者而言,趁 preview 免费去 OpenRouter 上跑几个自己的 agent 用例,是当下成本最低的验证方式——第三方复现数据出来之前,自己的负载就是最好的 benchmark。\n\n参考:[Hugging Face model card](https:\u002F\u002Fhuggingface.co\u002Fdots-studio\u002Fdots3-note-prev) · [GitHub 仓库](https:\u002F\u002Fgithub.com\u002Fstudio-dots-ai\u002Fdots3-note-prev)","https:\u002F\u002Fhuggingface.co\u002Fdots-studio\u002Fdots3-note-prev","3a5c4bb0-cd6f-4246-88ef-0fecc1e3f64d",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"d11f0044-8aef-487c-bebe-89ce4683a4a3","moe",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"ea2e4a39-9b6a-4b90-9b2d-c290784dcb94","en","Xiaohongshu's dots3-note: 280B MoE, 512K context, Apache 2.0","dots studio, the AI lab under Xiaohongshu (RedNote), released dots3-note preview on Hugging Face on August 14 — the first open-weight model of the dots3 family. It is a 280B-total\u002F16B-activated multimodal MoE with a 512K-token context window that understands text, images, video, and audio while producing text output. Weights ship under Apache 2.0, the FP8 variant runs on a single 8-GPU node via SGLang or vLLM, and self-reported scores include SWE-bench Verified 78.4 and MMMU Pro 79.1.","On August 14, dots studio — the AI lab under Xiaohongshu (RedNote) — published the first open-weight model of its dots3 family on Hugging Face: dots3-note preview, a 280B-total\u002F16B-activated MoE with a 512K-token context window, released under Apache 2.0. A content platform stepping into open-weight foundation models is itself worth a closer look.\n\n## Core Specs: A \"Big Total, Lean Activation\" MoE\n\nAccording to the official model card, dots3-note preview is configured as follows:\n\n- **Architecture**: Multimodal MoE with 1 dense layer + 45 MoE layers, hidden size 5120;\n- **Experts**: 256 routed experts + 1 shared expert, top-8 routing; 280B total parameters, 16B activated per token;\n- **Attention**: 13 DSA layers (Top-2048) + 33 sliding-window attention (SWA) layers, roughly a 1:3 ratio;\n- **Context**: 512K tokens, 152K vocabulary;\n- **Multimodal inputs**: A MoE ViT vision encoder (7B total, 1.2B activated) and an 800M dense audio encoder; accepts text, image, video, and audio input, produces text output;\n- **MTP**: 1 shared layer (1.13B) for speculative decoding;\n- **Precision**: BF16 \u002F FP8.\n\nThe model is positioned as \"the most lightweight member of the family\" — the dots3 family is designed to include multiple models with different trade-offs among capability, latency, and inference cost, and note preview is simply the first card played.\n\n## Self-Reported Evaluations: Coding and Multimodal Both Covered\n\nEvaluation numbers submitted to the Hugging Face model card include: SWE-bench Verified 78.4, SWE-bench Multilingual 75.7, SWE-bench Pro 61, MMMU Pro (standard 10-option) 79.1, WildClawBench 61.7, Claw-Eval 73.4, HLE 52.6, and Apex Agents 30.8.\n\nOne caveat worth stressing: these are all self-reported numbers, and third-party reproduction will take time. But the distribution tells a clear story — the narrative centers on coding agents and multimodal understanding, not on pure general-reasoning leaderboard climbing.\n\n## Deployment: Engineering Coordination Was Done Up Front\n\nThe FP8 checkpoint runs on a single 8-GPU node. vLLM main already has native support, while the SGLang (#33829) and Transformers (#47844) PRs are still under review; the team simultaneously shipped an SGLang Docker image and a vLLM recipe. MTP\u002FNEXTN speculative decoding is optional and, per the model card, can reduce TPOT by more than 50%. The weights are also on ModelScope, with a free entry point on OpenRouter.\n\nHaving the inference frameworks moving on day one means this is not a \"dump the weights and hope\" release — the engineering coordination was done in advance.\n\n## A Few Judgments\n\n**First, the 16B activation is the real selling point.** 280B total parameters sounds intimidating, but MoE inference cost is determined by activated parameters, and 16B puts this model in mid-tier pricing territory. Combined with the 512K context and multimodal inputs — and with tool use, multi-step agent workflows, and interactive tasks explicitly listed in the model card — this is a model designed as an interface for the agent era.\n\n**Second, both words in \"note preview\" deserve scrutiny.** Family's lightest member + preview build means dots studio is probing with a small cup: watch community feedback and inference-stack compatibility first, then decide how to release the bigger variants.\n\n**Third, a content platform building foundation models holds a card others don't: the corpus.** Xiaohongshu sits on massive, authentic multimodal data — images, text, and video — and the fact that the vision encoder is itself a MoE structure suggests it was built for heterogeneous inputs like these. Apache 2.0 licensing plus a free OpenRouter endpoint pushes the developer's cost of trying it to nearly zero.\n\n## So What\n\nFor the sequence of Chinese open-source models, this is just one more line item. But for the question of \"who gets to build open multimodal foundations,\" Xiaohongshu's answer is: content platforms with authentic multimodal corpora can take a seat at the table too. For developers, running a few of your own agent workloads through the free OpenRouter preview is now the cheapest possible validation — until third-party reproduction numbers arrive, your own workload is the best benchmark.\n\nReferences: [Hugging Face model card](https:\u002F\u002Fhuggingface.co\u002Fdots-studio\u002Fdots3-note-prev) · [GitHub repository](https:\u002F\u002Fgithub.com\u002Fstudio-dots-ai\u002Fdots3-note-prev)","dots3-note-preview-280b-open-weights","2026-08-18T23:10:00Z","2026-08-18T23:09:04.074166Z","2026-08-18T23:09:04.074175Z",true,"agent",145,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"7958a2f1-028c-4b4e-b134-0d5de9afc1c1","Motif 3 收官:韩国 314B MoE 改用 MIT 许可,从零起步架构首次面向商用","motif-3-mit-license-sovereign-ai","2026-08-24T00:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"fb97a60d-69a1-4988-8de6-d1540ba63359","2.4B 参数读懂整页 A4:Cohere Labs 把最小的多模态模型挂上了 Apache 2.0","cohere-north-micro-vision-open-vlm","2026-08-18T13:30:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00"]