[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-motif-3-mit-license-sovereign-ai":3,"news-related-7958a2f1-028c-4b4e-b134-0d5de9afc1c1":41},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"7958a2f1-028c-4b4e-b134-0d5de9afc1c1","Motif 3 收官:韩国 314B MoE 改用 MIT 许可,从零起步架构首次面向商用","韩国 Motif Technologies 把 314B MoE 模型 Motif 3 的最终权重换上 MIT 许可,加上自研 GDLA 注意力与 PolyNorm,首次允许全球商用、二次预训练与商业微调,目标是南韩 Dokpamo 主权 AI 计划。","## 一个没发新闻稿的发布\n\n8 月 13 日,韩国 Motif Technologies 把 Motif 3 的最终权重悄悄上到了 Hugging Face。没有博客,没有 X 上的官方通告,只是三个仓库在几分钟内同时出现 —— Motif-3-Base、指令微调的 Motif-3 和量化版 Motif-3-NVFP4,外加一份同时发布的 arXiv 技术报告 [arxiv.org\u002Fabs\u002F2608.09119](https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.09119)。\n\n真正让开发者社区炸锅的不是参数规模,而是许可证:三份权重全部从 7 月 beta 的\"仅供研究\"切换到 MIT,从此商业使用、二次预训练、再分发都不再需要首尔点头。\n\n## 数字看上去很熟悉,但架构是自研的\n\nMotif 3 是 314B 总参数 \u002F 13.2B 激活参数的 MoE,8 个专家 + 1 个共享专家、约 12.5 万亿训练 token、原生 256K 上下文。论文与官方模型卡 [huggingface.co\u002FMotif-Technologies\u002FMotif-3](https:\u002F\u002Fhuggingface.co\u002FMotif-Technologies\u002FMotif-3) 把规格摆得很细。\n\n值得划重点的是三个自研组件 —— 不是任何已有开源架构的再参数化:\n\n- **GDLA(Grouped Differential Latent Attention)**:把 Microsoft Research 与清华 2024 年的 Differential Transformer(用两组 softmax 注意力相减压制\"注意力沉没\")与 DeepSeek-V2 的 MLA(KV 压缩)结合起来,80 个 query 头配 16 个 KV 头,既降噪声又把 KV 缓存压住。\n- **Expert-Specific PolyNorm**:把 MoE 每个专家内部的 SiLU 换成可学习的逐专家多项式归一化,目标是缓解大模型训练中出 activation outlier 的老毛病。\n- **mHC(manifold-constrained hyper-connections)**:把普通残差加法换成四条并行残差流的 doubly-stochastic 混合。配上一个 MTP head,做 inference 时能自我推测解码。\n\n这个\"全部自己来\"的要求,是韩国政府 Dokpamo 计划的硬条件 —— 据 [TechTimes 报道](https:\u002F\u002Fwww.techtimes.com\u002Farticles\u002F324260\u002F20260813\u002Fmotif-3-final-release-mit-license-opens-koreas-sovereign-ai-builders.htm),Naver Cloud 与 NC AI 因为在模型里塞了 Qwen 冻结编码器权重被踢出局,1 月份就已出局。\n\n## 跑分、吞吐量与一句话盘点\n\nArtificial Analysis Intelligence Index 给 Motif 3 打 **47**,在 Dokpamo 四队里排第一,比 Upstage Solar Open 2(37)、SKT A.X K2(35)、LG K-EXAONE 2.0(31) 都高(口径见 BigGo Finance 的 Dokpamo 榜单)。厂商自报数据:τ²-Bench Telecom 94.7、SWE-Bench Verified 76.2、Terminal-Bench 2.1 74.9、GPQA Diamond 83.4 —— GPQA 这个分数接近闭源前沿 88-93 的下沿,SWE-Bench Verified 76.2 也能挤进开放权重第一梯队。\n\n但有两点要看见:\n\n1. ** 第三方独立复现还没完成**。OrcaRouter 在独立分析里写得很直白:\"这些指标是 Motif 自评的,外部还没跑完\"。这点对任何新模型都正常,但写在标题党之前很重要。\n2. ** Motif 输出非常啰嗦**。最终版跑 2.6 亿 token,对照同档开放权重中位数 1 亿 —— 推理成本估算要把这条算进去。\n\n部署要求 8 张 B200 或 H200,需要 Motif 自家 vLLM 分支(`MotifTechnologies\u002Fvllm`,Docker 镜像 `ghcr.io\u002Fmotiftechnologies\u002Fvllm:v0.20.2-motif3.rc3`),目前没有任何推理服务商把它挂上去。\n\n## 30 个人、5 个月与一份主权 AI 的最小可行证明\n\nMotif Technologies 母公司是 Moreh,2020 年成立,2023 年拿了 AMD 与韩国电信投的 2200 万美元 B 轮。Motif 本身今年 2 月才成立,5 月 B 轮融资 1690 万美元(韩元约 240 亿),整个团队约 30 人。韩国政府通过 Dokpamo 划了约 768 张 B200 给他们。\n\n5 月被选中,8 月 4 日向 NIPA 提交最终版 —— 从立项到能用的开放权重一共 5 个月。这大概是 2026 年\"小型团队 + 国家级算力 + 严格原研约束\"组合跑出来的最具说服力的样本之一。如果年底 Dokpamo 的两支胜出队伍要承担\"AI for All\"国家助手 50% 推理的硬性要求,那 Motif 的 MIT 权重已经是全球开发者可以现在就接住的入口。\n\n## 所以呢\n\n对一个被 DeepSeek \u002F Qwen \u002F Kimi \u002F Llama 把\"开放权重\"卷成标配的行业来说,Motif 3 真正的意义不在分数 —— 它给出了一个干净的、原研架构 + 顶级宽松许可证 + 国家主权背书的组合样本。对国内团队:别再死磕许可证分发了,Dokpamo 这种\"30 人 5 个月做出全球第一梯队\"的模式比千亿参数竞赛更值得复盘。","https:\u002F\u002Fwww.techtimes.com\u002Farticles\u002F324260\u002F20260813\u002Fmotif-3-final-release-mit-license-opens-koreas-sovereign-ai-builders.htm","4f2dc39f-0b6a-48e6-ad47-da9c3c15cbea",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"62a0408c-027e-4515-807d-5f19dc5e1390","korean",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"d11f0044-8aef-487c-bebe-89ce4683a4a3","moe",{"id":25,"name":26,"slug":26,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"f6566395-7b3a-4ef8-b92f-6cc35da0fbf2","en","Motif 3 ships under MIT: Korea's 314B MoE goes commercial with self-developed architecture","Motif Technologies of South Korea has switched the final Motif 3 314B MoE weights to MIT license, adding self-developed GDLA attention and PolyNorm, opening global commercial use, continued pretraining, and commercial fine-tuning — the build target is Korea's Dokpamo sovereign AI program.","## A launch without a press release\n\nOn August 13, 2026, South Korea's Motif Technologies quietly pushed the final Motif 3 weights to Hugging Face. No blog post, no announcement on X — just three repositories appearing within minutes of each other: Motif-3-Base, the instruction-tuned Motif-3, and the quantized Motif-3-NVFP4, plus an arXiv technical report released at the same time ([arxiv.org\u002Fabs\u002F2608.09119](https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.09119)).\n\nWhat blew up the developer community was not the parameter count but the license: all three repositories switched from the July beta's research-only terms to MIT, unlocking commercial use, continued pretraining, and redistribution without Seoul's blessing.\n\n## Familiar numbers, self-developed architecture\n\nMotif 3 is a 314B total \u002F 13.2B active Mixture-of-Experts model with 8 routed experts plus 1 shared expert, around 12.5 trillion training tokens, and a native 256K context window. The paper plus the official model card on [huggingface.co\u002FMotif-Technologies\u002FMotif-3](https:\u002F\u002Fhuggingface.co\u002FMotif-Technologies\u002FMotif-3) lay out the full specification.\n\nThree self-developed components are worth highlighting — none are reparameterizations of existing open-source architectures:\n\n- **GDLA (Grouped Differential Latent Attention)**: combines Microsoft Research \u002F Tsinghua's 2024 Differential Transformer (two softmax attention maps subtracted to cancel attention-sink noise) with DeepSeek-V2's MLA (KV-cache compression). The configuration uses 80 query heads with 16 KV heads, suppressing noise while compressing the KV cache.\n- **Expert-Specific PolyNorm**: replaces the SiLU activation inside each expert's feed-forward block with a learned per-expert polynomial normalization, designed to tame activation outliers — the well-known numerical instability at large scale.\n- **mHC (manifold-constrained hyper-connections)**: replaces standard residual additions with a doubly-stochastic mixing of four parallel residual streams. Combined with an MTP head, this enables self-speculative decoding at inference time.\n\nThe \"everything from scratch\" rule is the hard constraint of South Korea's government Dokpamo program. According to [TechTimes coverage](https:\u002F\u002Fwww.techtimes.com\u002Farticles\u002F324260\u002F20260813\u002Fmotif-3-final-release-mit-license-opens-koreas-sovereign-ai-builders.htm), Naver Cloud and NC AI were disqualified in January for incorporating frozen Qwen encoder weights.\n\n## Benchmarks, throughput, and a one-line scoreboard\n\nArtificial Analysis's Intelligence Index gives Motif 3 a **47**, leading the four Dokpamo teams: Upstage Solar Open 2 at 37, SK Telecom A.X K2 at 35, and LG AI Research K-EXAONE 2.0 at 31 (per BigGo Finance's Dokpamo ranking). Vendor-reported numbers: τ²-Bench Telecom 94.7, SWE-Bench Verified 76.2, Terminal-Bench 2.1 74.9, GPQA Diamond 83.4. The GPQA Diamond figure sits near the lower bound of the closed-source frontier (88–93), and the SWE-Bench Verified 76.2 puts Motif 3 in the top tier of open-weight models on software engineering tasks.\n\nTwo caveats worth flagging. First, **third-party independent reproduction is not complete**. OrcaRouter's independent write-up is direct about this: the figures are Motif's self-evaluation, not yet externally verified — a normal caveat for any newly released model, but worth stating before any headline. Second, **Motif 3 is verbose**: the final release generated 260 million output tokens in evaluation versus a comparable open-weight peer median of 100 million. Deployment teams should factor this into inference-cost estimates.\n\nRunning Motif 3 requires 8 NVIDIA B200 or H200 GPUs, a fork of vLLM (`MotifTechnologies\u002Fvllm`, Docker image `ghcr.io\u002Fmotiftechnologies\u002Fvllm:v0.20.2-motif3.rc3`), and at the moment no inference provider has any of the three repositories online.\n\n## 30 people, 5 months, and a minimum viable proof of sovereign AI\n\nMotif Technologies is a subsidiary of Moreh, founded in 2020 with a $22M Series B in 2023 from AMD and KT (Korea Telecom). Motif itself was spun out in February 2025 and raised approximately $16.9M (₩24B) in a Series B in May 2026 from NICE Investment Partners, Nautilus Investment, Ditto Investment, and Forest Ventures. The whole team is around 30 people, and the South Korean government allocated roughly 768 NVIDIA B200 GPUs through the Dokpamo program.\n\nSelected in May, submitted the final model to NIPA on August 4 — five months from selection to a usable open-weight release. That is one of the more concrete proof-of-concept results for \"small team + national compute + strict originality constraint\" combinations this year. The Dokpamo competition's two eventual winners will be the default supply layer for South Korea's \"AI for All\" mandate requiring at least 50% of inference for the national assistant to run on domestic models, covering all 51 million citizens. Motif's MIT weights are now the entry point the global developer ecosystem can pick up immediately.\n\n## So what\n\nFor an industry where DeepSeek \u002F Qwen \u002F Kimi \u002F Llama have turned \"open weights\" into table stakes, Motif 3's significance is not its benchmark numbers — it is a sample of the cleanest combination of self-developed architecture, top-tier permissive licensing, and national sovereignty endorsement. For Chinese teams: stop treating license distribution as the moat. The Dokpamo model — 30 people, 5 months, top-tier global benchmark — is worth more attention than yet another 100B-parameter race.","motif-3-mit-license-sovereign-ai","2026-08-24T00:00:00Z","2026-08-24T09:05:05.142212Z","2026-08-24T09:05:05.142223Z",true,"agent",39,{"items":42},[43,48,53,58,63,68],{"id":44,"title":45,"news_slug":46,"published_at":47},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"68072ee1-fc37-4064-ab18-09550ae72d1b","GLM-5.3-Flash 把 320B MoE 跑在国产芯片上:Flash 价位和 $0.15 API 的混合注意力栈","glm-5-3-flash-chinese-chips-hybrid-attention","2026-08-27T03:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"e75069c6-f15c-4ff9-8b11-404d705442e8","Upstage Solar Pro 4:把「agent 跑得稳」做成新一代闭源模型卖点","upstage-solar-pro-4-agent-reliability-closed-llm","2026-08-25T03:00:00+00:00"]