[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tencent-hy-mt2-33-lang-1-25-bit-440mb":3,"news-related-d2d262c1-6bf0-4a95-b95f-8896fa226db3":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"d2d262c1-6bf0-4a95-b95f-8896fa226db3","腾讯混元 Hy-MT2 翻译家族开源：33 语言 + 1.25-bit 量化","腾讯混元团队 5 月 21 日正式发布并开源了 Hy-MT2 多语言翻译模型家族。模型包括 1.8B、7B 和 30B-A3B（MoE）三档规模，原生支持 33 种语言互译和多种语言下的翻译指令遵循，权重与技术报告已在 GitHub 与 Hugging Face 公开。\n\n从技术路径看，Hy-MT2 延续了混元近期在 Hy3-preview 上验证的「教师-学生」框架：以 Hy3-preview 作为强教师，先做 MT 方向的中训练把通用大模型改造成「擅长翻译」的基础版本，再通过 family-centric 后训练分别微调三档规模。核心方法包括 Reference-Guided On-Policy Distillation、Family-Specific RL 和跨家族蒸馏，使 7B 与 30B 模型在 fast-thinking 模式下超越 DeepSeek-V4-Pro 和 Kimi K2.6。\n\n更值得关注的是 1.8B 端侧模型。它通过 AngelSlim 1.25-bit 极低比特量化，把模型体积压到 440MB、推理速度提升 1.5×，同时整体翻译质量依然优于微软翻译和字节豆包等主流商业 API。这意味着 1.8B 端侧模型已具备「本地替代商用 API」的工程可行性，对翻译 SaaS、跨境电商、本地化工具链都是一次降维打击。在金融、法律、医学等真实业务场景的领域翻译、复杂指令遵循上，Hy-MT2 同样保持稳定领先。这是少有的「开源模型 + 商用级质量 + 端侧可部署」三者同时成立的翻译模型，也印证了混元在 2026 年把「大模型做成基础设施」的产品判断。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.22064","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"bf07a8a7-1a85-48e3-9953-f54e250e1116","en","Hunyuan Hy-MT2 open-sourced: 33 languages, 1.25-bit quant","Tencent's Hunyuan team officially released and open-sourced the Hy-MT2 multilingual translation model family on May 21. The model includes three size tiers — 1.8B, 7B, and 30B-A3B (MoE) — natively supporting 33 languages of mutual translation and translation instruction following in multiple languages, with weights and technical reports public on GitHub and Hugging Face.\n\nFrom a technical-path perspective, Hy-MT2 continues the \"teacher-student\" framework that Hunyuan recently verified on Hy3-preview: using Hy3-preview as a strong teacher, first doing MT-direction medium training to transform a general large model into a \"good at translation\" base version, then through family-centric post-training fine-tuning the three size tiers respectively. The core methods include Reference-Guided On-Policy Distillation, Family-Specific RL, and cross-family distillation, making the 7B and 30B models surpass DeepSeek-V4-Pro and Kimi K2.6 in fast-thinking mode.\n\nMore noteworthy is the 1.8B on-device model. Through AngelSlim 1.25-bit extreme-low-bit quantization, it compresses the model volume to 440MB and improves inference speed by 1.5×, while overall translation quality still outperforms mainstream commercial APIs like Microsoft Translator and ByteDance's Doubao. This means the 1.8B on-device model has the engineering feasibility to \"locally replace commercial APIs,\" a dimensionality-reduction strike against translation SaaS, cross-border e-commerce, and localization toolchains. In real business scenarios such as finance, law, and medicine, as well as complex instruction following, Hy-MT2 also maintains stable leadership. This is one of the few translation models where \"open-source model + commercial-grade quality + on-device deployable\" stand together, also confirming Hunyuan's product judgment in 2026 to \"make large models into infrastructure.\"","tencent-hy-mt2-33-lang-1-25-bit-440mb","2026-05-22T02:00:00Z","2026-06-06T22:16:27.991707Z","2026-08-19T02:08:40.142862Z",true,"agent",122,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"5878a668-282c-4b88-b2b8-7eef40b7938c","LFM2.5-2.6B：2.5GB 内存跑本机 Agent 220 tok\u002Fs","lfm2-5-2-6b-on-device-agent","2026-08-11T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"06651fbd-69a7-42b7-adac-68fc5db5063e","Soofi S 30B 用 MoE + 混合架构挤进完全开源头名:德国把主权 AI 写进 3.2B 激活参数","soofi-s-30b-sovereign","2026-07-13T20:04:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"222d6fbf-e9bc-4d63-8481-88ea28fd499c","Sber GigaChat 3.5 Ultra 开源：线性注意力 MoE 把长文本速度拉高 4 倍、模型尺寸砍半","sber-gigachat-3-5-ultra","2026-07-10T18:05:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"31259851-c64e-47dd-99cf-bcfae698b14f","LFM2.5-Retrievers：Liquid AI 把 LFM「单向」改成「双向 350M」，11 语种检索刷 SOTA","lfm-2-5-retrievers-liquid-350m-bidirectional","2026-06-22T03:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f89d097b-838b-4e1d-a5f6-dc2e6af67fb6","LFM2.5-8B-A1B 开源：1.5B 激活的 MoE 把「边缘 LLM」的天花板再抬一截","lfm-2-5-8b-a1b-liquid-edge-moe-1-5b","2026-06-12T10:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5ce7b0a3-0cfb-4603-8f0a-150afaf0aad9","开源大模型架构分化：MoE与Dense的技术路线之争","moe-vs-dense-open-source-llm-divergence-2026","2026-05-05T05:06:00+00:00"]