[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen3-5-397b-a17b-8-6x-decoding":3,"topics-all":36,"news-related-8c28aa2a-10f7-4c0c-b976-e5bf7e781eed":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"8c28aa2a-10f7-4c0c-b976-e5bf7e781eed","Qwen3.5-397B-A17B发布：千亿MoE架构实现8.6倍解码吞吐提升","阿里发布Qwen3.5-397B-A17B，这是其Qwen开源家族的最新旗舰模型。该模型采用大规模MoE（混合专家）架构，融合多模态推理与超长上下文支持，成为当前开源领域最具竞争力的Agent与多模态工作负载模型之一。\n\n相比上一代Qwen3-Max，Qwen3.5在解码效率上实现了质的飞跃：官方数据显示，其解码吞吐量提升了8.6倍至19倍。这意味着在同等硬件条件下，Qwen3.5能够服务更多并发请求，对于大规模部署场景意义重大。\n\n在架构层面，Qwen3.5的另一个重要突破在于多模态推理的深度整合。不同于早期模型在文本backbone上外挂视觉模块的做法，Qwen3.5在架构更早阶段就将视觉与语言进行融合，使模型能够在文本、图像、视频和文档之间进行跨模态联合推理。这种「原生多模态」架构通常能带来更好的推理一致性和任务迁移能力。\n\n从行业角度看，8倍以上解码吞吐的提升直接回应了开源社区对高效推理的迫切需求。MoE架构通过条件激活减少计算冗余，而更早的多模态融合则让视觉理解成为语言模型的内生能力而非外挂功能。对于需要在边缘设备或成本敏感场景部署多模态AI的开发者而言，Qwen3.5提供了一条不需要在性能上做过多妥协的路径。当然，这一提升的实际表现还有待开源社区在真实应用场景中验证。","https:\u002F\u002Fwww.bentoml.com\u002Fblog\u002Fnavigating-the-world-of-open-source-large-language-models","4efc0816-0de0-4a2e-bffd-526b65850f91",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"98c852cb-4687-4027-9c83-07dcd57414a3","en","Qwen3.5-397B-A17B: MoE decode throughput up 8.6x","Alibaba released Qwen3.5-397B-A17B, the latest flagship of the Qwen open-source family. The model adopts a large-scale Mixture-of-Experts (MoE) architecture and integrates multimodal reasoning with ultra-long-context support, making it one of the most competitive open-source models for agent and multimodal workloads today.\n\nCompared to the previous-generation Qwen3-Max, Qwen3.5 achieves a qualitative leap in decoding efficiency: official data shows decoding throughput up 8.6× to 19×. On the same hardware, Qwen3.5 can serve far more concurrent requests — significant for large-scale deployment.\n\nArchitecturally, Qwen3.5's other key breakthrough is the deep integration of multimodal reasoning. Unlike early models that bolted vision modules onto a text backbone, Qwen3.5 fuses vision and language at an earlier stage of the architecture, enabling the model to perform cross-modal joint reasoning over text, image, video, and documents. This \"native multimodal\" architecture typically yields better reasoning consistency and task transferability.\n\nFrom an industry perspective, an 8×+ decoding throughput boost directly answers the open-source community's urgent demand for efficient inference. MoE reduces compute redundancy via conditional activation, while earlier multimodal fusion makes visual understanding an innate capability of the language model rather than an add-on. For developers who need to deploy multimodal AI on edge devices or in cost-sensitive scenarios, Qwen3.5 offers a path that doesn't require too many performance compromises. That said, real-world validation from the open-source community in production scenarios is still pending.","qwen3-5-397b-a17b-8-6x-decoding","2026-05-24T10:05:00Z","2026-05-24T10:12:39.995260Z","2026-08-19T02:08:40.142862Z",true,"agent",213,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"4f70d0fe-f0cb-443e-a4a4-27b9eb004441","1.3B参数多模态模型直接跑在手机上：MiniCPM-V 4.6开源，13亿参数覆盖iOS\u002F安卓\u002F鸿蒙","minicpm-v-4-6-1-3b-phone-multimodal","2026-05-17T13:10:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"bdde1d62-a4a4-4c52-8076-4cb55eef8ff3","Aria发布：全球首款开源多模态原生MoE模型，64K上下文重新定义效率边界","aria-rhymes-ai-25b-moe-multimodal-64k","2026-05-06T07:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"b26b4b63-5019-4344-a49e-95739f737376","NVIDIA Nemotron 3 Nano Omni：开源统一多模态模型能否颠覆AI Agent效率边界？","nvidia-nemotron-3-nano-omni-30b-moe-9x-throughput","2026-04-28T22:10:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"f8a0a967-dee7-4e2a-97de-f7b6bb38ae09","Mistral Small 4：119B MoE 一代三用，Apache 2.0 重新定义开源边界","mistral-small-4-119b-moe-6b-active-apache2","2026-04-25T23:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"c99751d5-418e-49d5-99d3-e43b84c80ec7","IBM与NASA开源月球基础模型:Lunar Foundation Model","nasa-ibm-lunar-foundation-model-sombench","2026-09-19T09:30:00+00:00"]