[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-aria-rhymes-ai-25b-moe-multimodal-64k":3,"topics-all":36,"news-related-bdde1d62-a4a4-4c52-8076-4cb55eef8ff3":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"bdde1d62-a4a4-4c52-8076-4cb55eef8ff3","Aria发布：全球首款开源多模态原生MoE模型，64K上下文重新定义效率边界","2026年5月，Rhymes AI团队正式发布Aria，这是全球首个开源的多模态原生混合专家（MoE）模型。与过往通过外挂视觉编码器拼接多模态能力的做法不同，Aria从架构设计之初就将视觉理解融入统一的海量Token空间，被视为多模态大模型开源生态的一次重要突破。\n\nAria总计拥有25.3B参数，但每次推理仅激活3.9B参数——这种稀疏激活机制让MoE架构的效率优势体现得淋漓尽致。相比同参数量级的稠密模型，Aria在保持高质量输出的同时，大幅降低了计算资源和显存占用，在单张A100（80GB）GPU上即可完成bfloat16精度的加载与推理。\n\n更值得关注的是其64K Token的多模态上下文窗口。传统多模态模型在处理长视频或大型文档时，往往受限于上下文长度或出现理解断层。Aria通过统一Token空间的设计，让文本、代码、图像和视频共享同一个语义表示体系，有效避免了跨模态信息丢失的问题。从实际评测看，无论是视频理解、文档分析还是多轮对话，Aria在多个基准测试中的表现都稳居开源多模态模型前列。\n\nAria不仅公开了模型权重，还同步释出了完整的技术报告和微调工具链，支持LoRA和全参数微调，开发者可以在消费级GPU上完成垂直场景的定制训练。这对医疗影像、工业文档理解等特定领域的需求降低门槛。\n\nAria的出现，回应了一个行业痛点：开源社区在多模态能力上长期落后于闭源模型，尤其是原生多模态——视觉和语言从架构层面深度融合而非简单拼接。Google的Gemini系列、OpenAI的GPT-4V在这点上构建了很高的技术壁垒。Aria以开源姿态首次在架构层面接近这一水准，对整个生态的推动意义不容小觑。不过，稀疏激活带来的路由开销、多模态统一表示的训练成本，都是Rhymes AI后续需要持续优化的方向。","https:\u002F\u002Fhuggingface.co\u002Frhymes-ai\u002FAria","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"3281693c-2890-4eb2-8cc3-692ea4690318","en","Aria: the first open natively-multimodal MoE, 64K context","In May 2026, the Rhymes AI team officially released Aria, the world's first open-source multimodal-native Mixture-of-Experts (MoE) model. Unlike the previous approach of bolting on multimodal capability via external vision encoders, Aria integrates visual understanding into a unified massive-Token space from the start of architecture design, seen as an important breakthrough for the multimodal-LLM open-source ecosystem.\n\nAria has 25.3B total parameters, but only activates 3.9B per inference — this sparse-activation mechanism lets the MoE architecture's efficiency advantage shine. Compared to dense models of the same parameter scale, Aria maintains high-quality output while significantly reducing compute resources and VRAM usage, with full bfloat16-precision loading and inference on a single A100 (80GB) GPU.\n\nMore noteworthy is its 64K Token multimodal context window. Traditional multimodal models often face context-length limitations or understanding breaks when processing long videos or large documents. Aria, through its unified-Token-space design, lets text, code, images, and video share the same semantic representation system, effectively avoiding cross-modal information loss. From actual evaluation, whether video understanding, document analysis, or multi-turn dialogue, Aria's performance in multiple benchmarks sits steadily at the front of the open-source multimodal-model pack.\n\nAria not only publicly released model weights, but also simultaneously released a complete technical report and fine-tuning toolchain, supporting LoRA and full-parameter fine-tuning, allowing developers to complete vertical-scenario customized training on consumer-grade GPUs. This lowers the threshold for needs in specific domains like medical imaging and industrial document understanding.\n\nAria's emergence responds to an industry pain point: the open-source community has long lagged behind closed-source models in multimodal capability, especially native multimodality — vision and language deeply integrated at the architecture level rather than simply concatenated. Google's Gemini series and OpenAI's GPT-4V have built high technical barriers in this regard. Aria, with an open-source posture, is the first to approach this level at the architecture level, and the promotion significance for the entire ecosystem cannot be underestimated. However, the routing overhead from sparse activation, and the training cost of unified multimodal representation, are directions Rhymes AI needs to continuously optimize.","aria-rhymes-ai-25b-moe-multimodal-64k","2026-05-06T07:00:00Z","2026-05-06T07:08:46.174669Z","2026-08-19T02:08:40.142862Z",true,"agent",204,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"8c28aa2a-10f7-4c0c-b976-e5bf7e781eed","Qwen3.5-397B-A17B发布：千亿MoE架构实现8.6倍解码吞吐提升","qwen3-5-397b-a17b-8-6x-decoding","2026-05-24T10:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"4f70d0fe-f0cb-443e-a4a4-27b9eb004441","1.3B参数多模态模型直接跑在手机上：MiniCPM-V 4.6开源，13亿参数覆盖iOS\u002F安卓\u002F鸿蒙","minicpm-v-4-6-1-3b-phone-multimodal","2026-05-17T13:10:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"b26b4b63-5019-4344-a49e-95739f737376","NVIDIA Nemotron 3 Nano Omni：开源统一多模态模型能否颠覆AI Agent效率边界？","nvidia-nemotron-3-nano-omni-30b-moe-9x-throughput","2026-04-28T22:10:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"f8a0a967-dee7-4e2a-97de-f7b6bb38ae09","Mistral Small 4：119B MoE 一代三用，Apache 2.0 重新定义开源边界","mistral-small-4-119b-moe-6b-active-apache2","2026-04-25T23:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"c99751d5-418e-49d5-99d3-e43b84c80ec7","IBM与NASA开源月球基础模型:Lunar Foundation Model","nasa-ibm-lunar-foundation-model-sombench","2026-09-19T09:30:00+00:00"]