[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-nemotron-3-nano-omni-30b-moe-9x-throughput":3,"topics-all":36,"news-related-b26b4b63-5019-4344-a49e-95739f737376":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"b26b4b63-5019-4344-a49e-95739f737376","NVIDIA Nemotron 3 Nano Omni：开源统一多模态模型能否颠覆AI Agent效率边界？","当大多数AI Agent系统还在用多个独立模型分别处理视觉、语音和语言时，NVIDIA直接把它们捏成了一个。4月28日，NVIDIA正式发布Nemotron 3 Nano Omni，这是一款开源的多模态统一模型，基于30B-A3B混合MoE架构，在单一系统内整合了视觉、音频和语言的感知与推理能力。效率是Nemotron 3 Nano Omni最核心的命题。当前Agent系统的典型做法是：为每种模态部署独立模型，推理时数据在多个模型之间来回传递，既增加了延迟，也容易丢失跨模态的上下文关联。NVIDIA用MoE架构将视觉编码器和音频编码器内嵌进同一个模型，用一次前向传播替代过去需要多次调用才能完成的多模态感知。官方数据显示，相比其他开源全模态模型，Nemotron 3 Nano Omni实现了9倍更高的吞吐量，同时保持了同等的交互响应速度。更值得关注的是它的原生高分辨率处理能力。H Company基于该模型构建的电脑使用Agent，使用1920×1080像素的原生输入分辨率进行视觉推理，在OSWorld基准测试中展现出对复杂图形界面的显著理解能力提升。Nemotron 3 Nano Omni在文档智能、音视频理解等6个基准测试leaderboard上位居榜首。模型以开源权重、开源数据集、开源训练技术的方式发布，这意味着整个社区可以验证、复现和定制。NVIDIA将Nemotron 3系列定位为一套完整的基础模型家族：Nano负责多模态感知、Super负责高频执行、Ultra负责复杂规划，三者可以协同工作组成完整的Agent工作流。Nemotron 3 Nano Omni的价值不只是又快又准，而是它代表了一种思路转变：过去我们用模型拼接来解决多模态问题，现在NVIDIA想用模型统一来彻底绕过这个工程债务。","https:\u002F\u002Fblogs.nvidia.com\u002Fblog\u002Fnemotron-3-nano-omni-multimodal-ai-agents\u002F","474eef8c-e0c3-46cf-adee-c089558220f9",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"7ca7ca32-06e4-407e-a7ca-0c1f27a3c229","en","Nemotron 3 Nano Omni: unified multimodal, open efficiency","While most AI Agent systems still use multiple independent models to separately handle vision, voice, and language, NVIDIA has directly merged them into one. On April 28, NVIDIA officially released Nemotron 3 Nano Omni, an open-source unified multimodal model based on 30B-A3B hybrid MoE architecture, integrating visual, audio, and language perception and reasoning within a single system. Efficiency is Nemotron 3 Nano Omni's most core proposition. The typical practice in current Agent systems is to deploy independent models for each modality, with data shuttling between multiple models during inference, increasing latency and easily losing cross-modal context associations. NVIDIA uses MoE architecture to embed the vision encoder and audio encoder into the same model, replacing the multi-modal perception that previously required multiple calls with one forward pass. Official data shows that compared to other open-source full-modal models, Nemotron 3 Nano Omni achieves 9× higher throughput while maintaining equivalent interactive response speed. More noteworthy is its native high-resolution processing capability. H Company's computer-use Agent built on this model uses 1920×1080 native input resolution for visual reasoning, showing significant improvement in complex GUI understanding on the OSWorld benchmark. Nemotron 3 Nano Omni tops the leaderboard on 6 benchmarks including document intelligence and audio-video understanding. The model is released as open-source weights, open-source datasets, open-source training techniques, meaning the entire community can verify, reproduce, and customize. NVIDIA positions the Nemotron 3 series as a complete foundation model family: Nano handles multimodal perception, Super handles high-frequency execution, Ultra handles complex planning, the three can work together to form a complete Agent workflow. Nemotron 3 Nano Omni's value isn't just being fast and accurate, but that it represents a thinking shift: in the past we solved multimodal problems with model stitching, now NVIDIA wants to use model unification to completely bypass this engineering debt.","nvidia-nemotron-3-nano-omni-30b-moe-9x-throughput","2026-04-28T22:10:00Z","2026-04-28T22:06:07.794932Z","2026-08-19T02:08:40.142862Z",true,"agent",177,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"8c28aa2a-10f7-4c0c-b976-e5bf7e781eed","Qwen3.5-397B-A17B发布：千亿MoE架构实现8.6倍解码吞吐提升","qwen3-5-397b-a17b-8-6x-decoding","2026-05-24T10:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"4f70d0fe-f0cb-443e-a4a4-27b9eb004441","1.3B参数多模态模型直接跑在手机上：MiniCPM-V 4.6开源，13亿参数覆盖iOS\u002F安卓\u002F鸿蒙","minicpm-v-4-6-1-3b-phone-multimodal","2026-05-17T13:10:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"bdde1d62-a4a4-4c52-8076-4cb55eef8ff3","Aria发布：全球首款开源多模态原生MoE模型，64K上下文重新定义效率边界","aria-rhymes-ai-25b-moe-multimodal-64k","2026-05-06T07:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"f8a0a967-dee7-4e2a-97de-f7b6bb38ae09","Mistral Small 4：119B MoE 一代三用，Apache 2.0 重新定义开源边界","mistral-small-4-119b-moe-6b-active-apache2","2026-04-25T23:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"c99751d5-418e-49d5-99d3-e43b84c80ec7","IBM与NASA开源月球基础模型:Lunar Foundation Model","nasa-ibm-lunar-foundation-model-sombench","2026-09-19T09:30:00+00:00"]