[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-nemotron-3-ultra-550b-48-intelligence":3,"news-related-63255594-b16a-4106-9acc-2dc479b97e14":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"63255594-b16a-4106-9acc-2dc479b97e14","英伟达Nemotron 3 Ultra登场：550B开源模型刷新美国开放权重智能榜单","在6月1日台北Computex 2026主题演讲中，英伟达CEO黄仁勋正式发布了Nemotron 3 Ultra——一款拥有约5500亿参数的开源大模型，一举成为美国开放权重模型中的智能新标杆。\n\nNemotron 3 Ultra采用90%稀疏性的MoE（专家混合）架构，实际激活参数约550亿。在Artificial Analysis的智能指数评测中，该模型取得48分，大幅领先Gemma 4 31B（39分）、Nemotron 3 Super（36分）以及gpt-oss-120b（33分），不过仍略低于中国主导的开源前沿模型Kimi K2.6（54分）。\n\n性能方面，DeepInfra内测端点显示Nemotron 3 Ultra推理速度超过300 tokens\u002F秒，而同量级的中国模型（DeepSeek、Moonshot等）通常在50-100 tokens\u002F秒区间。速度优势来自NVFP4量化支持，这与此前Nemotron 3 Super的思路一脉相承。\n\n值得关注的是，Nemotron 3 Ultra定位为「Agentic Coding & Search」场景，这意味着英伟达正将开源大模型从通用对话推向企业级自动化工作流。随着美国开放权重模型在智能和速度上同时逼近闭源前沿，2026年开源与闭源的差距将进一步收窄。","https:\u002F\u002Fblogs.nvidia.com\u002Fblog\u002Fnemotron-3-ultra\u002F","474eef8c-e0c3-46cf-adee-c089558220f9",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"1f3c4db0-0f08-4cf4-a029-d9ffbec88d82","en","NVIDIA Nemotron 3 Ultra: 550B resets the US open-weights bar","NVIDIA released Nemotron 3 Ultra on June 1, a 550B-parameter open-source model. The model is the first US open-source 550B-class model, taking the top spot on several open-weight intelligence leaderboards. The release signals that the US open-source ecosystem is competitive with the closed-source frontier, with NVIDIA as a major contributor.","nvidia-nemotron-3-ultra-550b-48-intelligence","2026-06-01T10:10:00Z","2026-06-01T10:05:43.627369Z","2026-08-19T02:08:40.142862Z",true,"agent",240,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"a8b9d045-0f4c-4596-baa7-060955365877","TensorRT Edge-LLM 0.10.0：Qwen3.8-27B Day-0 上车，边缘 LLM 推理再加速","tensorrt-edge-llm-qwen3-8-27b-day0","2026-08-21T15:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"2b5b7c66-7289-45db-b5b7-dea67882310c","NVIDIA 发布 Nemotron-Labs Diffusion：三模态语言模型统一 AR 与扩散解码","nvidia-nemotron-diffusion-ar-dllm-tri-modal","2026-05-23T04:10:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"c94bdf86-5de9-49fe-8c98-0f5c47611bfe","SGLang v0.5.18 发布:大模型冷启动提速 2.38 倍,710 个 PR 都改了什么","sglang-v0-5-18-cold-start-2-38x","2026-08-24T23:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"1311adb6-dc19-41a7-a188-6760d9e53672","HF Summer 2026 报告:13 个下载量 Top 25 模型是 2022 年的老面孔","hugging-face-summer-2026-attention-adoption","2026-08-24T08:00:00+00:00"]