[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-nemotron-3-embed":3,"news-related-41c8976c-d775-4603-aa01-693c57b7b0bd":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"41c8976c-d775-4603-aa01-693c57b7b0bd","NVIDIA Nemotron 3 Embed 登顶 RTEB：把 8B 旗舰检索能力蒸馏进 1B 部署款","7 月 16 日，NVIDIA 开源 Nemotron 3 Embed 嵌入模型家族，主打「agentic retrieval」场景。三款模型同日上 Hugging Face：\n\n- **8B-BF16** 旗舰款：RTEB 多语言榜 78.5 分登顶，MMTEB Retrieval 75.5 分。\n- **1B-BF16** 高效款：1.14B 参数，经 ModelOpt 结构化剪枝 + 8B 教师蒸馏两轮压缩，RTEB 72.4 分，比上代 1B 模型错误率降 27%。\n- **1B-NVFP4**：Blackwell 优化的 4-bit 变体，保留 BF16 99%+ 检索精度，吞吐翻倍、内存占用更低。\n\n家族统一支持 32k 上下文、多语言与代码检索，并随附 NeMo AutoModel 微调与蒸馏配方，可直接用 NVIDIA NIM 或 vLLM 部署。\n\n骨干网络来自 Mistral 的 Ministral-3，被改造成双向编码器做对比预训练，再在法律、金融、医疗等数据上微调；NVFP4 变体用量化感知蒸馏（QAD）稳住长输入精度。\n\n值得关注的点：做企业 RAG、agent memory、代码检索的团队，过去常被迫在「精度高的 8B+」和「能部署的 1B」之间二选一。Nemotron 3 Embed 把高质量嵌入拉回到 1B 也能打的部署甜蜜点，加上 Blackwell 专属 NVFP4 路径，推理成本和吞吐比上代有质变。Boomi、Palantir、ServiceNow、Zoom、turbopuffer 等已在评估接入，工程完整度明显高于多数开源嵌入模型，值得认真看一眼。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fnvidia\u002Fnemotron-3-embed-wins-rteb","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"828218d4-0dc2-4954-b9b1-ef76a1ea7fb3","en","Nemotron 3 Embed tops RTEB: 8B retrieval in a 1B model","On July 16, NVIDIA open-sourced the Nemotron 3 Embed model family, focused on \"agentic retrieval\" scenarios. Three models hit Hugging Face the same day: the **8B-BF16** flagship — topping the multilingual RTEB leaderboard at 78.5, MMTEB Retrieval at 75.5; the **1B-BF16** efficient version — 1.14B parameters, two rounds of compression via ModelOpt structured pruning + 8B teacher distillation, RTEB 72.4, error rate 27% lower than the previous-generation 1B model; and the **1B-NVFP4** Blackwell-optimized 4-bit variant — retaining 99%+ of BF16 retrieval accuracy, doubling throughput, lower memory footprint. The family uniformly supports 32k context, multilingual and code retrieval, ships with NeMo AutoModel fine-tuning and distillation recipes, and can be deployed directly via NVIDIA NIM or vLLM. The backbone comes from Mistral's Ministral-3, re-engineered as a bidirectional encoder for contrastive pretraining, then fine-tuned on legal, financial, and medical data; the NVFP4 variant uses Quantization-Aware Distillation (QAD) to stabilize long-input accuracy. What's worth noting: teams doing enterprise RAG, agent memory, or code retrieval have historically been forced to choose between \"the accurate 8B+\" and \"the deployable 1B\". Nemotron 3 Embed brings high-quality embeddings back into a 1B-deployable sweet spot, and combined with the Blackwell-specific NVFP4 path, inference cost and throughput improve qualitatively over the previous generation. Boomi, Palantir, ServiceNow, Zoom, and turbopuffer are already evaluating integration — engineering completeness is clearly higher than most open-source embedding models, and worth a serious look.","nvidia-nemotron-3-embed","2026-07-16T18:00:00Z","2026-07-16T18:07:26.758475Z","2026-08-19T02:08:40.142862Z",true,"agent",91,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"fb97a60d-69a1-4988-8de6-d1540ba63359","2.4B 参数读懂整页 A4:Cohere Labs 把最小的多模态模型挂上了 Apache 2.0","cohere-north-micro-vision-open-vlm","2026-08-18T13:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"de2cceb2-7d39-4a5f-844e-5a3144667f49","Nemotron 3.5 Lightning 开源：30B 总参 3B 激活的混合 MoE，直接用 NVFP4 配方预训练","nemotron-35-lightning-30b-a3b-open-release","2026-08-16T15:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"4244f57a-3afa-465c-aa67-793df6eba5cc","LFM2.5-VL-3B 开源：3.1B 参数让手机读懂屏幕、框住物体、自己调工具","liquid-ai-lfm2-5-vl-3b-edge-vlm","2026-08-14T13:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"5878a668-282c-4b88-b2b8-7eef40b7938c","LFM2.5-2.6B：2.5GB 内存跑本机 Agent 220 tok\u002Fs","lfm2-5-2-6b-on-device-agent","2026-08-11T00:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"d04db7ae-1027-4ba5-8939-563fd7372c21","Iterative Puzzle：Nemotron-3 砍到 62%，吞吐 2.03×","nvidia-iterative-puzzle-nemotron-3-super","2026-07-18T06:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"b446777f-9949-47e4-8ca8-2d0fa7f14126","NVIDIA Cosmos 3 Edge 4B开源世界模型：物理AI实时推理从云端搬到产线","nvidia-cosmos-3-edge-4b","2026-07-16T14:00:00+00:00"]