[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-vera-cpu-agent-orchestration":3,"news-related-30a147aa-e3ed-475c-b15f-9e5ffce6ffc9":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"30a147aa-e3ed-475c-b15f-9e5ffce6ffc9","英伟达 Vera CPU：DeepInfra 实测 Agent 编排提速 2.2 倍","7 月 21 日，英伟达发布 Vera Rubin 平台在 CoreWeave、Google Cloud、Microsoft Azure 和 Mistral 等合作伙伴处的实测成绩，把行业关注点从 GPU 单卡性能拉回到一个被忽视的角色——CPU。\n\nDeepInfra 的生产基准显示，Vera CPU 在同等 QoS 下可支撑 **1.6 倍并发 AI Agent**，编排速度比对比 CPU **快 2.2 倍**。这家 AI 云平台每周处理近 5 万亿 Token，其中 **约 30% 来自 Agentic 工作负载**——数字本身就解释了为什么 CPU 突然成了 AI 工厂的瓶颈：Agent 不是单次推理，而是「模型调用 + 工具调度 + 上下文切换 + 任务编排」的循环，每一步都在 CPU 上跑。\n\nVera CPU 的关键设计都瞄准这个新负载：自研 Olympus 核心带来 **2 倍单线程性能、3 倍核间带宽、40% 内存延迟降低**。这跟传统数据中心 CPU「核多频低」的路线完全相反——Agent 编排不需要堆核，需要的是单核能把一次推理前后的协调工作压到最短。\n\n平台层面，Vera Rubin NVL72 在 CoreWeave 的 DeepSeek-R1 实测中实现 **每兆瓦 10 倍 Token 吞吐**，相比 Grace Blackwell NVL72 是数量级跃迁。这个提升来自七颗芯片的 co-design：Rubin GPU + Vera CPU + NVLink 6 + ConnectX-9 + BlueField-4 + Spectrum-6 + Groq 3 LPX，不是「拼一台服务器」，是把整机柜当成一颗加速器。NVLink 6 提供 260 TB\u002Fs all-to-all 带宽，专门解决 MoE 架构的专家路由瓶颈；Spectrum-X 把 RDMA 带宽拉到比通用以太网高 1.6 倍。\n\n工程细节也透露了规模化野心：compute tray **无缆、无风扇、无软管**，装配时间从小时级压缩到 **1 分钟**；液冷进液温度做到 45°C，干冷器直冷免掉冷水机组。这不是炫技，是把「机柜级 AI 工厂」真正变成可批量复制的产品。微软与 Mistral 的数十亿美元合作也将基于数千颗 Vera Rubin GPU 落地，主打欧洲主权 AI。\n\n所以呢？当所有人盯着 GPU 算力时，Agent 工作负载第一次把 CPU 推到了 AI 工厂性能曲线的 C 位。Vera CPU 的真正信号不是「比对手快多少」，而是英伟达用一颗自研 CPU 告诉市场：未来 AI 基础设施的瓶颈不在「谁卡算得快」，在「谁能把每次工具调用、上下文组装、多步推理的协调开销压到最低」。OpenAI、Anthropic 在 Agent 框架上的军备竞赛，最终会反推到硅这一层。","https:\u002F\u002Fblogs.nvidia.com\u002Fblog\u002Fvera-rubin\u002F","474eef8c-e0c3-46cf-adee-c089558220f9",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":21,"name":22,"slug":22,"description":13,"color":13},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"52fc8608-de65-4eda-aa2c-048d9ede44b1","en","NVIDIA Vera CPU: DeepInfra logs 2.2x agent orchestration","On July 21, NVIDIA released Vera Rubin platform benchmark results from partners CoreWeave, Google Cloud, Microsoft Azure and Mistral, pulling industry attention from single-GPU performance back to a neglected role — the CPU. DeepInfra's production benchmarks show that, at equal QoS, the Vera CPU can sustain **1.6× concurrent AI Agents** and run orchestration **2.2× faster** than the comparison CPU. The AI cloud platform processes nearly 5 trillion tokens per week, **~30% from Agentic workloads** — the number itself explains why the CPU suddenly became the AI factory's bottleneck: an Agent isn't a single inference, it's a loop of \"model call + tool scheduling + context switching + task orchestration\", and every step runs on the CPU. The Vera CPU's key designs all target this new load: the proprietary Olympus core delivers **2× single-thread performance, 3× inter-core bandwidth, 40% memory-latency reduction**. This is the exact opposite of the traditional data-center CPU \"more cores, lower frequency\" road — Agent orchestration doesn't need more cores, it needs one core to compress the coordination work between successive inferences to the minimum. At the platform level, Vera Rubin NVL72 in CoreWeave's DeepSeek-R1 production test achieves **10× tokens-per-megawatt throughput**, an order-of-magnitude jump over Grace Blackwell NVL72. This improvement comes from a seven-chip co-design: Rubin GPU + Vera CPU + NVLink 6 + ConnectX-9 + BlueField-4 + Spectrum-6 + Groq 3 LPX — not \"stitching a server together\", but treating the whole rack as one accelerator. NVLink 6 provides 260 TB\u002Fs all-to-all bandwidth, specifically addressing the MoE expert-routing bottleneck; Spectrum-X pushes RDMA bandwidth to 1.6× ordinary Ethernet. The engineering details also betray scale ambition: the compute tray is **cable-less, fan-less, hose-less**, with assembly time compressed from hours to **1 minute**; liquid-cooling inlet temperature hits 45°C, with dry coolers eliminating the chiller. This isn't showing off — it's turning \"rack-scale AI factories\" into a product that can actually be mass-deployed. The multi-billion-dollar Microsoft-Mistral deal also lands on thousands of Vera Rubin GPUs, betting on European sovereign AI. So what? When everyone is staring at GPU compute, Agent workloads are pushing the CPU to the C-spot on the AI factory performance curve for the first time. The real signal from Vera CPU isn't \"how much faster than the rival\", but NVIDIA using a self-designed CPU to tell the market: the future AI infrastructure bottleneck isn't \"who can compute fast\", it's \"who can press the coordination overhead of every tool call, context assembly and multi-step reasoning to the minimum\". The OpenAI\u002FAnthropic arms race in Agent frameworks will, in the end, push back down to silicon.","nvidia-vera-cpu-agent-orchestration","2026-07-22T04:50:00Z","2026-07-22T04:53:26.033198Z","2026-08-19T02:08:40.142862Z",true,"agent",76,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"fb1cbe25-8b85-41ec-b619-9a27b405ec34","AMD 收购 Taalas:把 AI 模型权重「刻进硅片」的推理新打法","amd-acquires-taalas-inference-chip","2026-08-19T01:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"a235e2ab-5b61-47b2-a85a-5f8a1d438624","AMD 收购 Taalas:把\"为单一模型造芯\"的路子,搬进 Instinct 体系","amd-acquires-taalas-hardwired-inference-silicon","2026-08-09T06:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"ae924988-53ef-4f46-b1c2-c43c4e9866fe","欧盟掏 100 亿欧元建 7 座 AI 超级工厂：每座堆 10 万颗顶尖芯片，瞄准美国算力代差","eu-ai-gigafactory-7-factories-100000-chips","2026-07-31T04:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f4aad332-1f97-4cb2-96a4-37d8d2980728","英伟达BW首秀RTX Spark：笔记本本地跑120B大模型","nvidia-rtx-spark-120b-laptop","2026-07-12T12:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"66d66fa2-364e-4fa2-a9b2-e69e6f86dc8c","Vera Rubin 平台登陆 ISC 2026：144 张 GPU + 100% 液冷，把 TOP500 算力压进科研机柜","nvidia-vera-rubin-isc2026-144-gpu","2026-06-22T20:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"aadc3e5d-7b81-4996-aed5-12e5a26f6b21","英伟达RTX Spark平台曝光：N2X\u002FN3X路线图初现，端侧推理的芯片博弈","nvidia-rtx-spark-n1x-n2x-n3x-roadmap","2026-06-02T19:00:00+00:00"]