[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-risc-v-gpu-falcon-gsp":3,"topics-all":38,"news-related-509d88cb-d9df-4dab-8863-3832fc52897a":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"509d88cb-d9df-4dab-8863-3832fc52897a","英伟达每颗 GPU 藏 10-40 个 RISC-V 核心：Falcon 退役、GSP 上位","英伟达每颗 GPU 暗藏 10-40 个 RISC-V 微控制器内核，用 NV-RISCV32\u002F64\u002FRVV 系列替换 2005 年引入的 Falcon；GSP 接管驱动卸载与多租户调度，DLA 深度学习加速器同样基于 RISC-V。十年累计出货约 30 亿颗。","很多人买 GPU 关心 SM 数量、显存带宽、CUDA 核心，却很少有人注意到一件事——每一颗 NVIDIA GPU 里其实都藏着 10 到 40 个 RISC-V 微控制器。它们不参与图形渲染也不参与 AI 算力计算，而是负责视频编解码、电源管理、安全引擎、内核驱动卸载等看不见的工作。\n\n## Falcon 退役，RISC-V 上位\n\nNVIDIA 在 2005 年的 G98 GPU 上首次引入自研微控制器 Falcon（FAst Logic CONtroller）。Falcon 是 32 位核心、无数据缓存，已无法满足 AI 时代对芯片复杂度的要求；NVIDIA 在评估 Arm、MIPS 后，于 2015 年确定把 RISC-V 作为 Falcon 的下一代替代架构。\n\n到 2024 年，NVIDIA 当年 GPU 产品里包含的 RISC-V 核心总量估计约 10 亿颗。从 Turing（2018）到 Ampere（2020）、Ada\u002FHopper（2022）、Blackwell（2024），每一代都在加深对 RISC-V 的依赖。\n\n## 三类内核、一套子系统\n\nNVIDIA 至少有三个 RISC-V 微控制器内核：NV-RISCV32（RV32I-MU，顺序单发射，主频 1.8 GHz）、NV-RISCV64（RV64I-MSU，乱序双发射，主频 2 GHz，支持 SMP）、NV-RVV（在 NV-RISCV32 基础上叠加 1024 位向量扩展）。\n\n围绕这三个核心，NVIDIA 搭起一个叫 Peregrine 的统一子系统，把 RISC-V 内核、缓存、TCM、中断控制器、DMA、RSA\u002FPKA、AES 引擎按需组合，给 GPU 控制平面提供一套可参数化配置的积木。NVIDIA 还加了 20 多项自定义扩展，覆盖 64 位物理地址、2KB 页面、ICD 安全调试、ROM 内存保护等。\n\n## GSP 和 DLA：两块真正干 AI 活的应用\n\n**GPU 系统处理器（GSP）** 位于主机 CPU 和 GPU 之间，通过 PCIe 通信，直接控制 GPU 硬件单元和显存。GSP 用多个 RV64 内核加一致性互连和统一 TCM\u002F缓存，做三件事：内核驱动卸载、减少 GPU 对 CPU 的暴露面、封装 GPU 底层细节。它支持多分区模式——一个 GPU 切成多份 vGPU 分给不同租户的 Guest VM。这种架构直接服务了云端 LLM 推理的多租户场景，也支撑了机密计算。\n\n**深度学习加速器（DLA）** 是 NVIDIA 自研的固定功能推理加速 IP，也跑在 RISC-V 上——NV-RISCV32 跑标量、NV-RVV 跑向量，编译器能把多个算子融合到一个 RVV 内核里。Orin 等边缘 AI 芯片里 DLA 直接跑 TensorRT 模型的离 CPU\u002FGPU 卸载路径，跟大模型部署关系紧密。\n\n## 写在最后\n\nRISC-V Summit 2024 上披露的这些细节，揭示了一件被 AI 算力叙事盖过去的事：现代 GPU 早就是一颗 SoC，里面不仅有 CUDA 核心，还有一套基于开源 ISA 的控制平面和加速子系统。\n\n下次有人说 RISC-V 在数据中心还只是个实验品，把 NV-RISCV32\u002F64\u002FRVV、GSP、DLA、Peregrine 这一串名词丢过去就行。","https:\u002F\u002Fwww.phoronix.com\u002Fnews\u002FRISC-V-NVIDIA-One-Billion","d6b5a2f5-8327-40d2-9bc0-d8a3fd07b47f",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":22,"name":23,"slug":23,"description":14,"color":14},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"f8bfb96f-b1e8-469e-8a18-ac37a7207961","en","Nvidia GPUs Hide 10-40 RISC-V Cores: Falcon Retires","Nvidia GPUs hide 10-40 RISC-V cores per chip, replacing Falcon. GSP and DLA inference IP also run on RISC-V.","Most GPU buyers focus on SM count, memory bandwidth, and CUDA cores. Few realize that every Nvidia GPU actually hides 10 to 40 RISC-V microcontroller cores. They do not render graphics or do AI math; instead they handle video encode\u002Fdecode, power management, security engines, kernel driver offload, and other \"invisible work.\"\n\n## Falcon Retires, RISC-V Rises\n\nNvidia first introduced its proprietary Falcon (FAst Logic CONtroller) microcontroller on the 2005 G98 GPU. Falcon is a 32-bit core without a data cache, and can no longer keep up with the complexity AI workloads demand. After evaluating Arm and MIPS, Nvidia settled on RISC-V as Falcon's successor in 2015.\n\nBy 2024, Nvidia estimated its GPUs that year alone contained around 1 billion RISC-V cores. From Turing (2018) to Ampere (2020), Ada\u002FHopper (2022), and Blackwell (2024), every generation deepens the RISC-V dependency.\n\n## Three Cores, One Subsystem\n\nNvidia has at least three RISC-V microcontroller cores: NV-RISCV32 (RV32I-MU, in-order single-issue, 1.8 GHz, 1.8 CM\u002FMHz), NV-RISCV64 (RV64I-MSU, out-of-order dual-issue, 2 GHz, 5 CM\u002FMHz, supports SMP), and NV-RVV (NV-RISCV32 plus a 1024-bit vector extension).\n\nAround these three, Nvidia built a unified subsystem called Peregrine: it parameterizes the RISC-V core, caches, TCM, interrupt controller, DMA, RSA\u002FPKA, and AES engine into a configurable \"kit\" for the GPU control plane. Nvidia also added 20+ custom extensions covering 64-bit physical addresses, 2KB pages, ICD secure debug, ROM memory protection, and cache operations.\n\n## GSP and DLA: Two Pieces Doing Real AI Work\n\nThe two applications most relevant to AI inference are concrete.\n\nThe GPU System Processor (GSP) sits between the host CPU and the GPU, communicating over PCIe and directly controlling the GPU's hardware units and video memory. GSP uses multiple RV64 cores plus a coherent fabric and unified TCM\u002Fcache. It does three things: offloads kernel drivers, reduces the GPU's exposure surface to the CPU, and encapsulates GPU low-level details. It supports a multi-partition mode — one GPU is sliced into multiple vGPUs distributed across different tenants' Guest VMs. This architecture directly serves the multi-tenant scenario of cloud LLM inference and underpins confidential computing.\n\nThe Deep Learning Accelerator (DLA) is Nvidia's fixed-function inference IP, also running on RISC-V — NV-RISCV32 for scalar work, NV-RVV for vector. The compiler can fuse multiple operators into a single RVV kernel. In Orin and other edge AI chips, DLA runs TensorRT model offload paths away from the CPU\u002FGPU, tightly tied to large-model deployment.\n\n## Final Take\n\nWhat Nvidia disclosed at RISC-V Summit 2024 reveals something the AI compute narrative tends to obscure: the modern GPU is already a SoC. Beyond CUDA cores, it carries a control plane and acceleration subsystem built on an open ISA. The next time someone claims RISC-V is still a \"data center experiment,\" hit them with NV-RISCV32\u002F64\u002FRVV, GSP, DLA, and Peregrine.","nvidia-risc-v-gpu-falcon-gsp","2026-09-21T03:00:00Z","2026-09-21T03:08:34.703064Z","2026-09-21T03:08:34.703080Z",true,"agent",8,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"107277b2-2c3f-490b-bb68-a1432e723649","英伟达开源 PAIR：把家里闲置显卡串成一座个人 AI 数据中心","nvidia-pair-personal-ai-router","2026-09-05T06:25:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"b48106fa-b9c7-4317-a40b-61e4219a698a","IBM Z 双架构处理器:把 Arm 推进大型机,内置 AI 推理","ibm-z-dual-architecture-arm-ai-inference","2026-09-03T03:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"5d452086-ecb7-494d-82b9-d962664aa243","IBM 把 Arm 核塞进 Z 大型机:Hot Chips 2026 公布业界首款双指令集处理器","ibm-z-arm-dual-isa-hot-chips-aug-2026","2026-08-29T06:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"fb1cbe25-8b85-41ec-b619-9a27b405ec34","AMD 收购 Taalas:把 AI 模型权重「刻进硅片」的推理新打法","amd-acquires-taalas-inference-chip","2026-08-19T01:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"45375854-7739-4dd1-bc6a-30db4474652a","Taalas HC2:把单片参数拉到 200 亿,「模型刻进硅片」的第二章","taalas-hc2-20b-mxfp4-50-chips-1t-amd","2026-08-19T00:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"dfdc3216-52aa-4a78-9bf5-859affc37d17","AMD 收下 Taalas：把模型权重刻进芯片，推理的内存墙还剩多少？","amd-acquires-taalas-msic-etched-weights","2026-08-11T02:00:00+00:00"]