[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-pair-personal-ai-router-local-inference":3,"topics-all":38,"news-related-61cf8d85-c751-4da2-9aae-10b645415ec9":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"61cf8d85-c751-4da2-9aae-10b645415ec9","英伟达发布开源工具 PAIR,把家里电脑连成个人 AI 推理集群","英伟达推出开源工具 Personal AI Router (PAIR) 测试版,可将家中 RTX 显卡电脑、DGX Spark 工作站和苹果 M4 及以上 Mac 组成一个本地 AI 推理集群,跨设备路由 Ollama 或 LM Studio 的推理请求,数据全程留在局域网内。","最近 AI 圈都在卷云端大模型,但英伟达把目光放回了用户的客厅。2026 年 9 月初,英伟达上线了一款名为 Personal AI Router(PAIR)的开源测试版工具,核心卖点只有一句话:把你家里现有的电脑全部串起来,凑成一个本地 AI 推理集群(来源:https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85284)。\n\n## 它到底解决什么问题\n\n跑过本地大模型的人都知道,一张消费级显卡很快就不够用:32B 参数一上来显存就吃紧,多轮 Agent 任务排队时延迟陡升。PAIR 的思路不是把多台机器虚拟成一张「超级 GPU」——英伟达官方 FAQ 里明确说过,PAIR 不会合并设备成单一虚拟 GPU,而是把并行子任务分派给多台空闲机器,让它们各自消化一部分推理请求。这样一台机器跑不动的工作流,在几台机器协同下就能跑得动(来源:https:\u002F\u002Fwww.nvidia.com\u002Fen-us\u002Fai-on-rtx\u002Fpersonal-ai-router\u002F)。\n\n## 兼容什么设备\n\n按官方页面列出的「Validated Configurations」,PAIR 支持的硬件清单相当务实:\n\n- **GPU**:GeForce RTX 20 系列及更新型号,以及 DGX Spark\u002FGB10、Mac M4 或更新机型\n- **系统**:Windows 11、DGX OS、Ubuntu 14.04、macOS Tahoe\n- **内存**:8 GB 及以上\n- **磁盘**:推荐 20 GB 及以上\n\n换句话说,只要你家有一张 RTX 2060 以上的显卡,再加一台 M4 MacBook,理论上就能搭出一个像样的本地推理池。这对 Mac 用户尤其关键——之前想跑大模型基本只能外挂 eGPU 或者上云,现在苹果 M 系列芯片可以直接和英伟达设备协同。\n\n## 推理后端与隐私设计\n\nPAIR 当前内置支持 Ollama 和 LM Studio 两大本地推理框架,接入应用只要走 PAIR 的本地 endpoint,请求就会被智能代理到空闲节点,使用门槛几乎没有。除此之外,英伟达特意强调「Private Local Inference」卖点:prompt、文件、Agent 上下文全程留在家庭局域网,不上传云端。这对担心隐私的开发者、企业内网场景都是个明确加分项(来源:https:\u002F\u002Fwww.nvidia.com\u002Fen-us\u002Fai-on-rtx\u002Fpersonal-ai-router\u002F)。\n\n## 这事意味着什么\n\n把「家用计算资源」打包成推理算力,实际上是云端 LLM 路线的反向叙事。过去两年大家在卷「把模型做大、把数据中心堆满」,PAIR 反过来告诉你:用户手里的 RTX 显卡加上 Mac 芯片,本身就是一片巨大的分布式算力池,只是之前没人调度。\n\n它对开发者的现实意义有两个层面:一是本地 Agent、多模型路由工作流终于有了原生跨平台方案;二是给消费级 GPU 的「剩余价值」找到了去处——你那张闲置的 RTX 3080 终于可以在你睡觉时帮你跑点东西。\n\n当然,PAIR 目前还是 0.1.1 测试版,支持的推理后端只有 Ollama 和 LM Studio 两个,生态还很轻。但方向已经很清楚:大模型不必非得上云,家里那堆吃灰的显卡可能就是你的下一座「数据中心」。\n\n(本文素材综合自 Solidot 报道与英伟达 PAIR 官方页面:https:\u002F\u002Fwww.nvidia.com\u002Fen-us\u002Fai-on-rtx\u002Fpersonal-ai-router\u002F)","https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85284","474eef8c-e0c3-46cf-adee-c089558220f9",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":19,"name":20,"slug":20,"description":14,"color":14},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"3bade52b-3d7a-4104-bea6-410b573fd560","en","NVIDIA's open-source PAIR turns home PCs into a local AI cluster","NVIDIA has released an open-source beta of Personal AI Router (PAIR), which can pool a household's RTX GPUs, DGX Spark workstations, and Apple M4-or-newer Macs into a single local AI inference cluster, routing Ollama or LM Studio requests across devices while keeping all data on the local network.","While the AI world keeps scaling up cloud models, NVIDIA has shifted its gaze back to the living room. In early September 2026, NVIDIA quietly shipped an open-source beta called Personal AI Router (PAIR), whose pitch is simple: string together every PC you already own at home and turn it into a local AI inference cluster (source: https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85284).\n\n## What problem does it actually solve\n\nAnyone who has run large models locally knows the pain: a single consumer GPU runs out of memory fast when a 32B model is loaded, and multi-turn agent workloads pile up with brutal latency. PAIR's approach is NOT to virtualize multiple machines into a single \"super GPU\" — NVIDIA's official FAQ explicitly states that PAIR does not merge devices into one virtual GPU. Instead, it dispatches parallel sub-tasks to whichever machines are idle, letting each one chew through its share of inference requests. A workflow that one machine cannot handle becomes tractable when several cooperate (source: https:\u002F\u002Fwww.nvidia.com\u002Fen-us\u002Fai-on-rtx\u002Fpersonal-ai-router\u002F).\n\n## What hardware is supported\n\nThe \"Validated Configurations\" section on the official page lists a refreshingly pragmatic hardware matrix:\n\n- **GPU**: GeForce RTX 20 Series and newer, DGX Spark\u002FGB10, Mac M4 or newer\n- **Systems**: Windows 11, DGX OS, Ubuntu 14.04, macOS Tahoe\n- **RAM**: 8 GB or higher\n- **Disk**: 20 GB or higher recommended\n\nIn practice, if you have an RTX 2060-or-newer GPU in your home plus an M4 MacBook, you can theoretically assemble a respectable local inference pool. This is particularly meaningful for Mac users — until now, running large models locally usually meant an external GPU or a cloud round-trip, but Apple silicon can now cooperate directly with NVIDIA gear.\n\n## Inference backends and privacy design\n\nPAIR ships with built-in support for two of the most common local inference frameworks, Ollama and LM Studio. Any app pointed at the PAIR local endpoint will have its requests intelligently proxied to whichever node is idle, with almost zero onboarding friction. On top of that, NVIDIA leans heavily on the \"Private Local Inference\" selling point: prompts, files, and agent context stay on the home LAN and never leave for the cloud. That is a clear plus for privacy-conscious developers and enterprise-internal scenarios (source: https:\u002F\u002Fwww.nvidia.com\u002Fen-us\u002Fai-on-rtx\u002Fpersonal-ai-router\u002F).\n\n## What this actually means\n\nPackaging \"household compute\" into inference capacity is effectively the reverse narrative of the cloud-LLM story. For the past two years the industry has been racing to \"build bigger models and stack more data centers.\" PAIR turns that around and says: the RTX GPUs and Apple silicon sitting in users' homes are already a vast distributed compute pool — nobody was just scheduling it.\n\nFor developers, the practical implications cut two ways. First, local agent and multi-model routing workflows finally have a native cross-platform answer. Second, consumer-grade GPUs now have a use case for their \"residual value\" — that idle RTX 3080 in your closet can finally pull its weight while you sleep.\n\nThat said, PAIR is still at the 0.1.1 beta, only Ollama and LM Studio are supported at launch, and the ecosystem is still thin. But the direction is unmistakable: large models do not have to live in the cloud, and that pile of dusty GPUs at home might just be your next \"data center.\"\n\n(This article draws on Solidot coverage and the official NVIDIA PAIR page: https:\u002F\u002Fwww.nvidia.com\u002Fen-us\u002Fai-on-rtx\u002Fpersonal-ai-router\u002F)","nvidia-pair-personal-ai-router-local-inference","2026-09-09T02:00:00Z","2026-09-09T01:02:59.560464Z","2026-09-09T01:02:59.560473Z",true,"agent",111,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"367476a4-b9af-46f1-a6ab-3de1d83640ff","NVIDIA 把中国开发者日搬到苏州:10 月连开两天,AI 推理和物理 AI 是主菜","nvidia-china-developer-day-2026-suzhou","2026-09-16T03:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"107277b2-2c3f-490b-bb68-a1432e723649","英伟达开源 PAIR：把家里闲置显卡串成一座个人 AI 数据中心","nvidia-pair-personal-ai-router","2026-09-05T06:25:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"42b7939c-1b44-43b8-95cf-a8fc2204560d","NVIDIA 开源 Personal AI Router，把家里 RTX 与 Mac 拼成本地 AI 集群","nvidia-personal-ai-router-pair-beta","2026-09-04T03:20:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"bc0edcf9-caba-4d59-9ee8-f1c5b8469c92","英伟达 129 亿美元收购 Hugging Face:AI 算力霸主把开源仓库也吃了","nvidia-hugging-face-12-9b-acquisition","2026-08-27T06:30:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"30a147aa-e3ed-475c-b15f-9e5ffce6ffc9","英伟达 Vera CPU：DeepInfra 实测 Agent 编排提速 2.2 倍","nvidia-vera-cpu-agent-orchestration","2026-07-22T04:50:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"ab856198-bde1-4c2a-9ca2-77ccecf97cbd","ASR流式新标杆：NVIDIA Nemotron 3.5 ASR以600M参数覆盖40语种","nemotron-3-5-asr-nvidia-600m-40-lang","2026-06-11T06:30:00+00:00"]