[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tensorrt-edge-llm-qwen3-8-27b-day0":3,"news-related-a8b9d045-0f4c-4596-baa7-060955365877":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"a8b9d045-0f4c-4596-baa7-060955365877","TensorRT Edge-LLM 0.10.0：Qwen3.8-27B Day-0 上车，边缘 LLM 推理再加速","NVIDIA 边缘推理框架 TensorRT Edge-LLM 发 0.10.0：Day-0 支持阿里 Qwen3.8-27B，新增 Nemotron-3.5 Lightning 等模型，附免 ONNX 导出的实验性引擎构建与多轮 KV-cache 复用。","LLM 推理的主战场正在从数据中心外溢。车载助手、机器人、工业设备都要求把模型塞进本地算力：延迟要可控、断网要能跑、内存预算按 MB 计。NVIDIA 的回答是把 TensorRT Edge-LLM 做成一个开源 C++ 推理运行时，而 2026 年 8 月发布的 0.10.0 版本，把「新模型发布即支持」的节奏正式带到了边缘端。\n\n## 一个为「车上和机器上」定制的推理框架\n\nTensorRT Edge-LLM 是 NVIDIA 面向 Jetson、DRIVE 和 DGX Spark 平台的 C++ 推理运行时，覆盖文本、视觉、音频、语音乃至 action 模型。工作流分三段：先把 Hugging Face 检查点导出为 ONNX，再为目标硬件构建优化的 TensorRT 引擎，最后由 C++ 运行时执行推理。\n\n这个设计与数据中心框架的取向完全不同。NVIDIA 在技术博客里把边缘负载的特征讲得很直白：请求来自单个或少数用户、batch size 低（通常跨摄像头）、部署是任务关键型、且需要离线运行不更新。对应的工程要求是延迟最小且可预测、磁盘内存算力占用最小、符合生产标准、高鲁棒性。为此框架提供了 EAGLE-3 投机解码、NVFP4 量化、chunked prefill 等特性，依赖被刻意压到最少。\n\n## 0.10.0：Day-0 支持 Qwen3.8-27B\n\n这次更新最值得注意的一条，是 0.10.0 对阿里 Qwen3.8-27B 的 Day-0 支持——模型在 Hugging Face 上线，边缘框架同步跟进。同一版本还加入了：\n\n- Nemotron-3.5 Lightning（30B-A3B、NVFP4），支持 MTP 与 DFlash\n- Cosmos3-Edge 与 DiffusionGemma（26B-A4B、NVFP4）\n- Nemotron-3.5-ASR 流式语音识别（0.6B）\n- DSpark 投机解码\n- 实验性的直接引擎构建器：跳过 ONNX 导出，从检查点直接构建 TensorRT 引擎\n- 多轮对话 KV-cache 复用，以及实验性 OpenAI 兼容 server 的视频输入\n\n再往前看 7 月的 0.9.x，完整 Gemma 4 家族（E2B\u002FE4B\u002F12B\u002F26B-A4B\u002F31B，多模态文本+图像+音频，带 MTP）、Qwen3-Omni、Nemotron-3 NVFP4 也都已接入。文档里甚至有完整的 Qwen3-TTS 流水线指南，覆盖 CustomVoice、VoiceDesign 和 Base 检查点——文本、视觉、语音三类模型在边缘端的位面正在被补齐。\n\n## 产业链已经上车\n\n这个框架不是实验室项目。Bosch 联合 Microsoft 与 NVIDIA 基于 TensorRT Edge-LLM 打造了车载 AI 座舱，用端侧 ASR + TTS 配合 LLM 推理，再通过编排器与云端大模型协作；ThunderSoft 把它集成进基于 DRIVE AGX Orin 的 AIBOX 平台；MediaTek 的 CX1 SoC 用它加速座舱 AI 与驾驶员\u002F座舱监控，并向框架回馈了新的嵌入式推理方法。仓库以 Apache 2.0 许可证开源，目前约 512 star、112 fork、24 次 commit——工程节奏还早期，但产业抓手已经就位。\n\n## 所以呢\n\n两件事值得记住。第一，「Day-0 支持 Qwen3.8-27B」这个细节：边缘推理框架对中国开源模型的跟进速度，已经和 Hugging Face 上架节奏对齐——中国开放权重模型不止在下载量榜上领跑，也正在成为海外硬件栈的默认适配对象。第二，免 ONNX 的直接引擎构建器说明 NVIDIA 开始削减部署链路的中间环节：从检查点到边缘引擎的路径越短，车规和产线的确定性就越强。对做机器人、车载、离线设备的团队来说，现在盯这个仓库的 release note，比等季度报告更接近实时战场。\n\n（原文与发布细节见 GitHub 仓库：https:\u002F\u002Fgithub.com\u002FNVIDIA\u002FTensorRT-Edge-LLM）","https:\u002F\u002Fgithub.com\u002FNVIDIA\u002FTensorRT-Edge-LLM","998df6db-96e6-4b8e-8be1-cfa00a6cd177",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"04db9061-b59d-4eef-b29b-144291d7c724","en","TensorRT Edge-LLM 0.10.0: Qwen3.8-27B Gets Day-0 Support as Edge LLM Inference Accelerates","NVIDIA's edge inference framework TensorRT Edge-LLM ships 0.10.0 with Day-0 support for Alibaba's Qwen3.8-27B, new models including Nemotron-3.5 Lightning, an experimental ONNX-free engine builder, and multi-turn KV-cache reuse.","LLM inference is spilling out of the data center. In-car assistants, robots, and industrial equipment all demand models running on local silicon: latency must be controllable, offline operation must work, and memory budgets are counted in megabytes. NVIDIA's answer is TensorRT Edge-LLM, an open-source C++ inference runtime — and the 0.10.0 release from August 2026 brings \"support a new model the day it ships\" cadence to the edge.\n\n## A framework custom-built for cars and robots\n\nTensorRT Edge-LLM is NVIDIA's C++ inference runtime for the Jetson, DRIVE, and DGX Spark platforms, covering text, vision, audio, speech, and even action models. The workflow has three stages: export Hugging Face checkpoints to ONNX, build optimized TensorRT engines for the target hardware, then run inference with the C++ runtime.\n\nThis design points in the opposite direction from data-center frameworks. NVIDIA's technical blog states the characteristics of edge workloads plainly: requests come from a single user or a few users, batch sizes are low (typically across cameras), deployments are mission-critical, and systems must operate offline without updates. The corresponding engineering requirements are minimal and predictable latency, minimal disk\u002Fmemory\u002Fcompute footprint, compliance with production standards, and high robustness. To deliver this, the framework provides EAGLE-3 speculative decoding, NVFP4 quantization, and chunked prefill, with dependencies deliberately kept to a minimum.\n\n## 0.10.0: Day-0 support for Qwen3.8-27B\n\nThe most notable item in this update is Day-0 support for Alibaba's Qwen3.8-27B — the model lands on Hugging Face, and the edge framework follows in lockstep. The same release also adds:\n\n- Nemotron-3.5 Lightning (30B-A3B, NVFP4) with MTP and DFlash\n- Cosmos3-Edge and DiffusionGemma (26B-A4B, NVFP4)\n- Nemotron-3.5-ASR streaming speech recognition (0.6B)\n- DSpark speculative decoding\n- An experimental direct engine builder that skips ONNX export and builds TensorRT engines straight from checkpoints\n- Multi-turn KV-cache reuse, plus video input for the experimental OpenAI-compatible server\n\nLooking back at the 0.9.x releases from July, the full Gemma 4 family (E2B\u002FE4B\u002F12B\u002F26B-A4B\u002F31B — multimodal text + image + audio, with MTP), Qwen3-Omni, and Nemotron-3 NVFP4 are all on board. The docs even include a full Qwen3-TTS pipeline guide covering CustomVoice, VoiceDesign, and Base checkpoints — the text, vision, and speech columns of the edge-model matrix are being filled in.\n\n## The industry is already on board\n\nThis is not a lab project. Bosch, working with Microsoft and NVIDIA, built its AI-powered cockpit on TensorRT Edge-LLM, pairing on-device ASR + TTS with LLM inference, coordinated with larger cloud models through an orchestrator. ThunderSoft integrated it into its AIBOX platform based on NVIDIA DRIVE AGX Orin. MediaTek's CX1 SoC uses it to accelerate cabin AI and driver\u002Fcabin monitoring, and contributes new embedded-specific inference methods back to the framework. The repository is open-sourced under Apache 2.0, currently at roughly 512 stars, 112 forks, and 24 commits — early in its engineering cadence, but the industrial footholds are already in place.\n\n## So what\n\nTwo things are worth remembering. First, the \"Day-0 support for Qwen3.8-27B\" detail: edge inference frameworks now track Chinese open-weight models at the same pace as their Hugging Face releases — Chinese open-weight models are not just topping download charts, they are becoming default adaptation targets for overseas hardware stacks. Second, the ONNX-free direct engine builder signals NVIDIA trimming intermediate steps from the deployment chain: the shorter the path from checkpoint to edge engine, the stronger the determinism for automotive-grade and production-line requirements. For teams building robots, vehicles, or offline devices, watching this repo's release notes is now closer to a real-time front line than any quarterly report.\n\n(Original release details: https:\u002F\u002Fgithub.com\u002FNVIDIA\u002FTensorRT-Edge-LLM)","tensorrt-edge-llm-qwen3-8-27b-day0","2026-08-21T15:00:00Z","2026-08-20T21:08:34.080740Z","2026-08-20T21:08:34.080755Z",true,"agent",109,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"63255594-b16a-4106-9acc-2dc479b97e14","英伟达Nemotron 3 Ultra登场：550B开源模型刷新美国开放权重智能榜单","nvidia-nemotron-3-ultra-550b-48-intelligence","2026-06-01T10:10:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"2b5b7c66-7289-45db-b5b7-dea67882310c","NVIDIA 发布 Nemotron-Labs Diffusion：三模态语言模型统一 AR 与扩散解码","nvidia-nemotron-diffusion-ar-dllm-tri-modal","2026-05-23T04:10:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"c94bdf86-5de9-49fe-8c98-0f5c47611bfe","SGLang v0.5.18 发布:大模型冷启动提速 2.38 倍,710 个 PR 都改了什么","sglang-v0-5-18-cold-start-2-38x","2026-08-24T23:15:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"1311adb6-dc19-41a7-a188-6760d9e53672","HF Summer 2026 报告:13 个下载量 Top 25 模型是 2022 年的老面孔","hugging-face-summer-2026-attention-adoption","2026-08-24T08:00:00+00:00"]