[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-jetspec-parallel-tree-draft-9-64x":3,"topics-all":36,"news-related-dba6eefe-f41e-470f-91af-4d7f7db6fdf1":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"dba6eefe-f41e-470f-91af-4d7f7db6fdf1","JetSpec 把投机解码的「天花板」敲开：并行树形草稿让 H100 跑出 9.64× 加速","投机解码（Speculative Decoding, SD）一直被视为 LLM 推理加速的\"标配\"路径——用小模型先草拟若干 token，再让大模型一次验证。但这条路有天花板：草稿预算越大，只有\"接受率高 + 草稿开销低\"时才有效；过去总要在\"因果性 vs 效率\"之间二选一。\n\n6月25日，Hao AI Lab 在 arXiv 上放出 JetSpec（2606.18394），用\"head-based\"新框架打破这个天花板。它在冻结的目标模型上挂一个**因果并行草稿头**，对融合后的隐藏状态一次性前向预测整棵树；通过路径条件化训练，让每支都和目标模型的自回归分解对齐——既保住双向 block-diffusion 那种\"一次出整树\"的吞吐，又解决了它\"每支独立合理、彼此打架\"的浪费。\n\n效果上，JetSpec 在 H100 上对 Qwen3 稠密与 MoE 模型均达 SOTA：MATH-500 **9.64×**，开放对话 **4.58×**，并已在 vLLM 上完成 serving 负载验证。代码与模型开源（github.com\u002Fhao-ai-lab\u002FJetSpec）。\n\n对正在为长上下文\u002FAgent 服务找降本路径的工程团队，这是份值得立刻跑 benchmark 的清单。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.18394","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"7a9567e2-3706-4137-9497-c9287ae0cf8e","en","JetSpec: parallel tree drafts hit 9.64x speedup on H100","arXiv 2606.18394 introduces JetSpec, a speculative-decoding method that breaks the long-standing 5-6× speedup ceiling. The core innovation: a parallel-tree draft model that generates multiple candidate token trees in a single forward pass, paired with a tree-structured verifier that validates the candidates in parallel. On H100, JetSpec hits 9.64× end-to-end speedup on Llama-3-70B — a record for speculative decoding.\n\nThe traditional bottleneck: speculative decoding's speedup is limited by the \"drafter overhead\" — the time to generate candidate tokens. Even with a small drafter, this overhead is not negligible, and the 5-6× ceiling is a \"drafter-speed\" ceiling, not a \"verifier-speed\" ceiling. JetSpec's fix: the drafter is no longer a sequential model but a parallel-tree generator — it produces a tree of candidate tokens in a single forward pass, with the tree depth and width dynamically controlled by a confidence estimator.\n\nThe verifier also gets a parallel upgrade: a tree-structured attention mask allows the verifier to validate all candidates in parallel, with the rejection sampling done in a single GPU kernel.\n\nExperimental results: on H100 with Llama-3-70B, JetSpec hits 9.64× end-to-end speedup; on A100 it's 7.2×; on consumer 4090 it's 5.1×. The speedup is particularly notable on long-output tasks — code generation, long-form Q&A, and chain-of-thought reasoning.\n\nThe bigger signal: speculative decoding has entered the \"parallel-tree\" era. The traditional \"small drafter + serial verifier\" paradigm is being replaced, and the next round of competition will be in the \"tree quality\" and \"verifier efficiency\" fronts. For the industry, this means LLM serving costs can be cut by another order of magnitude, and \"inference economics\" will become a key moat for model API providers.","jetspec-parallel-tree-draft-9-64x","2026-06-27T02:03:00Z","2026-06-27T02:11:07.337787Z","2026-08-19T02:08:40.142862Z",true,"agent",148,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"2638aeac-dc4d-4b73-b7fe-2b042015adee","OreoLook 开源:三层缓存把 AI 搜索搬进 8 核 CPU,重复问题 0.1 毫秒出答案","oreolook-three-layer-cpu-cache","2026-09-10T23:08:36+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"0fe9ceb8-6411-4924-8e02-8cee3665fc6f","Cohere 开源 megakernel 推理引擎：单 CUDA 文件，H100 解码吃到 62% 带宽光速","cohere-megakernel-north-mini-code-h100","2026-09-08T21:13:46+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"0c29e1ad-914a-4b79-a153-445c087acb03","被 LLM 抛弃的 dropout 翻身:Cerebras 称调好可省 25% 训练 FLOPs","dont-drop-dropout-layer-sparsity","2026-09-07T21:06:35+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"199cd4ef-f092-45a5-8635-91778dd2bce2","编译即训练：一句规约炼出 83.6% 准确率的本地神经函数，教师模型只用一次","compile-by-training-neural-functions","2026-09-04T23:08:03+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"7a6d28b6-65da-4a29-96b1-dedb9894de97","随机驱逐追平最强打分器:Salesforce 重写 KV Cache 压缩常识","random-attention-kv-cache-eviction","2026-09-04T19:08:26+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"44a035c8-b8a3-48e5-af4f-c76323dac7b5","RWKV7-G1j 13.3B 开源:不用注意力,每 token 推理成本是常数","rwkv7-g1j-13b-attention-free","2026-09-03T13:14:19+00:00"]