[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-dels-spec-dflash-dual-expert":3,"topics-all":33,"news-related-7518634b-1961-4484-8047-9f787dcb2c18":52},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":20,"news_slug":26,"published_at":27,"created_at":28,"modified_at":29,"is_published":30,"publish_type":31,"image_url":13,"view_count":32},"7518634b-1961-4484-8047-9f787dcb2c18","DeLS-Spec 把 DFlash 拆成「长-短双专家」:加一个本地头就能再提速","DFlash 把整块一次性起草,把投机解码效率抬了一截,但块内每个位置缺显式因果;Domino、DSpark 想补短板,代价却是从头重训草稿模型。\n\n7 月 8 日挂在 arXiv 的 DeLS-Spec(arXiv:2607.07409) 给出更轻的路子:把现成 DFlash 冻住当「长上下文专家」,再额外训练一个轻量本地头作为「短上下文专家」。本地头只用标准 next-token 目标独立训练,既不要联合训练目标模型,也不绑定特定 DFlash 版本,训练成本几乎可忽略。推理时两路 logits 融合:长程依赖由 DFlash 兜底,块内一致性由本地头补齐。在 Qwen3 上的实验里,DeLS-Spec 在数学、代码、对话三类基准上一致超过 DFlash 原版,平均接受长度同步抬升。\n\n比起 Domino\u002FDSpark「重训一切」的思路,DeLS-Spec 更像插件式改造——对已经部署 DFlash 的团队几乎零迁移成本。它的价值不在炫技式的端到端重训,而在于给出高度模块化的设计模板:长程能力交给已成熟的专家,短程一致性用极轻的本地头补齐。这条「冻结主模型 + 插件子头」的范式,或许会被接下来一连串推理加速工作借用。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.07409","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[21],{"id":22,"lang":23,"title":24,"summary":25,"content":13},"9884725c-e04c-4fba-9862-d0cdf3edf2f6","en","DeLS-Spec splits speculation into long-short dual experts","DFlash drafts the whole block in one go, lifting speculative decoding efficiency a notch, but lacks explicit causality at each position within the block; Domino and DSpark try to patch the shortcoming, but the cost is retraining the draft model from scratch. DeLS-Spec (arXiv:2607.07409), posted to arXiv on July 8, gives a lighter path: freeze the existing DFlash as a \"long-context expert\", then additionally train a lightweight local head as a \"short-context expert\". The local head is trained independently with only the standard next-token objective, neither requiring joint training of the target model nor binding to a specific DFlash version, with training cost almost negligible. At inference, the two paths of logits are merged: long-range dependencies are backed by DFlash, intra-block consistency is patched by the local head. In experiments on Qwen3, DeLS-Spec consistently beats the original DFlash on math, code, and conversation benchmarks, with average accepted length also climbing. Compared with Domino\u002FDSpark's \"retrain everything\" approach, DeLS-Spec is more of a plug-in modification — for teams that have already deployed DFlash, migration cost is almost zero. Its value isn't an end-to-end retraining flex, but a highly modular design template: long-range capability goes to a mature expert, short-range consistency is patched by an extremely lightweight local head. This \"freeze main model + plug-in sub-head\" paradigm may be borrowed by a string of inference-acceleration works to come.","dels-spec-dflash-dual-expert","2026-07-10T00:08:00Z","2026-07-10T00:11:07.410278Z","2026-08-19T02:08:40.142862Z",true,"agent",185,[34,43],{"slug":35,"tag_slug":35,"title_zh":36,"title_en":37,"intro_zh":38,"intro_en":39,"id":40,"is_active":30,"created_at":41,"modified_at":42},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":44,"tag_slug":44,"title_zh":45,"title_en":46,"intro_zh":47,"intro_en":48,"id":49,"is_active":30,"created_at":50,"modified_at":51},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":53},[54,59,64,69,74,79],{"id":55,"title":56,"news_slug":57,"published_at":58},"c3814f7d-2649-4660-a798-28fb03aa2b6d","SwitchSD 让投机解码学会「该抄才抄」:读内部信号,EAGLE3 之上再快 15%","switchsd-copy-intent-speculative-decoding","2026-09-20T23:09:25+00:00",{"id":60,"title":61,"news_slug":62,"published_at":63},"3cecce90-70b9-4bb3-b9b7-93e6b0c05105","D-Quant 用熵编码压 KV:2.26bit 近无损","d-quant-entropy-coding-kv-cache","2026-09-20T17:10:42+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"813ad679-51dd-43d7-afcc-0baf48d2ef5f","When2Think:推理模型该想多久,先看题有多难","when2think-difficulty-aware-length-control","2026-09-19T19:08:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"9a3cd449-e29a-4730-814b-f1be5c2685c6","复旦FFD让Flash Attention退役？11.6× kernel提速把长上下文推到256K","fudan-ffd-long-context-attention-sparsity","2026-09-15T07:15:46+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"2638aeac-dc4d-4b73-b7fe-2b042015adee","OreoLook 开源:三层缓存把 AI 搜索搬进 8 核 CPU,重复问题 0.1 毫秒出答案","oreolook-three-layer-cpu-cache","2026-09-10T23:08:36+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"178aa5e5-2a4f-4a87-a97c-0da16295d96f","EMNLP 2026 OCGQuant:用通道配对治 NVFP4 陪葬误差,Qwen3-1.7B 接近 FP16","ocgquant-nvfp4-outlier-companion-grouping","2026-09-10T09:15:00+00:00"]