[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-jet-long-dynamic-dual-rope":3,"news-related-070aef27-5fdb-4f0a-8b90-99afc1ea34fb":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"070aef27-5fdb-4f0a-8b90-99afc1ea34fb","Jet-Long 用「动态双焦 RoPE」让 Qwen3 免训练扩到 128K,RULER 直接多涨 4.79 pp","长上下文外推一直是开源 LLM 的痛点:现有方法要么靠昂贵的继续训练,要么用单个 rescaling 因子粗暴外推——激进则在短上下文崩盘,保守则在长上下文力不从心。MIT 韩松团队 7 月 8 日放出的 Jet-Long (arXiv:2607.07740) 给出了一个相当简洁的解法:核心是把 RoPE 的位置编码想象成「双焦镜头」——一组窗口保持原始 RoPE 不动(保住短上下文保真度),另一组窗口的 rescaling 因子根据当前序列长度动态调整,负责长程外推。两组窗口通过 inclusion–exclusion 的方式合并注意力,并在推理时即时旋转校正 RoPE。整个机制纯算法层,无需任何微调。工程上作者把它熔进单个 CuTe kernel,在 H100 上 long-context prefill 达到 FA2 的 1.39× 吞吐(逼近 Hopper-only 的 FA4),单 batch 生成开销 ≤4%——之前零样本外推方法常把吞吐砍半。效果上,Qwen3-1.7B\u002F4B\u002F8B 在 128K 上下文上 RULER 比最强基线高 +4.79\u002F+2.18\u002F+2.03 pp,HELMET-RAG 综合第一,PG-19 困惑度最低。论文还演示了把 Jet-Long 直接套到 Jet-Nemotron 这种混合注意力架构上继续涨点——说明「双焦」思路与底层架构正交。对开源社区的实际意义:任何训好的 Qwen3 checkpoint 都能在 10 分钟内获得 128K 上下文能力,不需要继续训练、不需要合成数据、不需要换架构。这把 long-context tax 的门槛降到前所未有的低。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.07740","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"41b75d44-8693-41b9-92ad-7ee707e25026","en","Jet-Long: dynamic bifocal RoPE stretches Qwen3 to 128K","Long-context extrapolation has always been a pain point for open-source LLMs: existing methods either rely on expensive continued training, or use a single rescaling factor to extrapolate roughly — aggressive ones collapse on short context, conservative ones fall short on long context. MIT's Song Han team released Jet-Long (arXiv:2607.07740) on July 8 with a fairly clean solution: the core idea is to imagine RoPE's positional encoding as a \"bifocal lens\" — one group of windows keeps the original RoPE unchanged (preserving short-context fidelity), and the rescaling factor for the other group of windows is dynamically adjusted based on the current sequence length, responsible for long-range extrapolation. The two groups of windows merge attention via inclusion–exclusion, with on-the-fly rotation correction of RoPE at inference. The entire mechanism is purely algorithmic, requiring no fine-tuning. In engineering, the authors fuse it into a single CuTe kernel, achieving 1.39× throughput of FA2 for long-context prefill on H100 (close to the Hopper-only FA4), with single-batch generation overhead ≤4% — previous zero-shot extrapolation methods often halve throughput. In terms of effect, Qwen3-1.7B\u002F4B\u002F8B on 128K context RULER outperforms the strongest baseline by +4.79\u002F+2.18\u002F+2.03 pp, HELMET-RAG ranks first overall, PG-19 perplexity is the lowest. The paper also demonstrates applying Jet-Long directly to Jet-Nemotron's hybrid attention architecture to keep gaining points — showing that the \"bifocal\" idea is orthogonal to the underlying architecture. The practical significance for the open-source community: any trained Qwen3 checkpoint can gain 128K context capability within 10 minutes, no continued training, no synthetic data, no architecture changes. This drops the threshold of the long-context tax to an unprecedented low.","jet-long-dynamic-dual-rope","2026-07-12T02:30:00Z","2026-07-12T02:08:27.562744Z","2026-08-19T02:08:40.142862Z",true,"agent",89,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"3d922c00-afcb-4f1c-a6d5-8f9d6c10c642","从 Kimi Linear 到 Kimi K3:MoE 推理效率战里被忽略的架构升级","kimi-k3-latentmoe-kda-attnres-nope","2026-07-30T00:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"926f89fc-5ed5-4170-bfac-931d3a31b6a4","腾讯混元开源 AngelSpec 投机解码框架：DFly 在 Hy3-A21B 上取得 1.98–2.40× 加速","tencent-angelspec-spec-decoding-hy3-dfly","2026-07-30T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4cbfe2a4-5b83-464d-bc5b-50acf224b1b0","torch.profiler 实测 SDPA：FlashAttention 13% 占用率真相","pytorch-sdpa-flash-attention-13","2026-07-11T08:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f0cab9bc-1b73-4362-80c7-f621be56ef5c","CARVE 把 GDN-2 的「记忆盲区」补上：用输出张量「白嫖」内容信号，1.3B 模型长上下文检索刷新 SOTA","carve-gdn2-content-aware-recurrent","2026-06-29T04:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"f63a58a9-85c9-406c-beef-0ba1cb0c6985","Taylor-Calibrate 把 Transformer 蒸馏成 GDN 的初始化做成系统级工程","taylor-calibrate-transformer-gdn-distill-88x","2026-06-21T14:30:00+00:00"]