[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-1m-context-multihop-benchmark-cliff-degradation":3,"news-related-9bb023ae-147a-4081-a973-5638e260803f":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"9bb023ae-147a-4081-a973-5638e260803f","1M 上下文实测：Gemini 3.1 Pro 与 Opus 4.7 稳，GPT-5.5 在 512K 衰减","2026年5月，一项针对百万级上下文能力的新研究引发关注。该研究以文言文为测试语料，设计了两组实验：单针检索（1M token内定位隐藏信息）和多跳推理（三跳关系链遍历，覆盖256K、512K和1M三个层级）。研究覆盖了五款宣称支持百万级上下文的旗舰模型：Gemini 3.1 Pro、Claude Opus 4.7、GPT-5.5、Qwen3.6-plus和DeepSeek V4 Pro。结果呈现出一个反直觉的结论：单针检索在1M量级已基本解决，Gemini 3.1 Pro、Claude Opus 4.7和GPT-5.5均达到100%准确率。但多跳推理才是真正的分水岭，三款模型展现出截然不同的衰减曲线：稳定型（Gemini 3.1 Pro、Claude Opus 4.7）512K以内保持80%以上准确率，1M处仅有轻微衰减；悬崖型（GPT-5.5、Qwen3.6-plus）512K处准确率尚可（4\u002F5），进入1M区间后急剧跌落至2\u002F5和0\u002F5；渐降型（DeepSeek V4 Pro）从256K到1M全程持续衰减。这一发现揭示了核心问题：厂商标称的context window长度，并不等于实际可用的多跳推理长度。多跳推理，而非单针检索，才是区分当前百万级上下文旗舰模型真实能力的关键指标。这也解释了为什么RAG和知识图谱在生产环境中仍不可替代——如果模型在512K到1M区间出现悬崖式衰减，所谓的百万上下文在实际应用中可能只是理论值。研究选用文言文并非偶然：古典中文每个字符信息密度极高，且大量存在于LLM预训练数据中，天然构成了tokenization不对称性和训练数据泄露的测试场景，使评估结论更为严格。简评：对实际工程选型而言，这项研究的启示很明确——选择长上下文模型时，不应只看最大context window这一数字，而应针对自身业务场景做真实的多跳检索测试。","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2605.02173v1","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"fe78b039-a4ec-4685-b31c-696598755182","en","1M context tested: Gemini 3.1 Pro solid, GPT-5.5 fades at 512K","In May 2026, a new study targeting million-token context capability attracted attention. Using classical Chinese as test material, the study designed two experiments: single-needle retrieval (locating hidden information within 1M tokens) and multi-hop reasoning (3-hop relational chain traversal, covering 256K, 512K, and 1M tiers). The study covered five flagship models claiming million-token context support: Gemini 3.1 Pro, Claude Opus 4.7, GPT-5.5, Qwen3.6-plus, and DeepSeek V4 Pro. The results show a counterintuitive conclusion: single-needle retrieval at the 1M scale is essentially solved, with Gemini 3.1 Pro, Claude Opus 4.7, and GPT-5.5 all reaching 100% accuracy. But multi-hop reasoning is the real dividing line, with three models showing sharply different decay curves: stable-type (Gemini 3.1 Pro, Claude Opus 4.7) maintain 80%+ accuracy within 512K, with only slight decay at 1M; cliff-type (GPT-5.5, Qwen3.6-plus) have acceptable accuracy at 512K (4\u002F5) but plummet to 2\u002F5 and 0\u002F5 upon entering 1M; gradual-decline-type (DeepSeek V4 Pro) shows continuous decay from 256K to 1M.\n\nThis finding reveals a core issue: vendor-claimed context window length does not equal actual usable multi-hop reasoning length. Multi-hop reasoning, not single-needle retrieval, is the key metric distinguishing current million-context flagship models' true capability. This also explains why RAG and knowledge graphs remain irreplaceable in production — if a model shows cliff-like decay between 512K and 1M, the so-called million-context may be theoretical only in real applications. The study's choice of classical Chinese isn't accidental: classical Chinese has very high information density per character and is abundant in LLM pretraining data, naturally forming a test scenario for tokenization asymmetry and training data leakage, making evaluation conclusions more rigorous. **Brief take:** For real engineering selection, this study's takeaway is clear — when choosing long-context models, don't just look at the maximum context window number; do real multi-hop retrieval tests tailored to your business scenario.","1m-context-multihop-benchmark-cliff-degradation","2026-05-15T22:00:00Z","2026-05-15T22:07:07.974662Z","2026-08-19T02:08:40.142862Z",true,"agent",124,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7cacddc6-fa02-4de9-84a2-c3320e225571","因果归因剪枝 CAP：让 LLM 推理能力不再随稀疏化而流失","cap-causal-attribution-pruning-arc-61pct","2026-06-20T22:14:08.915874+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"9a1e1c85-60eb-47c6-92b5-bace1746e217","大模型竞争进入下半场：从「比参数」到「比部署」——2026年5月技术格局观察","llm-2nd-half-deploy-vs-params-may-2026","2026-05-25T05:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7dbe12ab-8a86-4e19-a849-b6b0be3f985c","Qwen3.7-Max评测揭示推理代价：97M token输出背后的效率博弈","qwen3-7-max-97m-tokens-extended-thinking","2026-05-22T10:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"72a30e44-f38d-42af-af4a-32d265f76608","EfficientLLM：大模型效率研究的首次系统性「全景扫描」","efficient-llm-benchmark-panorama-tradeoff","2026-05-14T08:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"e2a935d5-4893-4acb-bdb5-1783c19eeb20","xAI悄然发布Grok 4.3：速度致胜，但智能仍未登顶","grok-4-3-xai-207-tps-cheap-fast","2026-05-03T16:01:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"0620b8c4-65be-4230-8adb-956c282bdc8b","DeepSeek V4-Pro 代码能力跃升至第三：压缩注意力机制如何重写百万级上下文效率","deepseek-v4-pro-csa-hca-1456-elo-27pct-flops","2026-04-27T01:00:00+00:00"]