[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-dense-contexts-12k-needle-density-mattr":3,"topics-all":36,"news-related-7e7d7d95-a592-4ee7-a43f-c3d109c58405":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"7e7d7d95-a592-4ee7-a43f-c3d109c58405","LLM 长上下文「有效容量」被高估了：12K 词也能撑爆，密度才是隐藏分水岭","最近 arXiv 上一篇来自意大利都灵理工大学的论文《Dense Contexts Are Hard Contexts》给「百万上下文」叙事泼了一盆冷水。研究者用三组长度完全相同（约 12K tokens）的\"找针\"基准、严格控制信息位置，只改变信息密度——结果发现一个被行业长期忽视的现象：即便长度不变，模型的检索准确率会随密度上升断崖式下跌，原本在稀疏文本上几乎拿满分的开源模型，落到高密度场景直接掉到 60% 以下。\n\n这颠覆了「上下文窗口 = 有效容量」的隐含假设。过去两年，业界把\"长上下文\"等同于\"长注意力\"：把窗口从 128K 推到 1M、10M，benchmark 数字就好看，营销话术就响亮。但新论文指出，决定 LLM 表现的第三根轴是词项多样性（MATTR）——同样是 12K token，小说式散文（MATTR≈0.72）可以跳读，而配置型\u002F代码型\u002F检索拼装型文本（MATTR≈0.82）几乎每个 token 都要处理。长度相同，密度变了，难度天差地别。\n\n对所有在拼\"百万 token\"的厂商和工程团队来说，这是一次清醒提醒：RAG 检索后塞进 prompt 的资料、Agent 拼接的多段 tool 输出、用户长会话历史——只要本身就是高密度内容，1M 窗口的实际可用度可能还不如一个 64K 的\"清爽\"上下文。Llama 4、Gemini 3.x、Qwen3.7-Max 等新一代旗舰在 NIAH 上狂刷高分，并不代表它们在真实 Agent 工作流里就稳了。\n\n更值得关注的是下一步：能否训练模型对密度自适应，或者在推理端对高密度片段做动态摘要\u002F分块——这才是把\"长上下文\"从 benchmark 拉回生产价值的关键。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.06203","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"045c011e-e2bb-45ce-bdd6-0c927f8a3b87","token-efficiency",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"96b2a337-e142-45ba-9021-57a7e80d916d","en","Long-context capacity is overrated; density is the real line","A recent arXiv paper from Politecnico di Torino, \"Dense Contexts Are Hard Contexts,\" pours cold water on the \"million-token context\" narrative. Researchers used three sets of \"needle in a haystack\" benchmarks with exactly the same length (~12K tokens), strictly controlled the position of the information, and only changed the information density — and found a long-overlooked phenomenon: even if the length doesn't change, the model's retrieval accuracy drops off a cliff as density rises. Open-source models that nearly maxed out on sparse text fell directly below 60% on the high-density scenario.\n\nThis overturns the implicit assumption that \"context window = effective capacity.\" Over the past two years, the industry has equated \"long context\" with \"long attention\": pushing the window from 128K to 1M to 10M makes the benchmark numbers look good, the marketing copy sounds loud. But the new paper points out that the third axis that determines LLM performance is lexical item diversity (MATTR) — the same 12K tokens, novel-style prose (MATTR≈0.72) can be skip-read, while configuration-style \u002F code-style \u002F retrieval-assembled text (MATTR≈0.82) requires almost every token to be processed. Same length, different density, the difficulty is worlds apart.\n\nFor all the vendors and engineering teams piling on \"million tokens,\" this is a sobering reminder: the actual usability of a 1M window may be even less than a \"clean\" 64K context for RAG-retrieved material stuffed into the prompt, multi-segment tool output stitched together by Agents, and long user-session history — as long as the content itself is high-density. Llama 4, Gemini 3.x, Qwen3.7-Max and other new flagships pile on high scores on NIAH, but that doesn't mean they will be stable in real Agent workflows.\n\nWhat's more noteworthy is the next step: whether it is possible to train the model to be adaptive to density, or to do dynamic summarization\u002Fchunking on high-density segments at the inference end — this is the key to pulling \"long context\" from the benchmark back to production value.","dense-contexts-12k-needle-density-mattr","2026-06-08T08:00:00Z","2026-06-08T08:13:31.274948Z","2026-08-19T02:08:40.142862Z",true,"agent",151,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"13ed0c94-9bea-440c-93c4-771c38ba9834","sink 消失后，KV 驱逐改看 value 几何","valuediff-value-geometric-kv-eviction","2026-09-22T17:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"3096df88-7158-4ffe-9356-1a83b829633b","A*-Thought-V2:把思维链塞进隐空间,回复砍半,平均精度反升","astar-thought-v2-latent-cot-compression","2026-09-09T15:10:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"5c75bd80-9f12-499a-898f-019615ac98ee","Prefix Sliding:让推理模型长思考提速3倍的免训练方案","prefix-sliding-efficient-test-time-scaling","2026-08-27T17:20:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"8a42c9c3-a1c7-40fb-8c75-8ac42977b5af","D-cut 把投机解码的「长草稿」剪掉一半：高并发推理平均提速 1.65×、MoE 跑出 3×","d-cut-speculative-draft-cut","2026-07-18T10:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"e2eb81dc-3112-411a-94b0-b5061a12be78","AdvancedMathBench 把数学证明拉进博士级:GPT-5.5-xhigh 仍只 75.8","advanced-math-bench-phd-level","2026-07-14T16:15:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"fc1e888d-4bee-4633-86f4-edc76aa48161","BlockSearch 把语言模型变成「上下文检索器」：0.6B 在百万 token 上打平向量检索","blocksearch-context-retriever","2026-07-03T06:25:00+00:00"]