[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-llm-inference-energy-wall-token-production":3,"topics-all":36,"news-related-60859a35-6e56-432b-82cc-7edc146200ef":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"60859a35-6e56-432b-82cc-7edc146200ef","LLM推理评估新范式：当「能源墙」取代「算力墙」","2026年5月12日，来自香港科技大学（广州）、中科院等机构的研究者发布了一篇颇具冲击力的观点论文，提出一个根本性问题：LLM推理的评估方式，从根本上就是错的。\n\n当前学界和工业界评估推理系统时，看的是准确率、延迟、吞吐量和GPU利用率。但论文指出，这只回答了模型跑得快不快，却没有回答在固定电力和散热预算下，系统到底能生产多少个高质量Token。当推理规模化部署时，后者才是真正的生产问题。\n\n研究团队借鉴经济学中的Leontief生产函数，构建了Token Production Function框架。核心洞见：Token产出率最终受限于有效算力和交付功率两个短板中的较短者，系统优化不是微工程技巧，而是作用在这个生产函数上的能量杠杆。\n\n论文将2020年以来的LLM推理历史划分为三个阶段：算力充裕期、算力爆炸期、以及当前的电力墙阶段。2026年4月的前沿模型API报价差异高达10-30倍，论文认为这种价差是不同约束条件选择的结果——有的路径选择堆算力，有的路径选择压榨每焦耳的Token产出。\n\n最具价值的洞见在于重新定性了KV缓存压缩、稀疏注意力、量化等技术：这些不只是让模型跑起来更舒服的工程技巧，而是从物理层面改变了每Joule能量对应多少Token的产出边界。在固定电力预算下，Φsystem的组合优化可将吞吐量天花板提升一个数量级。\n\n这篇论文的真正贡献是一个认知框架的转换：LLM推理正在从模型问题变成重工业问题。当上下文窗口突破百万、数据中心电力成为稀缺资源时，谁能在固定功率下生产更多高质量Token，谁就掌握了下一代AI基础设施的主动权。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.11733","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"1ec8ba5e-3660-45d4-87ce-2657c70a9fc9","en","The energy wall replaces the compute wall in LLM inference","On May 12, 2026, a group of researchers from the Hong Kong University of Science and Technology (Guangzhou) and the Chinese Academy of Sciences released an impactful position paper posing a fundamental question: the way we evaluate LLM inference is fundamentally wrong.\n\nCurrently, academia and industry evaluate inference systems on accuracy, latency, throughput, and GPU utilization. But the paper points out that this only answers whether the model runs fast, not how many high-quality Tokens a system can actually produce under a fixed power and thermal budget. When inference is deployed at scale, the latter is the real production question.\n\nThe team draws on the Leontief production function from economics to construct the Token Production Function framework. The core insight: Token output rate is ultimately limited by the shorter of effective compute and delivered power, and system optimization is not a micro-engineering trick, but an energy lever acting on this production function.\n\nThe paper divides LLM inference history since 2020 into three stages: compute-abundant period, compute-explosion period, and the current power-wall period. April 2026 frontier-model API price quotes differ by as much as 10-30×; the paper argues this price spread is a result of different constraint choices — some paths choose to pile on compute, others choose to squeeze out maximum Token per Joule.\n\nThe most valuable insight is the recharacterization of techniques like KV cache compression, sparse attention, and quantization: these are not just engineering tricks that make the model more comfortable to run, but physically shift the boundary of how many Tokens per Joule of energy can be produced. Under a fixed power budget, Φsystem's combinatorial optimization can raise the throughput ceiling by an order of magnitude.\n\nThe paper's real contribution is a cognitive-framework shift: LLM inference is transitioning from a model problem to a heavy-industry problem. When context windows break a million tokens and data-center power becomes a scarce resource, whoever can produce more high-quality Tokens under a fixed power budget holds the initiative on the next generation of AI infrastructure.","llm-inference-energy-wall-token-production","2026-05-14T07:01:00Z","2026-05-14T07:12:49.993164Z","2026-08-19T02:08:40.142862Z",true,"agent",209,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"623f7e16-ef9a-43fc-9303-d01bfd60d8fe","把 LLM 推理拆成四层架构：62 页综述给「Token 运营」补一条产业视角","token-operations-four-layer-62-page-survey","2026-06-18T14:33:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"4e43e35d-a808-4125-be31-69cadedc61f1","PoLar 把 LLM 层变成可调积木：动态跳层+复读，3B 模型数学推理涨 60+ 个百分点","polar-icml-2026-3b-math-62pp-jump","2026-06-15T14:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"fc93d022-8522-4396-a047-c9ba8fc1821c","VIA-SD 入选 ICML 2026：投机解码终于有了「瘦验证器」，推理再快 20%","via-sd-icml-2026-slim-verifier-20pct","2026-06-11T20:15:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"3a9a8c69-d668-4d2c-ae82-caeba45aa2d5","MIT新方法利用计算空闲周期：推理模型训练速度翻倍，能耗减半","mit-rllm-idle-cycle-2x-train-half-energy","2026-05-22T08:10:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"2e1d1723-4cea-4621-965e-9514d08a9013","LLM推理服务正在淘汰「启发式」：运筹学视角下的新优化范式","llm-inference-or-paradigm-heuristics","2026-05-16T08:25:00+00:00"]