[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-token-operations-four-layer-62-page-survey":3,"topics-all":36,"news-related-623f7e16-ef9a-43fc-9303-d01bfd60d8fe":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"623f7e16-ef9a-43fc-9303-d01bfd60d8fe","把 LLM 推理拆成四层架构：62 页综述给「Token 运营」补一条产业视角","当 LLM 服务真正进入大规模商用，「省 token、稳 token」已经比「刷 benchmark」更值钱。arXiv:2606.20295 在 6 月 18 日放出一份 62 页综述，首次把「面向 Token 运营（Token-Operations-Oriented）」的推理优化整理成一套四层架构：多模型融合、模型优化、计算-模型融合、计算-网络-模型融合，依次对应「在多模型间切流量」「单模型内的量化\u002F蒸馏\u002F投机解码」「算力调度与模型协同」「网络栈参与推理」四件事。\n\n它真正值得关注的不是任何单点 trick，而是视角切换：把 token 当成产线上的零件，关注它的「生产、供应、稳定性」，而不是只盯着模型本身打榜。论文把业内散落在 PD 分离、KV cache 复用、Speculative Decoding、MoE 路由、RDMA\u002FInfiniband 协同推理等不同圈层的优化技术，重新归位到一条纵向价值链上，并给出 36 张图系统对照「业内到底做到了哪一层」。\n\n对工业团队来说，这套框架最大的用处是诊断「成本到底卡在哪一层」——若 80% 预算花在「算力-网络协同」，继续做 INT8 量化收益有限，反而应补齐 prefill\u002Fdecode 分离部署和 RDMA 拓扑；对个人开发者来说，则提示了一个常被忽略的趋势：未来 LLM 工程师很可能要兼懂网络栈，「算法优化」和「系统工程」之间的边界在迅速收窄。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.20295","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"944ab20a-3353-48c5-a307-d166a724a4d3","en","Four-layer LLM inference: a 62-page industry survey","arXiv 2606.20295 introduces a 62-page survey that decomposes LLM inference into a \"four-layer architecture,\" providing an industry-oriented view of \"Token Operations\" (the discipline of serving LLM inference at scale). The survey is co-authored by leading inference infrastructure researchers from OpenAI, NVIDIA, Anthropic, and academic institutions.\n\nThe four layers: (1) the **Model Layer** — the neural network itself (Transformer, SSM, etc.); (2) the **Scheduling Layer** — request batching, routing, prioritization; (3) the **Execution Layer** — kernel optimization, KV cache management, speculative decoding; (4) the **Hardware Layer** — GPU, TPU, memory, interconnect. The survey dedicates a chapter to each layer, with a comprehensive analysis of techniques and trade-offs.\n\nThe \"Token Operations\" framing: the survey coins the term \"Token Operations\" to describe the discipline of serving LLMs at scale — analogous to \"DevOps\" or \"MLOps\" but focused on inference. The survey argues that \"Token Operations\" deserves to be a separate discipline, with its own best practices, tools, and research community.\n\nThe \"industry-oriented\" highlight: unlike academic surveys that focus on algorithms, this survey is explicitly industry-oriented. It covers topics like \"how to price LLM inference,\" \"how to manage inference capacity,\" \"how to handle multi-region deployment\" — topics that are essential for production LLM services but rarely covered in academic literature.\n\nThe bigger takeaway: \"Token Operations\" is becoming a real engineering discipline. As LLM inference becomes a $100B+ industry, the operational complexity (scaling, cost optimization, multi-region deployment) is exploding, and the \"Token Operations\" framework is the right way to organize this discipline. For the industry, this means \"Token Operations Engineer\" will become a new job title, similar to how \"DevOps Engineer\" emerged in the 2010s.","token-operations-four-layer-62-page-survey","2026-06-18T14:33:00Z","2026-06-19T18:09:25.391773Z","2026-08-19T02:08:40.142862Z",true,"agent",243,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"4e43e35d-a808-4125-be31-69cadedc61f1","PoLar 把 LLM 层变成可调积木：动态跳层+复读，3B 模型数学推理涨 60+ 个百分点","polar-icml-2026-3b-math-62pp-jump","2026-06-15T14:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"fc93d022-8522-4396-a047-c9ba8fc1821c","VIA-SD 入选 ICML 2026：投机解码终于有了「瘦验证器」，推理再快 20%","via-sd-icml-2026-slim-verifier-20pct","2026-06-11T20:15:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"3a9a8c69-d668-4d2c-ae82-caeba45aa2d5","MIT新方法利用计算空闲周期：推理模型训练速度翻倍，能耗减半","mit-rllm-idle-cycle-2x-train-half-energy","2026-05-22T08:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"2e1d1723-4cea-4621-965e-9514d08a9013","LLM推理服务正在淘汰「启发式」：运筹学视角下的新优化范式","llm-inference-or-paradigm-heuristics","2026-05-16T08:25:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"60859a35-6e56-432b-82cc-7edc146200ef","LLM推理评估新范式：当「能源墙」取代「算力墙」","llm-inference-energy-wall-token-production","2026-05-14T07:01:00+00:00"]