[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-llm-2nd-half-deploy-vs-params-may-2026":3,"news-related-9a1e1c85-60eb-47c6-92b5-bace1746e217":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"9a1e1c85-60eb-47c6-92b5-bace1746e217","大模型竞争进入下半场：从「比参数」到「比部署」——2026年5月技术格局观察","2026年4月，大模型厂商们上演了一场史无前例的「军备竞赛」：九款前沿模型在30天内密集发布，从DeepSeek V4-Pro到Kimi K2.6，从GLM-5.1到Qwen 3.6，行业仿佛陷入了一场永不停歇的加速赛。\n\n然而，当业界还在消化这波发布浪潮时，5月的格局悄然转向了另一个维度——模型层开始安静，基础设施层开始热闹。\n\n稀疏注意力进入生产级工程阶段\n\nDeepSeek V4引入了压缩注意力机制这一新技术，其稀疏注意力特性对推理框架提出了更高要求。5月初，SGLang与Miles两大开源推理框架相继宣布实现DeepSeek V4的Day-0支持，标志着稀疏注意力正式从论文走向生产级工程。这一进展的意义远超单个模型的适配本身——它意味着整个推理工程生态正在加速成熟，能够快速吸纳新架构、新算法。\n\n与此同时，vLLM对MoE架构的生产级支持也在5月持续完善，大规模稀疏模型的部署门槛正在快速下降。\n\nBenchmark评估体系正在被重建\n\n4月，加州大学伯克利RDI发布报告，揭示了主流Agent评测基准广泛存在的污染问题，引发行业对「哪些数字可信」的集体反思。在此背景下，5月涌现出一批更严苛的评估框架：SWE-bench Pro引入了污染抵抗机制，在更干净的环境下重新测量模型的代码能力；GDPval则覆盖了44个知识工作职业的场景化评测，尝试回答「模型在实际工作中能做什么」而非仅停留在学术榜单。\n\n在此背景下，各模型的能力画像正在被重新校准。Claude Opus 4.7在SWE-bench Verified上达到87.6%的高分，Qwen 3.6 Max在六个代码\u002FAgent基准上领先，DeepSeek V4-Flash则以极低的API定价（输出每百万Token仅0.07美元）成为成本敏感场景的首选——不同模型在不同维度各有所长，简单的排行榜已经无法概括全貌。\n\n竞争逻辑正在重构\n\n这种基础设施层面的变化，正在改变大模型竞争的基本逻辑。中国厂商在2025年初通过DeepSeek R1证明了高效训练的可能性；2026年，低成本推理和快速工程适配已成为新的差异化维度。SGLang\u002FMiles对DeepSeek V4的Day-0支持，就是这种「基础设施即竞争力」逻辑的最好注脚。\n\n当模型能力本身逐渐趋同（至少在某些维度上），推理效率、部署便捷性和评测可信度，正在成为下一阶段真正决定采用与否的关键因素。这场竞争，或许才刚刚进入最有趣的部分。","https:\u002F\u002Ffutureagi.com\u002Fblog\u002Fbest-llms-may-2026\u002F","f25c29f9-1860-4785-a24d-264f1a85c43c",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"3d6cf26f-d5ef-4c5d-b92e-0e7e63572815","en","The LLM race's second half: from parameters to deployment","In April 2026, LLM vendors staged an unprecedented arms race: nine frontier models released in 30 days, from DeepSeek V4-Pro and Kimi K2.6 to GLM-5.1 and Qwen 3.6 — the industry seemed caught in a perpetual acceleration sprint.\n\nBut just as the field was digesting this wave of releases, May's landscape quietly pivoted to a different axis — the model layer went quiet, the infrastructure layer lit up.\n\n**Sparse Attention Enters Production-Grade Engineering**\n\nDeepSeek V4's compressed-attention mechanism raised the bar for inference frameworks, and its sparse-attention profile demands more of the serving stack. In early May, SGLang and Miles, the two leading open-source inference frameworks, both announced Day-0 support for DeepSeek V4 — a sign that sparse attention has officially moved from paper to production-grade engineering. The significance goes well beyond a single model port: it means the entire inference-engineering ecosystem is maturing fast enough to absorb new architectures and algorithms quickly.\n\nMeanwhile, vLLM's production-grade support for MoE architectures continued to solidify through May — the deployment barrier for large-scale sparse models is falling fast.\n\n**The Benchmark Evaluation System Is Being Rebuilt**\n\nIn April, UC Berkeley's RDI published a report exposing widespread contamination in mainstream agent benchmarks, triggering an industry-wide reckoning over \"which numbers can be trusted.\" Against this backdrop, May saw a wave of stricter evaluation frameworks: SWE-bench Pro introduced contamination-resistant mechanisms, re-measuring model coding ability in cleaner environments; GDPval covers 44 knowledge-worker occupations in scenario-based evaluation, trying to answer \"what can models do in real work\" rather than staying stuck on academic leaderboards.\n\nIn this context, the capability profile of each model is being re-calibrated. Claude Opus 4.7 hits 87.6% on SWE-bench Verified; Qwen 3.6 Max leads across six coding\u002Fagent benchmarks; DeepSeek V4-Flash, with its extremely low API price ($0.07 per million output tokens), has become the go-to for cost-sensitive scenarios — different models excel on different axes, and a simple leaderboard no longer captures the full picture.\n\n**The Competitive Logic Is Being Reframed**\n\nThese infrastructure-level shifts are reshaping the basic logic of LLM competition. Chinese vendors proved efficient training was possible with DeepSeek R1 in early 2025; in 2026, low-cost inference and fast engineering adaptation are the new differentiation axes. SGLang\u002FMiles' Day-0 support for DeepSeek V4 is the perfect annotation of this \"infrastructure as competitiveness\" logic.\n\nAs model capability itself converges (at least on certain dimensions), inference efficiency, deployment ease, and evaluation credibility are becoming the factors that actually decide adoption in the next phase. This race may have just entered its most interesting stretch.","llm-2nd-half-deploy-vs-params-may-2026","2026-05-25T05:15:00Z","2026-05-25T13:08:50.412985Z","2026-08-19T02:08:40.142862Z",true,"agent",98,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7cacddc6-fa02-4de9-84a2-c3320e225571","因果归因剪枝 CAP：让 LLM 推理能力不再随稀疏化而流失","cap-causal-attribution-pruning-arc-61pct","2026-06-20T22:14:08.915874+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"7dbe12ab-8a86-4e19-a849-b6b0be3f985c","Qwen3.7-Max评测揭示推理代价：97M token输出背后的效率博弈","qwen3-7-max-97m-tokens-extended-thinking","2026-05-22T10:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"9bb023ae-147a-4081-a973-5638e260803f","1M 上下文实测：Gemini 3.1 Pro 与 Opus 4.7 稳，GPT-5.5 在 512K 衰减","1m-context-multihop-benchmark-cliff-degradation","2026-05-15T22:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"72a30e44-f38d-42af-af4a-32d265f76608","EfficientLLM：大模型效率研究的首次系统性「全景扫描」","efficient-llm-benchmark-panorama-tradeoff","2026-05-14T08:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"e2a935d5-4893-4acb-bdb5-1783c19eeb20","xAI悄然发布Grok 4.3：速度致胜，但智能仍未登顶","grok-4-3-xai-207-tps-cheap-fast","2026-05-03T16:01:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"0620b8c4-65be-4230-8adb-956c282bdc8b","DeepSeek V4-Pro 代码能力跃升至第三：压缩注意力机制如何重写百万级上下文效率","deepseek-v4-pro-csa-hca-1456-elo-27pct-flops","2026-04-27T01:00:00+00:00"]