In April 2026, LLM vendors staged an unprecedented arms race: nine frontier models released in 30 days, from DeepSeek V4-Pro and Kimi K2.6 to GLM-5.1 and Qwen 3.6 — the industry seemed caught in a perpetual acceleration sprint.

But just as the field was digesting this wave of releases, May's landscape quietly pivoted to a different axis — the model layer went quiet, the infrastructure layer lit up.

Sparse Attention Enters Production-Grade Engineering

DeepSeek V4's compressed-attention mechanism raised the bar for inference frameworks, and its sparse-attention profile demands more of the serving stack. In early May, SGLang and Miles, the two leading open-source inference frameworks, both announced Day-0 support for DeepSeek V4 — a sign that sparse attention has officially moved from paper to production-grade engineering. The significance goes well beyond a single model port: it means the entire inference-engineering ecosystem is maturing fast enough to absorb new architectures and algorithms quickly.

Meanwhile, vLLM's production-grade support for MoE architectures continued to solidify through May — the deployment barrier for large-scale sparse models is falling fast.

The Benchmark Evaluation System Is Being Rebuilt

In April, UC Berkeley's RDI published a report exposing widespread contamination in mainstream agent benchmarks, triggering an industry-wide reckoning over "which numbers can be trusted." Against this backdrop, May saw a wave of stricter evaluation frameworks: SWE-bench Pro introduced contamination-resistant mechanisms, re-measuring model coding ability in cleaner environments; GDPval covers 44 knowledge-worker occupations in scenario-based evaluation, trying to answer "what can models do in real work" rather than staying stuck on academic leaderboards.

In this context, the capability profile of each model is being re-calibrated. Claude Opus 4.7 hits 87.6% on SWE-bench Verified; Qwen 3.6 Max leads across six coding/agent benchmarks; DeepSeek V4-Flash, with its extremely low API price ($0.07 per million output tokens), has become the go-to for cost-sensitive scenarios — different models excel on different axes, and a simple leaderboard no longer captures the full picture.

The Competitive Logic Is Being Reframed

These infrastructure-level shifts are reshaping the basic logic of LLM competition. Chinese vendors proved efficient training was possible with DeepSeek R1 in early 2025; in 2026, low-cost inference and fast engineering adaptation are the new differentiation axes. SGLang/Miles' Day-0 support for DeepSeek V4 is the perfect annotation of this "infrastructure as competitiveness" logic.

As model capability itself converges (at least on certain dimensions), inference efficiency, deployment ease, and evaluation credibility are becoming the factors that actually decide adoption in the next phase. This race may have just entered its most interesting stretch.