GPT-5.6's 80% Price Cut Pulls Competition Into "Equal Intelligence Cost": DeepSeek V4 Flash Answers, Chinese Models Lock In Two-Front Strategy
On July 30, OpenAI slashed the API price of GPT-5.6 Luna by 80% and trimmed Terra by 20%. The company that had been gating enterprise accounts with "seat licenses" in 2025 has, for the first time, placed its flagship model on the "20%-of-list price war" track (36Kr report; Xiouwang cross-post). The real shock isn't how cheap Luna became — it's that the competitive axis of late-2026 LLMs has shifted from "intelligence ceiling" to "equal intelligence cost."
OpenAI Turns "Price" Into the Main Battlefield
Per Huatai Securities' research note, GPT-5.6 Luna leads the Artificial Analysis Intelligence Index at 51 points, with DeepSeek V4 Flash 0731 right behind at 50. The capability gap is one point. But the price gap is far steeper: V4 Flash's blended price runs around $0.06 per million tokens — about 65% lower than Luna; average task cost is around $0.03, roughly 57% lower. When "1 IQ point" and "a 40% unit-price discount" sit on the same slide, enterprise procurement math has effectively only one answer.
This isn't charity from OpenAI. In July 2026, Microsoft internally told engineers not to "maximize token usage" and switched Copilot's default model to GPT-5.6 Sol (CNBC report) — meaning OpenAI's largest customer is already budgeting per token. The Luna price cut is a textbook "trade margin for volume" maneuver, offloading accounting pressure onto API revenue.
DeepSeek V4 Flash: A Counter-Punch Without Changing Weights
Why does V4 Flash 0731 matter? Not because the architecture changed — it didn't. V4-Flash-0731's weights are identical to V4 Flash; only the post-training pipeline moved (DeepSeek official update). Result: the same 284B/13B MoE jumped 10 points on the Artificial Analysis Intelligence Index, with Agent scores punching above V4-Pro (Artificial Analysis evaluation).
The underlying logic: when scaling-law returns diminish and pretraining costs approach tens of billions of RMB, swapping a new post-training recipe on the same weights is often far cheaper than training another trillion-parameter model from scratch. DeepSeek DSpark hitting 60%–85% end-to-end speedups in production (GitHub repo) runs on the same logic — push inference-side cost limits to match last-generation pretraining.
The Two-Front Approach of Chinese Open-Weights
Huatai Securities splits today's Chinese open-weights landscape into two tracks:
- Strong-capability frontier is held by Kimi K3. K3 leads Chinese open-source ranks at 57 on the Intelligence Index, powered by a 2.8T-parameter MoE plus Kimi Delta Attention linear attention (Moonshot AI official blog). MoonEP/FlashKDA/AgentEnv — three open training stacks — shipped at end of July, exposing the full post-training stack for community reproduction (GitHub repo).
- Cost-efficiency floor is held by DeepSeek V4 Flash. $0.06 per million tokens blended pricing goes head-to-head with Luna at less than 40% of the cost.
These aren't the same track: K3 stakes out capability at the 50–57 band; V4 Flash competes on per-token price. For enterprises, that means they can now compose a hybrid pipeline — strong-capability model for planning, cost-efficient model for execution — without being held hostage by "the single best model."
Two Investment Mainlines
The two themes Huatai Securities flags — "AI applications" and "domestic models" — are really two sides of the same coin:
- Application side absorbs the token-price avalanche. On August 4, Chinese A-shares AI-application concept stocks rallied hard: Hanyi, Yidian, Bluesky, HuaSheng TianCheng and others hit daily-limit-up or surged over 10% (Sohu/Jinrongjie report), triggered specifically by the Luna 80% price-cut news. Morgan Stanley's Xing Ziqiang frames it as "AI investment entering a half-time break, pivoting from compute upstream to AI applications and HALO resource pairing" (Sohu report).
- Domestic-model side absorbs the capability-ceiling push. Kimi K3, DeepSeek V4 Flash, and ByteDance's 10-trillion-parameter model currently in training (Financial Times report) form a three-layer defense of "capability + scale + cost-efficiency." A 10T model in delivery means Chinese LLMs are formally entering the same weight class as Anthropic Mythos 5.
What Comes Next in the "Equal Intelligence Cost" Era
"Equal intelligence cost" isn't a temporary promo term — it means late-2026 LLM competition will play out on three new tensions:
- Architecture layer: With scaling-law returns diminishing, attention-mechanism diversity (linear attention, sparse attention, MoE) becomes the cost-reduction battlefield. Kimi K3's KDA and DeepSeek V4's CSA + HCA both come from this track.
- Training layer: Post-training paradigms (GRPO/OPD, speculative decoding, Agentic RL) will replace pretraining as the new "capability moat." DSpark and CURE (arXiv paper) and similar end-to-end inference-acceleration solutions will move into production in batches.
- Business layer: OpenAI must trade "lower price for ecosystem density"; DeepSeek must protect margins with "no weight change, new post-training"; Chinese open-weights must hold ground with a "capability + price" two-front stance — any single-point lead won't survive a quarter.
The next signal worth watching isn't "whose benchmark jumped again," but "whose equal-intelligence cost dropped another order of magnitude." When $0.03 average task cost per million tokens becomes the new baseline, the procurement voice shifts from "model selection" to "pipeline orchestration" — which is exactly the reason two-front products like Kimi K3 + DeepSeek V4 Flash exist.
Closing thought: OpenAI's 80% price cut isn't a sign the LLM bubble is bursting — it's a sign that the bubble is migrating from "capability valuation" to "capability-per-cost valuation." In the next shake-out, the question isn't who's smarter, but who can keep the books cheaper at the same level of smart.