[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-july-2026-five-frontier-models":3,"news-related-5bfdf32b-44eb-4eb5-a98b-39e921168182":39},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":26,"news_slug":32,"published_at":33,"created_at":34,"modified_at":35,"is_published":36,"publish_type":37,"image_url":13,"view_count":38},"5bfdf32b-44eb-4eb5-a98b-39e921168182","九天内连发五款前沿模型:7 月的大模型军备赛,真正决胜负的不再是 benchmark","7 月 8 日到 16 日,Grok 4.5、GPT-5.6、Muse Spark 1.1、Inkling、Kimi K3 五款前沿模型相继亮相,密度史上最高。但当五家实验室的能力都跑到了够用的水平线之上,benchmark 表就已经不是排名,而是营销选择。Build Fast with AI 通过 API 与真实工作流测试这五款模型后发现:真正决定选型的,变成了单 token 价格、许可证条款、能否微调,以及你究竟在跑哪种任务。\n\nKimi K3 把\"规模买能力\"推到极致,2.8 万亿参数、1M 上下文、原生视频输入,跑出 91.2% BrowseComp 和 93.5% GPQA Diamond 开源档第一,代价是 \u002F5 的高定价。GPT-5.6 Sol 在自研 Cerebras 硬件上跑到 750 tok\u002Fs,Tier 拆分( Sol\u002FTerra\u002FLuna)直接覆盖从批量到旗舰的全价格带。Inkling 是榜单上唯一 Apache 2.0 真开源权重(975B 总参 \u002F 41B 激活),附带可调思考强度拨杆,Thinking Machines 自己说它\"不是最强模型\"——明牌定位为基础底座,不是榜单选手。Meta Muse Spark 1.1 把价格砸到 .25\u002F.25,却同时拿下 MCP Atlas 88.1 分的工具调用 SOTA,支持文本\u002F图像\u002F视频\u002F音频\u002FPDF 单端点接入,是迄今最便宜也最多模态的智能体大脑。Grok 4.5 走 X 平台原生搜索 + 500K 上下文的差异化路线,\u002F 价格配上独立 Artificial Analysis 验证的智能指数,稳坐价值档。\n\n价格跨度已经超过 12 倍但能力差距只几个百分点。Build Fast with AI 的测试里,Muse Spark 在 3D 房屋生成任务上花 4 美分,Fable 5 要 1.82 美元——同样的输出,45 倍价差。当单一榜单不再能定义胜负,选型逻辑就要重写:把 80% 日常流量路由到 Muse Spark 1.1、Grok 4.5 或 GPT-5.6 Luna,把 Sol 或 Fable 5 留给那 20% 真正的硬骨头;只要任务可重复,Inkling 的微调路径就是比租用闭源模型更划算的长线投资。\n\n七月的真正信号不是某个模型封神,而是\"能力门槛\"被五家同时跨过之后,大模型竞争从卷分数正式切换到卷成本、卷许可、卷生态——开源权重重新回到牌桌中央,价格战从 token 蔓延到按完成任务计费,定制化反而成了最稀缺的护城河。","https:\u002F\u002Fwww.buildfastwithai.com\u002Fblogs\u002Fnew-ai-models-july-2026","5d6d4f6b-780a-419e-8057-20a8c1bbaa60",[10,14,17,20,23],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":24,"name":25,"slug":25,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[27],{"id":28,"lang":29,"title":30,"summary":31,"content":31},"b341cf21-d8eb-49e6-8c91-eea2f5aa2685","en","Five frontier models in nine days: July's real race","From July 8 to 16, five frontier models — Grok 4.5, GPT-5.6, Muse Spark 1.1, Inkling, Kimi K3 — appeared in succession, the highest density ever. But when all five labs' capabilities have crossed the \"good enough\" line, the benchmark table is no longer a ranking, but a marketing choice. Build Fast with AI, after testing these five models through APIs and real workflows, found that what truly decides selection has become per-token price, license terms, fine-tunability, and the actual task you're running. Kimi K3 pushes \"buy capability with scale\" to the limit — 2.8T parameters, 1M context, native video input, scoring 91.2% BrowseComp and 93.5% GPQA Diamond as the open-source #1, at the cost of high pricing. GPT-5.6 Sol runs at 750 tok\u002Fs on its self-developed Cerebras hardware, with the tier split (Sol\u002FTerra\u002FLuna) directly covering the full price band from batch to flagship. Inkling is the only true Apache 2.0 open-weight on the leaderboard (975B total \u002F 41B active), with an adjustable thinking-intensity knob, and Thinking Machines itself says it's \"not the strongest model\" — openly positioned as a base layer, not a leaderboard entrant. Meta Muse Spark 1.1 slams price down to $0.25 \u002F $0.25, yet simultaneously takes the tool-calling SOTA on MCP Atlas with 88.1, supporting text \u002F image \u002F video \u002F audio \u002F PDF single-endpoint access — to date the cheapest and most multimodal agent brain. Grok 4.5 takes the differentiated path of X-platform-native search + 500K context, with $0.50 \u002F $1.50 pricing plus independently Artificial-Analysis-verified intelligence index, holding the value tier. The price spread has already exceeded 12x, but the capability gap is only a few percentage points. In Build Fast with AI's test, Muse Spark spent 4 cents on a 3D house generation task, Fable 5 spent $1.82 — same output, 45x price difference. When a single leaderboard no longer defines winners, selection logic has to be rewritten: route 80% of daily traffic to Muse Spark 1.1, Grok 4.5, or GPT-5.6 Luna; save Sol or Fable 5 for the 20% real hard problems; as long as the task is repeatable, Inkling's fine-tuning path is a more cost-effective long-term investment than renting a closed-source model. July's real signal isn't that some model became a god, but after \"the capability threshold\" was crossed by all five at the same time, LLM competition has formally switched from competing on scores to competing on cost, license, and ecosystem — open weights are back at the center of the table, the price war has spread from per-token to per-completed-task, and customization has become the rarest moat.","july-2026-five-frontier-models","2026-07-23T12:00:00Z","2026-07-23T04:03:09.506920Z","2026-08-19T02:08:40.142862Z",true,"agent",104,{"items":40},[41,46,51,56,61,66],{"id":42,"title":43,"news_slug":44,"published_at":45},"6f1f105b-8e80-4b2c-b88c-b392556952aa","2026年本地LLM深度评测：开源模型性能全解析","local-llm-2026-deep-eval-swe-bench-aime","2026-04-25T11:15:00+00:00",{"id":47,"title":48,"news_slug":49,"published_at":50},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":52,"title":53,"news_slug":54,"published_at":55},"d4fa7e14-8fbd-4940-93a6-3dd6f0a3991d","DeepSeek V4 Pro 正式版：1.6T MoE，1M 上下文","deepseek-v4-pro-0813-ga-1m-context-moe","2026-08-13T02:00:00+00:00",{"id":57,"title":58,"news_slug":59,"published_at":60},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"80315de0-7eb3-491a-b2e6-103a691a8bd7","Nanbeige4.2-3B 用 Looped Transformer 在 11 项基准上跑赢 Qwen3.5-9B","nanbeige-4-2-3b-looped-transformer-agentic-3b-beats-qwen3-5-9b","2026-07-30T10:30:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"dfc3dec4-2211-4c7e-b6ff-9e0d9a479ec4","微软与 Mistral 签下数十亿美元协议:Vera Rubin GPU 上的「欧洲主权云」开始落地","microsoft-mistral-vera-rubin-sovereign","2026-07-22T02:00:00+00:00"]