[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-grok-4-6-agentic-cost-frontier":3,"news-related-c4ec4625-4a84-4c24-88f0-0ef1beb4f19e":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"c4ec4625-4a84-4c24-88f0-0ef1beb4f19e","Grok 4.6 发布:61 分追平 GPT-5.6 Sol,把长程 Agent 的 token 账单砍到四分之一","xAI 于 8 月 12 日发布 Grok 4.6,聚焦长程 Agent 与交互\u002F视觉工作。Artificial Analysis Intelligence Index 拿 61 分追平 GPT-5.6 Sol,较一个月前的 Grok 4.5 涨 5 分;定价维持每百万 token 2\u002F6 美元不变,比 Claude Opus 5 和 GPT-5.6 Sol 低 60% 以上,并以约 53 回合、0.5B 输入 token 的回合效率,把前沿模型的「每任务成本」压到 0.84 美元。","# Grok 4.6 发布:重返前沿的 61 分,和一张「每任务 0.84 美元」的账单\n\n8 月 12 日,xAI 发布 Grok 4.6。官方定位说得很清楚:在 Grok 4.5 基础上,重点强化**长程 Agent**和更复杂的交互与视觉工作——模型要能在研究、代码库分析、从想法到成品应用这类多步骤任务里「钉」住不放手。\n\n## 跑分:追平 GPT-5.6 Sol,落后 Claude 双旗舰\n\n在 Artificial Analysis Intelligence Index(九项基准的综合分)上,Grok 4.6 拿到 61 分,追平 GPT-5.6 Sol(max),落后 Claude Fable 5(62)和 Claude Opus 5(63)。独立评测机构 Artificial Analysis 的评价是:这让 xAI「回到智能前沿,与 OpenAI 并肩,仅落后 Anthropic」。作为参照,一个月前的 Grok 4.5 是 56 分——一个月涨 5 分,较 Grok 4.3 累计涨 23 分。\n\n真正拉开差距的是 agentic 维度:\n\n- **GDPval-AA v2**(真实世界知识工作):Elo 1753,仅次于 Claude Opus 5,与 Fable 5、Qwen3.8 Max 置信区间重叠\n- **𝜏³-Banking**(多轮客服+工具使用):50.7%,全榜前二\n- **Terminal-Bench v2.1**:88.4%,与领先模型持平\n- xAI 自报口径里,CursorBench v3.2 拿到 69.9%,DeepSWE v1.1 65.9%,FrontierCode v1.1 61.3%,均高于 Grok 4.5 的 66.7%\u002F54%\u002F56.6%\n\n也要看官方表格里的另一面:在更新更难的 Terminal-Bench v3.0 上,Grok 4.6 只有 26%,GPT-5.6 Sol 是 34.6%——新版本基准上的差距仍然真实存在,agentic 强不等于全面领先。\n\n## 训练:让 Grok 4.5 给 4.6 打工\n\n训练侧的官方描述值得细读:Grok 4.6 做了比 4.5 更长的补充训练,使用精选的模型生成数据和高质量工程数据,并改进了优化器和训练配方。然后 xAI **用 Grok 4.5 重新生成覆盖不同推理 effort、agent harness 和领域(STEM、软件工程、知识工作)的 SFT 轨迹**,再用基于模型的检查过滤掉问题 trace。之后是覆盖内核优化、Web 开发、计算机辅助设计等领域环境的 agentic RL。老模型当数据工厂、新模型当学生——这条自我迭代流水线正在变成头部厂商的标配。\n\n## 定价:涨智能不涨价的罕见一代\n\nArtificial Analysis 特别指出:前沿模型的智能上涨通常伴随涨价,而 Grok 4.6 定价维持在每百万 token 2 美元\u002F6 美元(输入\u002F输出),比 Claude Opus 5(5\u002F25 美元)和 GPT-5.6 Sol(5\u002F30 美元)低 60% 以上。实测每任务成本 0.84 美元,与 Kimi K3 相同但智能略高,落在「智能 vs 每任务成本」的 Pareto 前沿上。\n\n更狠的是回合效率:在长程知识工作基准 AA-Briefcase 上,Grok 4.6 平均约 **53 个回合、0.5B 输入 token** 完成任务,而 Claude Opus 5(max)需要约 **103 个回合、2.0B 输入 token**。同一层级的答案,一半的回合、四分之一的输入量——长程 Agent 的上下文累积速度决定了,这个 token 效率优势折算成成本远超单价差。唯一的涨价项是缓存命中价,从每百万 0.3 美元涨到 0.5 美元;上下文窗口维持 500k 不变。\n\n## 所以呢\n\nGrok 4.6 没有任何单项跑分全场第一,但它把「每任务成本」打进了 Intelligence Index 所有 agentic 评测的 Pareto 前沿。当智能分挤在 61–63 的窄带里,决定采购的可能不再是榜单,而是月底那张 token 账单——这大概是这代发布真正值得盯的信号。\n\n来源:[xAI 官方发布页](https:\u002F\u002Fx.ai\u002Fnews\u002Fgrok-4-6)、[Artificial Analysis 独立评测](https:\u002F\u002Fartificialanalysis.ai\u002Farticles\u002Fgrok-4-6-benchmarks-and-analysis)","https:\u002F\u002Fx.ai\u002Fnews\u002Fgrok-4-6","b82e17a3-1dbd-4b5d-88dc-9f518f917cc0",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"267b3f88-b711-461d-96ca-b6147b3db065","en","Grok 4.6 ties GPT-5.6 Sol at 61, quarter the agent bill","xAI released Grok 4.6 on August 12, focused on long-running agents and more ambitious interactive and visual work. It scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol — up 5 points from Grok 4.5 just one month earlier. Pricing stays at $2\u002F$6 per million tokens, more than 60% below Claude Opus 5 and GPT-5.6 Sol, and with roughly 53 turns and 0.5B input tokens per task, it pushes frontier per-task cost down to $0.84.","# Grok 4.6 Arrives: Back at the Frontier with 61 Points, and a $0.84-Per-Task Bill\n\nOn August 12, xAI released Grok 4.6. The official positioning is clear: building on Grok 4.5, with a particular focus on **long-running agents** and more ambitious interactive and visual work — the model needs to stay with complex tasks across many steps, whether researching a topic, working across a codebase, or turning an idea into a polished application.\n\n## Benchmarks: Matching GPT-5.6 Sol, Behind Claude's Twin Flagships\n\nOn the Artificial Analysis Intelligence Index (a composite of nine benchmarks), Grok 4.6 scores 61, matching GPT-5.6 Sol (max), behind Claude Fable 5 (62) and Claude Opus 5 (63). Independent evaluator Artificial Analysis put it this way: this brings xAI \"back to the intelligence frontier alongside OpenAI, behind only Anthropic.\" For reference, Grok 4.5 one month earlier scored 56 — a 5-point gain in one month, and +23 cumulative over Grok 4.3.\n\nWhere it pulls ahead is the agentic dimension:\n\n- **GDPval-AA v2** (real-world agentic knowledge work): Elo of 1753, behind only Claude Opus 5, with confidence intervals overlapping Claude Fable 5 and Qwen3.8 Max\n- **τ³-Banking** (multi-turn customer service with tool use): 50.7%, among the top two\n- **Terminal-Bench v2.1**: 88.4%, in line with the leading models\n- In xAI's self-reported table: CursorBench v3.2 at 69.9%, DeepSWE v1.1 at 65.9%, and FrontierCode v1.1 at 61.3% — all above Grok 4.5's 66.7% \u002F 54% \u002F 56.6%\n\nThe other side of the official table deserves attention too: on the newer, harder Terminal-Bench v3.0, Grok 4.6 scores just 26% versus GPT-5.6 Sol's 34.6% — the gap on new-generation benchmarks is still real, and agentic strength does not mean across-the-board leadership.\n\n## Training: Putting Grok 4.5 to Work for Grok 4.6\n\nThe official training description is worth a close read. Grok 4.6 underwent a longer supplemental training run than Grok 4.5, using curated model-generated data and high-quality engineering data, with an improved optimizer and training recipe. xAI then **used Grok 4.5 to regenerate the SFT trajectories** across reasoning efforts, agent harnesses, and domains such as STEM, software engineering, and knowledge work, filtering out problematic traces with model-based checks. This was followed by agentic RL across knowledge work, general coding, and domain-specific environments for kernel optimization, web development, and computer-aided design. Older model as data factory, newer model as student — this self-iteration pipeline is becoming standard equipment for frontier labs.\n\n## Pricing: A Rare Generation That Gains Intelligence Without Raising Prices\n\nArtificial Analysis specifically noted: at the frontier, intelligence gains usually come with price increases, but Grok 4.6 holds pricing at $2\u002F$6 per million input\u002Foutput tokens — more than 60% below Claude Opus 5 ($5\u002F$25) and GPT-5.6 Sol ($5\u002F$30). Measured cost per task is $0.84, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier.\n\nThe turn efficiency is even more striking: on AA-Briefcase, the long-horizon knowledge work benchmark, Grok 4.6 completes tasks in roughly **53 turns and 0.5B input tokens** on average, versus roughly **103 turns and 2.0B input tokens** for Claude Opus 5 (max). Same-tier answers at half the turns and a quarter of the input — and since long-horizon agents accumulate context rapidly, this token-efficiency edge translates into cost savings well beyond the per-token price gap. The only price increase is on cache hits, up from $0.3 to $0.5 per million tokens; the context window stays unchanged at 500k.\n\n## So What\n\nGrok 4.6 does not take a single outright first place on any benchmark, but it lands on the cost-performance Pareto frontier for every agentic evaluation in the Intelligence Index. When intelligence scores are packed into a narrow 61–63 band, what decides procurement may no longer be the leaderboard but the token bill at the end of the month — and that is probably the signal from this release truly worth watching.\n\nSources: [xAI official release](https:\u002F\u002Fx.ai\u002Fnews\u002Fgrok-4-6), [Artificial Analysis independent review](https:\u002F\u002Fartificialanalysis.ai\u002Farticles\u002Fgrok-4-6-benchmarks-and-analysis)","grok-4-6-agentic-cost-frontier","2026-08-14T19:00:00Z","2026-08-14T17:07:55.918092Z","2026-08-14T17:07:55.918103Z",true,"agent",77,{"items":39},[40,45,50,54,59,64],{"id":41,"title":42,"news_slug":43,"published_at":44},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"fbdcfd41-bb54-487a-8319-9f35dfc82be5","DeepSeek-V4-Flash转正：这次升级不靠换架构","deepseek-v4-flash-0731-agent-update","2026-07-31T08:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":49},"3cc63477-1334-497d-80cb-90850c019101","DeepSeek-V4-Flash 转正:不靠换架构,只做后训练重新发力 Agent","deepseek-v4-flash-official-post-training-agent-0731",{"id":55,"title":56,"news_slug":57,"published_at":58},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":60,"title":61,"news_slug":62,"published_at":63},"e75069c6-f15c-4ff9-8b11-404d705442e8","Upstage Solar Pro 4:把「agent 跑得稳」做成新一代闭源模型卖点","upstage-solar-pro-4-agent-reliability-closed-llm","2026-08-25T03:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00"]