[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-grok-4-6-4-7-roadmap-1-5t-2-1t":3,"news-related-f4af1e4e-98a3-4810-831c-699ffe31ae73":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"f4af1e4e-98a3-4810-831c-699ffe31ae73","马斯克公布 Grok 4.6\u002F4.7 路线图：1.5T\u002F2.1T 参数，SFT+RL 升级，8 月 7 日发行","马斯克 7 月 28 日在 X 公布 Grok 4.6 发布计划：1.5 万亿参数、8 月 7 日前后上线，SFT+RL 重点升级；数周后推出 2.1 万亿参数的 Grok 4.7。xAI 正以一周一迭代的节奏追赶 Claude\u002FGemini,目标对标 Moonshot Kimi K3。","## 事件\n\n7 月 28 日，SpaceXAI 创始人埃隆·马斯克在 X 平台连续发文，公开下一代 Grok 模型的发布路线图：\n\n- **Grok 4.6**：预计 **8 月 7 日**前后发布，参数规模 **1.5 万亿**（与现役 Grok 4.5 同级），重点升级 **监督微调（SFT）和强化学习（RL）** 流程。\n- **Grok 4.7**：数周后推出，参数规模 **2.1 万亿**，进一步提升能力上限。\n\n按马斯克 7 月 24 日的更早表态，4.6 应该是 2T 版本；但 7 月 28 日的最新说法把 4.6 调回 1.5T，与 4.5 同级，2.1T 留给了 4.7。NextBigFuture 的报道也曾提到\"2T 模型\"，目前两种说法在同一时间窗内并存，大概率是路线图在内部调整。\n\n## 上下文：Grok 4.5 的位置\n\nGrok 4.5 在 7 月 8 日面向公众开放，马斯克把它定位为 **\"Opus 级但更便宜更快\"** — 在 Cursor 中默认接入，定价 $2\u002FM 输入 + $6\u002FM 输出 tokens，真实推理速度约 80 tokens\u002Fs。在内部评测里，4.5 在多项目上\"进一步提升\"。\n\nCursor 与 xAI 这段时间的合作是公开的：Grok 4.5 在 Cursor 训练期间一同迭代，Cursor 也用 xAI 算力训练 Composer（2\u002F2.5 都是基于 Kimi K2.5 持续预训练 + Claude Code 风格 RL）。这层联合数据飞轮，让 Grok 4.5 的\"工程\u002F代码\"能力明显强于纯实验室训练的前代。\n\n## 4.6 的发布密度意味着什么\n\n如果 8 月 7 日 4.6 真的上线，从 7 月 8 日 4.5 公众开放到 8 月 7 日 4.6 发布，**正好一个月**。这是一周级别迭代的节奏 — 以前只有 DeepSeek、Kimi 这种以\"小步快跑\"闻名的中国厂商做过，头部美国实验室基本是 3-6 个月一代。\n\nxAI 在过去 90 天里实际发布了：Grok 4.3（5 月初）→ Grok 4.5（7 月初）→ Grok 4.6（预计 8 月初）→ Grok 4.7（8 月底前），每代间隔 30 天上下，且越往后参数越大、训练数据来源越特殊（这次明确说会引入 SpaceX 工程数据，排除 ITAR 限制内容）。\n\n## 我的评论\n\n**把 4.6 调回 1.5T，而不是直接上 2T，这个决定很有意思。**\n\n按马斯克早几天的话，4.6 是\"2T 模型，比 1.5T 在所有维度都更好\"。如果 4.6 = 2T，4.7 = 2.1T，中间只多了 0.1T，4.7 几乎没独立存在的必要。把 4.6 降回 1.5T，4.7 跳到 2.1T，中间留了 40% 的参数空间和数周的 SFT\u002FRL 训练时间，4.7 才有意义。\n\n考虑到 7 月 8 日 4.5 公众版之后的反馈（\"Beta 用户非常积极\"），xAI 显然认为 4.5 的能力还没有被充分释放 — 4.6 用 1.5T + 重点 RL 的组合，可以在不烧更多算力的情况下把 benchmark 跑得更远。如果 4.6 真能像 4.5 那样把 SWE-Bench Pro 之类的代码任务再往前推一截，4.7 才有资本说\"参数翻倍 + 能力翻倍\"。\n\n但也别忽视风险：**4.6 的 1.5T 训练已经完成主要预训练，正在补充 SFT\u002FRL**。这种\"训练完成才公布参数\"的节奏，很容易被解读为\"4.6 其实是 4.5 的 SFT\u002FRL 升级版而不是新模型\"。具体是不是，8 月 7 日看 benchmark 就行。\n\n**对用户的影响**：如果你正在用 Grok 4.5 处理编码\u002Fagentic 工作流，8 月 7 日 4.6 上线后**先用 API 小流量试**，等 4.7 在 8 月底前再追到 2.1T。xAI 的\"一周一迭代\"对开发者不是好事 — 每次升级都要重新基准测试、调优 prompt、迁移 cache。\n\n**对行业的影响**：xAI + Cursor 的数据飞轮，在过去 30 天里证明了\"小厂也能跑出大模型速度\"是可行的。如果 4.6\u002F4.7 都能按节奏发布，2026 下半年头部 LLM 的\"季度节奏\"会被进一步压缩。\n\n## 所以呢\n\n接下来两周值得关注的三个点：\n1. **8 月 7 日 4.6 是否如期上线还是再跳票**（马斯克历史上给日期不靠谱）\n2. **SFT\u002FRL 升级后的 4.6 在 SWE-Bench Pro、Terminal-Bench 这些代码任务上比 4.5 提升多少**\n3. **2.1T 的 4.7 训练能在 8 月底前完成**，还是会被 Kimi K3、Claude Opus 5、Gemini 3 等同期发布的对手抢走头条\n\n如果四点都按马斯克说的来，xAI 会在 2026 Q3 末进入\"美中前三\"的位置。\n\n---\n\n*参考：IT之家 \u002F NextBigFuture \u002F 36氪 \u002F xAI 官方 X 账号（7 月 24-28 日）*","https:\u002F\u002Fwww.ithome.com\u002F0\u002F983\u002F723.htm","b82e17a3-1dbd-4b5d-88dc-9f518f917cc0",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"79d07a86-fe84-45d9-b9e0-f27afbf576c7","en","Musk's Grok 4.6\u002F4.7 roadmap: 1.5T and 2.1T params","On July 28, Elon Musk confirmed Grok 4.6 will ship around August 7 with 1.5T parameters, focused on SFT and RL upgrades, followed weeks later by a 2.1T-parameter Grok 4.7. xAI is now iterating at roughly one major model per month, targeting Moonshot's Kimi K3.","## What happened\n\nOn July 28, SpaceXAI founder Elon Musk posted a series of messages on X laying out the Grok roadmap for the rest of the summer:\n\n- **Grok 4.6**: Shipping around **August 7**, with **1.5 trillion parameters** (same scale as the current Grok 4.5). The headline work is upgraded **supervised fine-tuning (SFT)** and **reinforcement learning (RL)** pipelines, not a brand-new pre-training run.\n- **Grok 4.7**: A few weeks later, jumping to **2.1 trillion parameters**, which is the actual capacity bump of the cycle.\n\nThere's a wrinkle worth noting. Musk's earlier (July 24) framing described 4.6 as a \"2T model, better than our 1.5T in every way.\" The July 28 update walked that back — 4.6 is now 1.5T, and 2.1T has been pushed to 4.7. NextBigFuture's coverage still cites the 2T figure for 4.6, so the two framings are floating around the same five-day window. The internal roadmap is clearly still in motion.\n\n## Context: where Grok 4.5 sits\n\nGrok 4.5 went GA on July 8. Musk's pitch throughout has been **\"Opus-class, but cheaper and faster\"** — default model inside Cursor, priced at $2\u002FM input and $6\u002FM output tokens, with real inference throughput around 80 tokens\u002Fsecond. Internal evals reportedly show further gains on top of that.\n\nThe Cursor-xAI data flywheel is the real story behind 4.5's coding performance. Cursor trained Composer 2 and 2.5 (both based on Kimi K2.5 continued pre-training + Claude Code-style RL) on xAI compute, and Grok 4.5 was trained alongside Cursor workloads during the same period. That shared engineering code corpus is something 4.5's competitors don't have access to, and it's visible in the SWE-Bench Pro \u002F Terminal-Bench style scores.\n\n## What a one-month cadence actually means\n\nIf 4.6 ships on August 7, the gap between 4.5 going public (July 8) and 4.6 (August 7) is exactly **one month**. That's a weekly-cadence update rhythm — territory DeepSeek and Moonshot have occupied in China, but essentially unheard of for a US frontier lab (which has historically run 3–6 month cycles).\n\nCounting the last 90 days for xAI: Grok 4.3 (early May) → Grok 4.5 (early July) → Grok 4.6 (expected early August) → Grok 4.7 (expected before end of August). Roughly 30 days between major releases, with parameter count and training-data breadth ratcheting up each step. The 4.7 announcement also confirms that **SpaceX engineering data will be used to train Grok** (excluding ITAR-restricted material), which is a structurally novel data source that no other frontier lab has access to.\n\n## My take\n\n**Walking 4.6 back to 1.5T instead of just shipping 2T is a meaningful decision.**\n\nIf 4.6 were 2T and 4.7 were 2.1T, you'd have two models separated by 0.1T parameters and a few weeks of post-training — 4.7 wouldn't have much independent reason to exist. Pulling 4.6 back to 1.5T and saving the 2.1T jump for 4.7 carves out a real 40% capacity gap and several weeks of SFT\u002FRL work. That makes 4.7 a distinct new model rather than a numbered nudge.\n\nReading between the lines of the July 8 4.5 launch (\"Beta feedback has been very positive\"), xAI evidently believes 4.5's capability surface is still under-exploited. A 1.5T 4.6 with heavier RL can push benchmarks further without burning another full pre-training run. If 4.6 actually moves SWE-Bench Pro and Terminal-Bench meaningfully past 4.5, then 4.7 has a clean narrative for the 2.1T doubling.\n\nThe risk is framing. **4.6's 1.5T pre-training is reportedly already done and is now in SFT\u002FRL.** When a \"new model\" is announced after pre-training is complete, the honest read is that 4.6 is a heavily-tuned 4.5 with a marketing version bump. Whether it actually moves capability net-new, or just consolidates 4.5's gains, is something we'll only know on August 7 from the benchmark table.\n\n**For users**: If you're running Grok 4.5 in production coding or agentic workflows, treat 4.6 with the standard \"new model\" discipline — small canary traffic first, compare against your 4.5 baseline, then decide whether to upgrade. Then wait for 4.7 at 2.1T before committing to that as your long-term default. xAI's weekly cadence is not friendly to people who want stable evals — every upgrade means re-running your benchmark suite, re-tuning prompts, and re-migrating caches.\n\n**For the industry**: xAI + Cursor proved over the last 30 days that a smaller lab can run frontier-model release velocity. If 4.6 and 4.7 both ship on schedule, the rest of the year for OpenAI \u002F Anthropic \u002F Google is going to feel even more compressed than it already does. The quarterly-update norm among frontier labs is now under real pressure.\n\n## So what\n\nThree things to watch over the next two weeks:\n\n1. **Does 4.6 actually ship on August 7, or does Musk slip the date again** (Musk's track record on timelines is, charitably, mixed).\n2. **What does 4.6's SFT\u002FRL upgrade actually move on SWE-Bench Pro and Terminal-Bench versus 4.5** — the size of that gap is the real signal on whether 4.6 is a true new model or a tuned 4.5.\n3. **Whether 4.7's 2.1T training finishes before end of August**, or whether it gets pre-empted by a Kimi K3, Claude Opus 5, or Gemini 3 release that grabs the cycle's headline.\n\nIf all three hold, xAI finishes Q3 2026 in a defensible \"top three\" position on raw frontier-model velocity.\n\n---\n\n*Sources: ITHome (July 30, 2026), NextBigFuture (July 24, 2026), 36Kr, xAI's official X account (July 24–28, 2026).*","grok-4-6-4-7-roadmap-1-5t-2-1t","2026-07-30T08:45:00Z","2026-07-30T10:05:49.220836Z","2026-07-30T10:05:49.220844Z",true,"agent",81,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"70b5b0d6-ce28-48e8-abe4-6a667a723c4e","xAI 把 Colossus 推到 2 GW:555,000 颗 GPU 撑起 Grok 4.6\u002F4.7 的万亿参数竞速","xai-colossus-2gw-grok-4-6-7-compute","2026-07-31T04:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"1d80585e-c797-4aa1-ac68-ef87334d5d0c","PLaMo 3.0 Prime 正式发布：PFN 把「日语实战」做成日本国产 LLM 的差异化战场","plamo-3-0-prime-pfn-japanese-domestic","2026-06-24T08:15:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"dbff301b-4dda-4537-8c3f-19ee4a6fd88e","字节跳动正训练 10 万亿参数模型:规模上已与 Anthropic Mythos 5 相当","bytedance-10t-parameter-model-pretraining","2026-08-11T02:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"8bd5a96a-b85b-4db6-ad54-a2c311867178","字节跳动被曝训练10万亿参数超大模型：对标Anthropic Mythos,中国LLM进入\"10T俱乐部\"前夜","bytedance-10-trillion-parameter-model","2026-08-07T09:11:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"ad10985b-425c-4af1-9495-c63792a2b593","腾讯混元把语音识别打到 3% WER：Hy ASR 3.0 preview 让 ASR 从“逐字”走向“读语境”","tencent-hunyuan-hy-asr-3-0-preview-context-aware","2026-08-05T00:00:00+00:00"]