What happened
On July 28, SpaceXAI founder Elon Musk posted a series of messages on X laying out the Grok roadmap for the rest of the summer:
- Grok 4.6: Shipping around August 7, with 1.5 trillion parameters (same scale as the current Grok 4.5). The headline work is upgraded supervised fine-tuning (SFT) and reinforcement learning (RL) pipelines, not a brand-new pre-training run.
- Grok 4.7: A few weeks later, jumping to 2.1 trillion parameters, which is the actual capacity bump of the cycle.
There's a wrinkle worth noting. Musk's earlier (July 24) framing described 4.6 as a "2T model, better than our 1.5T in every way." The July 28 update walked that back — 4.6 is now 1.5T, and 2.1T has been pushed to 4.7. NextBigFuture's coverage still cites the 2T figure for 4.6, so the two framings are floating around the same five-day window. The internal roadmap is clearly still in motion.
Context: where Grok 4.5 sits
Grok 4.5 went GA on July 8. Musk's pitch throughout has been "Opus-class, but cheaper and faster" — default model inside Cursor, priced at $2/M input and $6/M output tokens, with real inference throughput around 80 tokens/second. Internal evals reportedly show further gains on top of that.
The Cursor-xAI data flywheel is the real story behind 4.5's coding performance. Cursor trained Composer 2 and 2.5 (both based on Kimi K2.5 continued pre-training + Claude Code-style RL) on xAI compute, and Grok 4.5 was trained alongside Cursor workloads during the same period. That shared engineering code corpus is something 4.5's competitors don't have access to, and it's visible in the SWE-Bench Pro / Terminal-Bench style scores.
What a one-month cadence actually means
If 4.6 ships on August 7, the gap between 4.5 going public (July 8) and 4.6 (August 7) is exactly one month. That's a weekly-cadence update rhythm — territory DeepSeek and Moonshot have occupied in China, but essentially unheard of for a US frontier lab (which has historically run 3–6 month cycles).
Counting the last 90 days for xAI: Grok 4.3 (early May) → Grok 4.5 (early July) → Grok 4.6 (expected early August) → Grok 4.7 (expected before end of August). Roughly 30 days between major releases, with parameter count and training-data breadth ratcheting up each step. The 4.7 announcement also confirms that SpaceX engineering data will be used to train Grok (excluding ITAR-restricted material), which is a structurally novel data source that no other frontier lab has access to.
My take
Walking 4.6 back to 1.5T instead of just shipping 2T is a meaningful decision.
If 4.6 were 2T and 4.7 were 2.1T, you'd have two models separated by 0.1T parameters and a few weeks of post-training — 4.7 wouldn't have much independent reason to exist. Pulling 4.6 back to 1.5T and saving the 2.1T jump for 4.7 carves out a real 40% capacity gap and several weeks of SFT/RL work. That makes 4.7 a distinct new model rather than a numbered nudge.
Reading between the lines of the July 8 4.5 launch ("Beta feedback has been very positive"), xAI evidently believes 4.5's capability surface is still under-exploited. A 1.5T 4.6 with heavier RL can push benchmarks further without burning another full pre-training run. If 4.6 actually moves SWE-Bench Pro and Terminal-Bench meaningfully past 4.5, then 4.7 has a clean narrative for the 2.1T doubling.
The risk is framing. 4.6's 1.5T pre-training is reportedly already done and is now in SFT/RL. When a "new model" is announced after pre-training is complete, the honest read is that 4.6 is a heavily-tuned 4.5 with a marketing version bump. Whether it actually moves capability net-new, or just consolidates 4.5's gains, is something we'll only know on August 7 from the benchmark table.
For users: If you're running Grok 4.5 in production coding or agentic workflows, treat 4.6 with the standard "new model" discipline — small canary traffic first, compare against your 4.5 baseline, then decide whether to upgrade. Then wait for 4.7 at 2.1T before committing to that as your long-term default. xAI's weekly cadence is not friendly to people who want stable evals — every upgrade means re-running your benchmark suite, re-tuning prompts, and re-migrating caches.
For the industry: xAI + Cursor proved over the last 30 days that a smaller lab can run frontier-model release velocity. If 4.6 and 4.7 both ship on schedule, the rest of the year for OpenAI / Anthropic / Google is going to feel even more compressed than it already does. The quarterly-update norm among frontier labs is now under real pressure.
So what
Three things to watch over the next two weeks:
- Does 4.6 actually ship on August 7, or does Musk slip the date again (Musk's track record on timelines is, charitably, mixed).
- What does 4.6's SFT/RL upgrade actually move on SWE-Bench Pro and Terminal-Bench versus 4.5 — the size of that gap is the real signal on whether 4.6 is a true new model or a tuned 4.5.
- Whether 4.7's 2.1T training finishes before end of August, or whether it gets pre-empted by a Kimi K3, Claude Opus 5, or Gemini 3 release that grabs the cycle's headline.
If all three hold, xAI finishes Q3 2026 in a defensible "top three" position on raw frontier-model velocity.
Sources: ITHome (July 30, 2026), NextBigFuture (July 24, 2026), 36Kr, xAI's official X account (July 24–28, 2026).