On June 28, Elon Musk announced via X that xAI's newest large language model, Grok 4.5, has begun a beta test inside SpaceX and Tesla — the first time the Grok series has run an internal test in a true enterprise-grade production environment.
The model itself: built on a V9 base of 1.5 trillion parameters, 50% larger than the 1T Grok 4.4 from late May; supplementary training mixes in Cursor developer-workflow data, baking real IDE usage into alignment; RL is still being optimized, and Musk claims every round of RL significantly improves capability.
Internal early evaluation shows "close to, possibly exceeding Claude Opus" (pointing at Opus 4.6) — the first time Grok has publicly benchmarked itself against a frontier closed model. But there are still no independent external benchmarks, and the conclusion remains a vendor-side claim.
Musk also announced: SpaceX will release a brand-new foundation model, trained completely from scratch, every month for the rest of 2026 — versus the 3-6 month major-release cadence of peers. This is xAI's most aggressive translation of Colossus-cluster compute directly into delivery cadence.
Commentary: putting a 1.5T model in production beta at SpaceX and Tesla is, in effect, running a massive agent evaluation with two world-class engineering teams. The Cursor-data inclusion is notable: once IDE workflow enters training, the model goes from "knows how to write code" to "writes code the way a Cursor user does" — SaaS-tool vendors are pulled from the "application layer" into the "data provider" layer. Whether the monthly from-scratch-training promise holds depends on Colossus capacity and training-pipeline maturity.