On September 24, Google DeepMind chief Koray Kavukcuoglu confirmed at an event hosted by The Information that the next flagship model, Gemini 4, has entered the early "post-training" phase of development, and that Google hopes to release the first version "well before" the end of the year (CLS, Sina Finance).
What post-training means
Post-training is the stage where a base model, after large-scale pre-training, is further optimized so that its reliability, behavioral consistency, and real-task performance reach a shippable level. According to Sina Finance, Google has already seen internal test results and plans to ship an early post-training build rather than waiting until every optimization is done, then iterate rapidly afterwards. In other words, the Gemini 4 release strategy is closer to "ship a usable version first, keep upgrading", shortening the gap between training completion and launch.
Skipping Gemini 3.5 Pro
Google had originally planned to launch Gemini 3.5 Pro in June — CEO Sundar Pichai announced that timeline at I/O in May, but the model never appeared. Kavukcuoglu explained that after Gemini 3 and 3.1, Google chose to "take a step back" and focus more on the Flash family to speed up learning and iteration. He did not say whether 3.5 Pro was officially cancelled, but confirmed the focus has now shifted to Gemini 4.
Already on internal duty
Gemini 4 is not just sitting in lab benchmarks. According to Sina Finance, Google is developing safety guardrails and running safety tests for the model, and engineers already use it to run the internal Antigravity AI coding tool. Putting the model into a real engineering environment early exposes problems in complex coding tasks, tool use, and long-horizon execution before the public release.
Catch-up context and the TPU loop
The acceleration comes amid renewed frontier competition: Anthropic launched its new Mythos model family this spring and kept upgrading it this month, while OpenAI began rolling out its GPT-6 series in early September; reports suggest both companies' latest flagships already surpass Google's strongest current models in complex reasoning. Asked whether Google is behind, Kavukcuoglu said he retains full confidence in the DeepMind team.
The other easily missed dimension is hardware. Sina Finance reports that Google started selling TPUs directly to customers this year, and that DeepMind's model roadmap feeds the hardware team early signals about future compute needs, letting it plan the next two to three TPU generations accordingly. Models define chips, chips feed back into models — a structural loop that pure model labs do not have.
How to read it
Three takeaways:
- Skipping a generation is a resource-concentration signal. Dropping 3.5 Pro and betting directly on Gemini 4 means the flagship line aims at generational leaps while the Flash family handles high-frequency iteration — a cleaner division of labor than "saving up for one big release".
- "Usable first, iterate fast" compresses evaluation windows. Once early post-training builds go public, downstream teams face a continuously rolling model state rather than one-off version comparisons; procurement and benchmarking cadence has to keep up.
- AGI narrative gives way to trustworthy agents. Kavukcuoglu downplayed the "has AGI been achieved" debate and emphasized building agents you can actually trust — consistent with the industry shift from leaderboard chasing to reliable long-horizon task execution.
So what: if your technology shortlist still has a slot for "waiting on Gemini 3.5 Pro", rewrite it as "evaluate an early Gemini 4 before year-end, then track rapid iterations" — Google itself has already removed the mid-cycle version from its own roadmap.