If last year's Google Veo 3 made people exclaim that AI video was nearly indistinguishable from reality, then this year's launch of Gemini Omni Flash pushes that "nearly" to a different extreme — so close it makes your skin crawl. Omni Flash is the first officially released multimodal-generation model of Google's Omni family, now live on Google AI's video platform Flow. Three improvements over the previous Veo stand out: multimodal input can take both video and text prompts as the starting point; real-world knowledge fusion helps the model maintain object consistency and scene logic over long videos; and instruction-aware editing lets users request changes in natural language, with the model genuinely incorporating them into the result.
The most striking finding from reporters' hands-on tests: the personal deepfake videos generated by Omni Flash are basically indistinguishable to ordinary people. In a "selfie video eating spaghetti," a husband — given no prior warning — fully believed it was real; the only tell was that "the bowl looked unfamiliar." A clip in front of the Eiffel Tower, while slightly cartoony, was nearly impossible to identify as AI-generated on its own. This suggests the "realism" bottleneck of video generation is shifting from the technical to the psychological.
But Omni is far from perfect — object-property drift, broken physical consistency, and unstable instruction-following all show that current multimodal-generation models still struggle with temporal consistency and precise instruction execution; they're not yet reliable creative tools. The race to deploy SynthID-style watermarking tech is more urgent than ever. Beyond the tech showcase, the question worth asking is: are we ready to live in a world where "seeing is no longer believing"?