As the entire AI industry bets on language models, New York-based startup Runway has chosen another path — building intelligence from video data rather than text.
Runway co-founder Anastasis Germanidis offers a provocative take: "Language models are trained on the internet, forums, social media, and textbooks — distilling existing human knowledge. But to surpass them, we need to leverage data with fewer biases." This statement underpins a world-model ambition: observing the world operate directly through video, rather than indirectly describing it through human language.
Runway was founded in 2018, with all three co-founders from NYU's Tisch School of the Arts, not the typical Silicon Valley background. They started in video generation, and Gen-4.5 is their fourth-generation model, serving top-tier clients like Lionsgate and AMC Networks, and has been used in productions like "Everything Everywhere All at Once." The company is valued at $5.3 billion, with $40 million in new annual recurring revenue added in Q2 2026.
But the real bet is the architectural choice. While GPT-5.5 and Claude Opus 4.7 chase each other on language benchmarks, Runway believes the next generation of intelligence will come from video, not text. If this judgment is correct, Google won't be the only opponent that needs to worry — the entire language-based AI route could face reevaluation.
This isn't without risks. Google has deeper pockets and more data reserves. But Germanidis believes that precisely because Runway isn't from Google or Meta, it's possible to make a different bet. The second half of the AI competition may not be just about parameter counts and benchmark scores, but a fundamental divide over "what data intelligence is built upon."