Google Veo, Runway AI Pitched as World Models for Next-Gen Robotics

The world’s most popular AI video-generation platforms are racing into robotics, reframing their creative tools as rehearsal spaces for next-generation machines.

Google DeepMind and Runway just made near back-to-back announcements, claiming breakthroughs in video-generation that could accelerate the push toward general-purpose robotics. Another popular platform, Luma AI, recently signaled its video models are being developed as large-scale world models for downstream use cases like training robots in simulation.

They’re entering a race toward so-called world models, which are AI systems that can stand in for reality itself. It’s a crowding field with well-funded players like World Labs, NVIDIA, OpenAI and Tencent. Billions of dollars are flowing into the emerging field as consensus shifts toward large language models (LLMs) being insufficient as enabling layer in embodied AI.

“Language is fundamentally a purely generated signal. You don’t go out in nature and there’s words written in the sky for you,” World Labs founder Fei-Fei Li said recently on The A167 Show podcast.

They’re basically stepping on NVIDIA‘s turf, similar to how the GPU maker expanded beyond graphics and into AI computing with its CUDA platform in the mid-2000s. They’re taking systems built to render video and retrofitting them as infrastructure for simulating reality itself. Since their popularity took off a few years ago, the developers say their models have gotten increasingly uncanny in their ability to model the real world.

World Models for Robots?

The idea of “world models” has been around for decades but its modern usage is usually traced to a 2018 paper by David Ha and Jürgen Schmidhuber that showed neural networks learning internal simulations of environments directly from visual data. Its emergence as a serious foundation for robotics and physical AI is very recent. There’s no clear leader so far, and the technology could fragment into several approaches like video-based, physics-first, and hybrid systems.

Video-based models learn by watching enormous amounts of footage and predicting what the next moments will look like. The technology has advanced dramatically since it first emerged in the 2010s, when academic AI models could only generate short, blurry clips with jittery motion and little sense of cause and effect. Researchers describe advanced video-generation systems as having world-model-like behavior because they learn underlying dynamics and can imagine future scenes conditioned on actions.

Raw Video for World Models

Video-based models learn to predict outcomes by watching enormous amounts of footage. Raw video is becoming a hot commodity for robotics firms, and entrepreneurs are going to unusual lengths to get it.

Eighteen-year-old Eddy Xu, for example, convinced thousands of factory workers to wear head mounted cameras during their daily jobs. Using Shenzhen as his base, he worked directly with factory networks across Southeast Asia, producing a massive egocentric dataset with more than 100,000 hours of footage, nearly 11 billion frames, and over 2 million clips.

Eddy Xu, founder of Build AI
Eddy Xu, founder of Build AI, convinced thousands of factory workers in southeast Asia to wear headsets to capture video of them working to train AI models (Source: @eddybuild/X)

It’s the exact kind of data developers are seeking for next-gen video generation, world models, and robotics systems, but it’s small compared to what Google has access to since it owns YouTube.

Google Veo as a Robotics Simulator

Google DeepMind, Alphabet’s AI research wing, explicitly reframes its Veo video-generation model as a robotics evaluation simulator that can predict how robot policies behave, and misbehave, before they’re deployed on real hardware.

“Veo is a surprisingly strong world simulator,” Thomas Kipf, a DeepMind senior staff researcher, said on X. “We fine-tuned Veo on action-conditioned, multi-view robotics data. Key result: running a policy in the world model is strongly correlated with real-world results.”

DeepMind’s new paper, titled “Evaluating Gemini Robotics Policies in a Veo World Simulator,” details how Veo can simulate robot actions, identify specific unsafe outcomes, and then see those same failures occur when the scenarios are recreated with real hardware. The simulations are detailed enough to expose safety failures like a robot colliding with a human hand or damaging an object due to poor situational judgment.

Google Veo robotics test
Veo’s robotics world model predicts a dangerous grasp. Right: the same unsafe behavior occurs during a real robot test.

The researchers stress that the failures are not hypothetical. After Veo flagged unsafe behaviors during test simulations, they rebuilt the same scenes physically and observed the same dangerous outcomes. The team says the approach surfaced edge cases that would be too impractical and hazardous to discover through traditional hardware testing. According to the paper, the system also effectively ranked robot policies and predicted how performance would degrade when conditions changed. Veo’s predictions aligned closely with real-world results seen across more than 1,600 robot trials.

“Ultimately, this work demonstrates the massive impact of video models in robotics,” the team said in its conclusion. “The ability to evaluate robots in an infinitely rich and varied proxy of the world provides necessary infrastructure for developing generalist embodied agents that operate usefully, capably, and safely in real-world environments.”

The next logical step for DeepMind is a tighter integration within its Gemini Robotics models, which are already being used by leading robotics firms like Apptronik, Agile Robots, Agility Robotics, Boston Dynamics, and Enchanted Tools for next-gen industrial, logistics, and service applications.

Runway’s General World Models

Runway, one of the most widely used AI video-generation services, is also moving beyond filmmaking with the release of its first general-purpose world model.

The New York-based startup, launched in 2018 and valued at $3 billion, just launched a GWM-1, a new family of models designed to generate and maintain interactive, persistent environments rather than single-shot video clips. GWM-1 differs from traditional video generators by allowing virtual agents to act inside the simulated worlds and observe consequences over time.

Runway GWM-1 robotics model
A synthetic robot view generated by Runway’s GWM Robotics model, simulating object interaction in a virtual environment.

According to Runway, GWM-1 includes three variants:

  • GWM-Worlds creates explorable environments with an understanding of geometry, lighting, and basic physics.
  • GWM-Robotics generates synthetic data and simulated scenarios for training and testing robots.
  • GWM-Avatars creates interactive digital characters.

Runway paired the announcement with an update to its flagship video model, which now supports native audio and longer, more complex sequences. The startup is offering a Python SDK for its robotics world model API that supports multi-view video generation and long-context sequences. Runway says its interface is designed for “seamless integration into modern robotic policy models.”

OpenAI’s Sora as a World Model?

OpenAI has not publicly positioned its Sora video-generation model as a robotics system, though the nonprofit-turned-tech-giant is widely reported as pushing deeper into hardware. Sora has been presented as an AI video model that learns space, motion, and cause and effect from large amounts of data. Since those abilities are crucial for robotics, researchers view Sora as an implicit world model with potential applications in simulation and embodied AI.

World Labs Ships Commercial Model

While video models learn how scenes tend to change visually, Palo Alto-based World Labs is trying to extract what the world actually looks like in three dimensions. The unicorn startup, launched by ‘Godmother of AI’ Fei-Fei Li in 2024, relies heavily on computer vision, 3D reconstruction, and spatial reasoning models rather than video diffusion.

World Labs has commercialized its first world model, called Marble, which was launched in November 2025 after a limited beta. It’s now available in free and paid tiers for users to generate, edit, and export 3D environments from text, images, video, and 3D layouts. According to World Labs, the editable 3D worlds can be used in video games, simulations, design, and other workflows.

The Future of World Models + Robotics

Nobody knows where this technology is headed or whether enthusiasm will continue into 2026. To stay up to date on whatever it manifests into, make sure to subscribe.