How Can Video Generation Lead to Robots Interacting With The Physical World?

Reese Watson - Author
By

Updated Aug. 22 2026, 8:19 p.m. ET

Diego Marti Monso
Source: Ho Tin Fan

Video Generation Leading to Robots?

Diego Marti Monso is one of the inventors of the algorithms that power the AI models generating the videos we see all over social media. What many people perceive as mere AI Slop might actually be the key to developing new AI models capable of interacting with the physical world.

Article continues below advertisement

Artificial general intelligence (AGI) refers to the idea of having AI models that can solve any general task, just like humans do. Although AI has evolved dramatically in recent years and has become a standard tool for most white-collar workers, it still falls short of human flexibility. For example, you still can’t rely on AI to manage a business or perform blue-collar work. There are not yet currently generally useful robots that humans can use.

The term "AGI" is often limited to the digital domain, but Marti Monso argues that giving AI an embodiment to interact with the physical world is also a requirement for true AGI. As one of the most impactful researchers of video and world models in recent years, Marti Monso explains how video generation methodologies have recently evolved and how they could unlock the next frontier of physical AI.

Article continues below advertisement
Diego Marti Monso
Source: Ho Tin Fan

Video Generation Leading to Robots?

How AI Videos Became "Good"

We’ve all seen AI-generated videos on social media: fruit talking, world leaders dancing together, and even dogs being interviewed by babies. In the early days of AI content, these videos were mostly generated using models such as OpenAI’s Sora 1, which was announced in February 2024, and Veo 1, which was announced by Google DeepMind at Google I/O in May 2024. However, these early models had a severe limitation: they could only generate short bursts of fixed-length frames from a single text prompt. This is why, until recently, AI-generated videos always consisted of short clips stitched together, making the videos look unrealistic and inconsistent over long time horizons. It was only very recently that video generation capabilities leapt and became commonplace on the internet.

Article continues below advertisement

The Infinite-length Breakthrough

In July 2024, Marti Monso and his collaborators at the MIT Computer Science and Artificial Intelligence Laboratory published Diffusion Forcing, marking a major turning point in video generative modeling. Marti Monso introduced a new training paradigm that enabled video models to generate videos of any length, effectively removing the constraint.

Over the following months, DeepMind released Veo 2 in December 2024 and Veo 3 in May 2025. OpenAI followed with Sora 2 in September 2025. These are their flagship models that can generate videos of any length and serve millions of users worldwide. Other open-source releases soon followed, such as Sand AI’s MAGI-1 and Skywork AI’s Skyreels v2, which were trained directly with Diffusion Forcing. The entertainment industry has quickly adopted long video generation to produce films, advertisements, and social media content. There are even video games built on this technology.

Article continues below advertisement

From Video Generation To World Modeling

The infinite-length breakthrough sparked by Marti Monso also sparked a revolution in the field of world modeling. A world model is a video model that can respond to inputs, such as video game controls or robot commands, turning the video model into a simulator of the world. With his latest release, Open Dreamer, Marti Monso and his collaborators demonstrate how DeepMind’s frontier world model with diffusion forcing (Dreamer 4) works. The model is trained on recordings of humans playing Minecraft, and it learns a playable real-time simulator of the game (“dream”).

How World Models Could Reach the Real World

Marti Monso believes that vision is the most important sense for intelligence. One of the core mechanisms of human and animal learning is incorporating life experiences into decision-making processes. We mostly perceive experiences by observing the results of our interactions with the world — just how world models are trained. If you could take a world model and simply train it in the physical world, not the digital world of Minecraft, you would effectively have a reactive intelligence that could adapt and exist in the real world.

Article continues below advertisement
Diego Marti Monso
Source: Ho Tin Fan

Video Generation Leading to Robots?

Video Models Are Better Suited Than Language Models

Today, consumers are most familiar with Large Language Models (LLMs). LLMs such as ChatGPT are trained using a vast amount of human knowledge written down throughout history. This includes everything from the earliest records of classical literature to the latest scientific publications and codebases. These AIs learn facts and patterns of intelligence from the text they see during training. This results in the AI models we use every day. However, we are already reaching the limits of available text data with which to continue feeding the expensive scaling laws that drive the frontier of intelligence.

Article continues below advertisement

"Text is actually quite expensive to produce," says Marti Monso. "Imagine how long it would take you to write a full description of a day in your life without skipping any detail. Now, consider how much easier and richer it would be to use a camera to record your day instead.” Furthermore, videos can pack physical information much more densely than text can. "As the saying goes, a picture is worth a thousand words." This architecture is much more well-suited for the real world - it's simply much easier to anticipate a baseball based on watching its arc than it is to try to to write out its trajectory in text.

The Future

The AI videos that people scroll past are a byproduct of AI learning how the physical world works. As researchers like Marti Monso continue to improve the algorithms that generate these videos, we are getting closer to a world in which embodied AI inhabits the same spaces as humans and becomes a helpful assistant in the physical world. Little do most people know, these two things are deeply, inextricably connected.

Advertisement

Latest Business News News and Updates

    © Copyright 2026 Engrost, Inc. Distractify is a registered trademark. All Rights Reserved. People may receive compensation for some links to products and services on this website. Offers may be subject to change without notice.