New Physical Intelligence Study Intensifies Race for Raw Video for Robot Brains

Physical Intelligence human-to-robot studdy
A researcher wears cameras mounted to his head and wrists to collect manipulation data (Source: Physical Intelligence)

The latest paper from the $5.6B embodied AI startup Physical Intelligence is likely to intensify the growing race to collect raw video data to train next-generation robots to take over human jobs.

The San Francisco startup says its research shows that robots can learn new tasks from first-person human demonstration videos, but with an important caveat. According to the researchers, the approach only works once the robot’s underlying artificial intelligence has been trained on enough diverse real-world experiences. Once that critical mass is met, the researchers say, videos of humans become a powerful training resource rather than background noise.

Physical Intelligence detailed its findings in the paper, “Emergence of Human to Robot Transfer in Vision-Language-Action Models.” The study was conducted in collaboration with the Georgia Institute of Technology.

“Much like large language models, larger VLAs may not only improve performance but also unlock entirely new capabilities,” the researchers wrote. “Using human video may be just one such capability, and it’s exciting to imagine what others might emerge as we continue to scale up robotic foundation models.”

Launched in 2024, Physical Intelligence has raised more than $1 billion to date to develop foundational models for robots across form factors to perform complex, contact-rich tasks without custom programming. Its software has demonstrated early success automating chores long considered too unstructured or unpredictable for modern robots, including folding laundry, box assembly, kitchen tasks, and industrial manipulation. Alphabet, Google’s parent company, led the startup’s recent $600 million Series B funding round.

Physical Intelligence video collection illustration
Physical Intelligence researchers used cameras mounted to the head and wrist to collect raw video of manipulation tasks (Source: Physical Intelligence)

To conduct their new study, researchers collected egocentric human demonstration data by having people wear head-mounted cameras and, in some cases, cameras strapped to each wrist while performing everyday manipulation tasks like tidying up, bussing tables, and sorting objects. The actions were recorded as short, task-focused video clips, similar in structure to robot teleoperation data. The team then used visual SLAM to label each video with what the person was trying to do.

The researchers stressed that they didn’t try to match human joints to robot joints. Instead, they mixed the annotated human video data before feeding it into Physical Intelligence’s π0.5 vision-language-action (VLA) model. They tested the system on specific challenges in unfamiliar homes.

In the real-world tests, the researchers noted a clear, measurable difference once the robot’s AI was experienced enough. They used an ARX mobile manipulator, focusing on tasks the robot had never been trained to do on its own. In one instance, the robot’s original training only covered placing eggs into cartons, but human videos show how to sort by them by color. Without human data, the robot sorted correctly 57 percent of the time; after adding human videos, accuracy jumped to 78 percent.

Physical Intelligence human-to-robot transfer on generalization tasks
Comparison of human-to-robot transfer on generalization tasks with and without human video data (Source: Physical Intelligence)

In household cleanup tasks, success rates nearly doubled. Tidying a spice rack improved from 32 to 71 percent, while cleaning a dresser surface rose from 25 to 50 percent. For a more complex table-bussing job, performance improved a modest 10 points, from 53 to 63 percent. The researchers said the key pattern across tests was that human video only helped once the model had enough prior knowledge.

Physical Intelligence consistently argues that learning grounded in real-life physical interaction matters more than relying on simulation training and synthetic data. Its robot brain is trained on large amounts of teleoperation data, where human operators control robots as they move, grasp objects, collide with their environment, succeed, and fail in the real world. Looking ahead, the researchers said the next step is fueling its VLA with broader real-world experiences, not just more of the same data.

As raw video becomes a hot commodity for robotics, companies are going to unusual lengths to obtain it. Some startups, like Mountain View-based Sunday Robotics, are paying people to record themselves do chores while wearing custom gear. Another Silicon Valley startup, Figure AI, has partnered with real estate giant Brookfield to use its massive portfolio of residential, commercial and logistics properties to collect real-world data for training its humanoid robots. A newer startup, Build AI, has secured $15 million in funding to scale production and deployment of its devices used to collect 100,000 hours data from real factory workers.

Factory workers wearing Build AI headsets for data collection
Factory workers wear Build AI headsets for data collection (Source: @eddybuild/X)

Popular AI video generation services like Google Veo and Runway are also exploring use cases for robot training. Runway recently shared synthetic video examples created using its world model for embodied AI, GWM Robotics, integrated with Physical Intelligence’s flagship model. Others, like the National University of Singapore’s Show Lab, are coming up with ways to make existing footage more useful for robot training using open-source video generation tools like Wan 2.2.