
The latest paper from the $5.6B embodied AI startup Physical Intelligence is likely to intensify the growing race to collect raw video data to train next-generation robots to take over human jobs.
The San Francisco startup says its research shows that robots can learn new tasks from first-person human demonstration videos, but with an important caveat. According to the researchers, the approach only works once the robot’s underlying artificial intelligence has been trained on enough diverse real-world experiences. Once that critical mass is met, the researchers say, videos of humans become a powerful training resource rather than background noise.
Physical Intelligence detailed its findings in the paper, “Emergence of Human to Robot Transfer in Vision-Language-Action Models.” The study was conducted in collaboration with the Georgia Institute of Technology.
“Much like large language models, larger VLAs may not only improve performance but also unlock entirely new capabilities,” the researchers wrote. “Using human video may be just one such capability, and it’s exciting to imagine what others might emerge as we continue to scale up robotic foundation models.”
Launched in 2024, Physical Intelligence has raised more than $1 billion to date to develop foundational models for robots across form factors to perform complex, contact-rich tasks without custom programming. Its software has demonstrated early success automating chores long considered too unstructured or unpredictable for modern robots, including folding laundry, box assembly, kitchen tasks, and industrial manipulation. Alphabet, Google’s parent company, led the startup’s recent $600 million Series B funding round.

To conduct their new study, researchers collected egocentric human demonstration data by having people wear head-mounted cameras and, in some cases, cameras strapped to each wrist while performing everyday manipulation tasks like tidying up, bussing tables, and sorting objects. The actions were recorded as short, task-focused video clips, similar in structure to robot teleoperation data. The team then used visual SLAM to label each video with what the person was trying to do.
The researchers stressed that they didn’t try to match human joints to robot joints. Instead, they mixed the annotated human video data before feeding it into Physical Intelligence’s π0.5 vision-language-action (VLA) model. They tested the system on specific challenges in unfamiliar homes.
In the real-world tests, the researchers noted a clear, measurable difference once the robot’s AI was experienced enough. They used an ARX mobile manipulator, focusing on tasks the robot had never been trained to do on its own. In one instance, the robot’s original training only covered placing eggs into cartons, but human videos show how to sort by them by color. Without human data, the robot sorted correctly 57 percent of the time; after adding human videos, accuracy jumped to 78 percent.

In household cleanup tasks, success rates nearly doubled. Tidying a spice rack improved from 32 to 71 percent, while cleaning a dresser surface rose from 25 to 50 percent. For a more complex table-bussing job, performance improved a modest 10 points, from 53 to 63 percent. The researchers said the key pattern across tests was that human video only helped once the model had enough prior knowledge.
Physical Intelligence consistently argues that learning grounded in real-life physical interaction matters more than relying on simulation training and synthetic data. Its robot brain is trained on large amounts of teleoperation data, where human operators control robots as they move, grasp objects, collide with their environment, succeed, and fail in the real world. Looking ahead, the researchers said the next step is fueling its VLA with broader real-world experiences, not just more of the same data.
As raw video becomes a hot commodity for robotics, companies are going to unusual lengths to obtain it. Some startups, like Mountain View-based Sunday Robotics, are paying people to record themselves do chores while wearing custom gear. Another Silicon Valley startup, Figure AI, has partnered with real estate giant Brookfield to use its massive portfolio of residential, commercial and logistics properties to collect real-world data for training its humanoid robots. A newer startup, Build AI, has secured $15 million in funding to scale production and deployment of its devices used to collect 100,000 hours data from real factory workers.

Popular AI video generation services like Google Veo and Runway are also exploring use cases for robot training. Runway recently shared synthetic video examples created using its world model for embodied AI, GWM Robotics, integrated with Physical Intelligence’s flagship model. Others, like the National University of Singapore’s Show Lab, are coming up with ways to make existing footage more useful for robot training using open-source video generation tools like Wan 2.2.