1X: NEO Starts Learning on Own with Video Data

NEO the humanoid robot can now see into the future and manifest it.

That’s according to 1X Technologies, the Palo Alto-based startup that’s preparing to begin shipping its $20,000 humanoid robot for households. Though 1X is working toward full autonomy, it will rely on teleoperators for initial deployments when NEO encounters actions it can’t handle on its own. The OpenAI-backed startup says it could go full auto by 2027 with help from early adoption data.

As skeptics question whether Jetsons like living is within reach, 1X just announced what it describes as a major breakthrough in robotic intelligence. 1X says its inhouse world model, under development since 2024, lets Neo perform tasks it’s never been explicitly trained on.

1X NEO humanoid robot world model
Still from 1X Technologies’ marketing video on its World Model progress (Source: 1X Technologies)

The 1X World Model is an AI system that acts as a stand-in for physical reality. It’s trained using hundreds of hours of first-person human video to understand how people get things done.

The underlying video data comes from three main sources. The core model is pre trained on large web scale datasets, similar to those used for generative video models. The footage shows people and objects interacting in everyday situations. Next the model is trained on egocentric human video of people performing tasks. Finally, the model is finetuned using real video and motion data from Neo itself, which teaches the system how those behaviors map to the robot’s body and physical limits.
which teaches the system how those behaviors map to the robot’s body and physical limits.

Neo basically generates short video clips as it decides what to do next. Using the system, Neo had the most success with simple well defined object interactions. The success rate dropped as tasks became more physically complex and contact rich.
Experimenting with generating multiple futures and picking the best one improved results but also computing requirements.

In tests, the imagined videos produced by the model closely matched what wound up happening in the real world. Using the system, NEO had the most success with simple well defined object interactions. It had a close to 95 percent success rate steaming a shirt, 80 percent grabbing chips, and 75 percent opening a sliding door. NEO struggled most with drawing a smiley face and pouring cereal.

1X researchers also found that when the generated video looks physically correct, the robot is much more likely to succeed. In cases where the video prediction is clearly wrong, real-world execution usually failed. This led to the team experimenting with generating multiple possible futures and choosing the best one. The system at times misjudges depth or contact dynamics because it relies on single camera perspectives.

1X said the key takeaway is that Neo improves best by learning from experience. The startup says NEO is designed for safety in households with its soft exterior, pinch proof joints, near silent noise level, and its light build of 30 kg (66 lbs).