U.S. robotics company Dyna Robotics has unveiled DYNA-2, a robot-based model trained on about 1 million hours of first-person human video without robot data. It learns from footage of people handling real objects to build robots’ physical action capabilities, an approach aimed at solving a shortage of training data seen as a key hurdle in developing general-purpose robots.
Cryptopolitan, a blockchain media outlet, reported on Aug. 11 (local time) that Dyna Robotics unveiled DYNA-2 pre-trained using only first-person human video.
In the development of general-purpose robots, a widely used method has been to collect data by having people directly teleoperate robots. People control movements of a robot arm or a humanoid and repeat specific tasks to secure training data.
But the process requires substantial time and cost. There are also limits to large-scale expansion because new data must be collected again as task types increase or robot forms change.
Dyna Robotics approached the problem in the opposite way. It excludes robots entirely during training and extracts action information robots need from videos showing people interacting with objects.
DYNA-2 trained on about 1 million hours of first-person human video. The company described that as a scale equivalent to about 170 years of continuous human behavioral experience. Dyna Robotics co-founder Jason Ma (제이슨 마) said, "Action data is scarce, but video is everywhere," stressing that physical intuition can be learned from human video. He said robots can learn information needed for physical interaction from scenes of people grasping and moving objects, without directly operating robot arms for millions of hours.
The company also presented performance gains. Dyna Robotics said success rates in precision manufacturing tasks rose from about 20 percent to 80 to 90 percent. In an evaluation across 15 benchmark tasks, the company said a model trained on more human video consistently performed better than one trained on less.
The model’s structure also differs from existing robot AI. DYNA-2 is designed as a video-generation-based World-Action Model rather than a typical vision-language model. In pre-training, it does more than understand the situation in video, predicting the next video frame and the next action at the same time.
Dyna Robotics said this training method helps robots internally represent physical changes and spatial relationships that occur when they make contact with objects.
The company ultimately aims for generality that is not tied to a specific robot. It said knowledge learned from human video can be transferred to a fixed robotic arm, a humanoid prototype and a multi-jointed robotic hand that were not directly used in DYNA-2 pre-training.
The company said it can also significantly shorten fine-tuning for new robot platforms. It aims to reduce tasks that took weeks to a matter of hours. As an example, it cited a case in which a robotic hand learned a bottle-cap twisting action with only 13 minutes of additional data.
Dyna Robotics also stressed that it is not staying only at the research stage and is deploying robots in industrial settings. The company’s DYNA-1 robot is known to be used in various environments, including hotels, restaurants, laundries and gyms. The company was co-founded by Lyndon Gao (린든 가오), York Yang (요크 양) and Jason Ma, who previously worked at Google DeepMind.
It plans to expand the scale of human video training from the current 1 million hours to 10 million hours. Dyna Robotics sees the bottleneck for general-purpose robots in securing data rather than in spreading robot devices themselves. It judges that training can be scaled up much faster by using human action videos that already abound online and in the real world, instead of deploying thousands or tens of thousands of expensive robots to collect data directly.
Still, it remains to be verified how reliably action knowledge learned from human video can be reproduced in real robot environments. That is because people and robots differ in body structure, joints, strength and field of view.
Ultimately, the success of DYNA-2 is expected to depend less on how much human video it learns and more on how reliably it can transfer the physical intuition gained in the process to different robots and real work environments. Attention is focused on whether the 10 million-hour video training pursued by Dyna Robotics can take hold as a new way to secure data for general-purpose robot development.
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws: • world-action models exhibit scaling law on human data across four orders of magnitude, from 1000… pic.twitter.com/wZamR0axzS