During pre-training, DYNA-2 used a video prediction objective — learning to predict future video frames from past observations — combined with a video co-training algorithm that jointly learns visual representations and action prediction . This approach contrasts with the VLA architecture of DYNA-1, which directly maps perception to robot actions without intermediate video prediction
.
The company demonstrated scaling across four orders of magnitude, from 1,000 hours up to 1,000,000 hours of pre-training data .
The company ran direct head-to-head comparisons between DYNA-2 and DYNA-1 under matched training conditions and real-world deployments.
Mean normalized performance rose monotonically with pre-training scale: 20% → 28% → 45% → 53% of the attainable maximum as data scaled from 1k → 10k → 100k → 1M hours . At the 1M-hour scale, DYNA-2 achieved the strongest real-world performance on 9 of 14 tasks
. Key task-level improvements included:
Dyna Robotics reported what it calls "the first true scaling law in robotics powered entirely by human data" . Key characteristics:
The company explicitly states that this is a "first-of-its-kind" demonstration that scaling human observational data, not robot action data, can serve as a reliable axis for improving physical robot intelligence .
If the claims hold up to independent verification, Dyna Robotics has shown a practical path to scaling robot intelligence without the bottleneck of collecting expensive robot action data. Human video exists at effectively unlimited scale. The company says it is already working toward 10 million hours of training data, which would enable robots to master new physical tasks with just hours of local fine-tuning .