The shift to video-based training addresses the primary constraint in generalist robotics: the difficulty of acquiring high-quality physical action data. While previous iterations, such as the DYNA-1 model, relied on Vision-Language-Action architectures, DYNA-2 operates as a World-Action Model. This system predicts physical motion and contact physics by processing 170 years of human waking experience, enabling cross-embodiment transfer between stationary arms, humanoids, and dexterous hands.
Performance metrics from customer deployments highlight a significant leap in capability. In high-precision manufacturing, the model boosted task success rates from 20% to nearly 90% without requiring post-training adjustments. During head-to-head physical evaluations, DYNA-2 completed tasks 1.55 times more frequently than its predecessor, demonstrating an 87% pass rate compared to the 46% achieved by the VLA baseline. The model also shows resilience in real-world environments, recovering from physical disturbances during dexterous tasks like food preparation without human intervention.




Comments (0)
No comments yet. Be the first!