Developed by the University of Hong Kong’s MMLab alongside nearly 20 international academic institutions, RoboDojo pushes robots beyond simple demonstration-based training. VPP2 distinguishes itself from standard video generation models by focusing on object dynamics and the translation of high-level instructions into precise physical movements. By integrating video prediction with action generation, the system allows robots to navigate unfamiliar surroundings and complex manipulation tasks with increased reliability.
Performance metrics highlight the model’s versatility. On the ALOHA platform, VPP2 outperformed existing baselines in nine out of ten task categories, achieving a 58.5% success rate. The integration of a vision-language model for high-level task planning further improved performance, more than doubling success rates on the LIBERO-Pro benchmark from 27.6% to 57.6%.


Comments (0)
No comments yet. Be the first!