The model functions as a closed-loop system where robots generate candidate actions based on language instructions and visual history, predict the resulting states, and evaluate whether those outcomes achieve specific goals. During execution, Motus2 employs Best-of-N planning to compare potential actions in real-time, executing the option with the highest predicted success score. This architecture effectively bridges the gap between pre-trained human manipulation knowledge and physical robotic control.
Performance benchmarks highlight the system's efficacy. In real-robot trials involving tasks such as screwing in a light bulb and multi-finger manipulation, Motus2 achieved an average success rate of 84%. Researchers observed that incorporating a lightweight tactile expert, which processes contact feedback to refine movements, improved performance on tasks like tearing paper by 12.5 percentage points. The model is trained on a massive foundation of 130,000 hours of egocentric human video data, subsequently refined with over 100 hours of robot-specific trajectories.



Comments (0)
No comments yet. Be the first!