Ant Group's subsidiary, Ant Lingbo, has made a bold decision to abandon existing digital pre-trained models and build a native visual foundation for robots from scratch. Chief Scientist Shen Yujun emphasizes that while traditional visual models focus on image quality and scene recognition, robots require an understanding of how to interact with objects in real-time, necessitating a complete overhaul of training methods.
This approach is significant as it addresses the limitations of existing models that cannot process real-world scenarios effectively. The second-generation model has seen a substantial increase in pre-training data from 20,000 hours to 60,000 hours, expanding robot configurations and improving efficiency in training. The new model is built on three technological pillars, including a mixture of experts architecture and a causal unidirectional attention mechanism, which have been developed to meet the demands of the physical world.
Looking ahead, Ant Lingbo aims to gather millions of hours of native pre-training data to enhance its models. Shen believes that China will surpass the U.S. in data scale due to its robust hardware supply chain and increased production of robots. No further timeline was disclosed at the time of publication.
Editor's Note
Ant Lingbo's innovative approach to robot vision could reshape the landscape of robotics by prioritizing real-time interaction over traditional image recognition. This shift may influence how companies approach the development of intelligent systems and the integration of AI in robotics, particularly in manufacturing and automation sectors.
Leave a comment