At the World Artificial Intelligence Conference (WAIC), a notable trend emerged in the robotics sector, with an increasing number of VLA models attempting to enable robots to execute longer task chains. A significant challenge has been the fundamental conflict between 'contextual memory' and 'real-time inference costs' in VLA models, which has hindered the completion of long-range tasks.
During WAIC, Mianbi Intelligence introduced and open-sourced its first series of embodied intelligence results, the MiniCPM-Robot, which includes the general VLA model MiniCPM-RobotManip and the tracking navigation model MiniCPM-RobotTrack. This model, with only 1.3 billion parameters, successfully completed long-range tasks like making sandwiches and achieved a significant lead over well-known models in the RMBench leaderboard, showcasing its advanced contextual memory capabilities.
The release of MiniCPM-Robot signifies Mianbi Intelligence's transition of its accumulated multimodal technology from the digital realm to the physical world. The MiniCPM-Robot series addresses a structural shortcoming in current embodied intelligence, allowing for efficient visual token compression that enhances real-time inference speed while retaining memory, thus overcoming a long-standing dilemma in VLA model development.
Editor's Note
The advancements in embodied intelligence showcased at WAIC highlight a critical shift in the robotics landscape, particularly in addressing the limitations of VLA models. The ability to maintain contextual memory without sacrificing speed could significantly enhance the efficiency and effectiveness of robotic applications across various sectors, from manufacturing to service industries.
Leave a comment