In the past six months, the focus of the domestic embodied intelligence sector has shifted from hardware competition to the deeper challenges that define the intelligence limits of robots. Luo Jianlan, an associate professor at Shanghai Chuangzhi Academy and chief scientist at Zhiyuan Robotics, argues against the prevailing notion that robots can replicate large language models through sheer data accumulation. He emphasizes that the core issue in embodied intelligence is not about breakthroughs in isolated components but rather the ability to create a closed-loop system in real-world deployments.
Luo, who has a background in both academia and industry, including roles at Google X and DeepMind, believes that many teams in the sector are not genuinely pre-training models but are instead engaged in mid-training or fine-tuning due to the scarcity of high-quality interaction data. He asserts that true embodied intelligence requires a scalable closed-loop system, where deployment leads to data collection, which in turn enhances model capabilities.
His current focus includes developing scalable online post-training infrastructure, enabling robots to learn continuously in real-world environments, and creating a world model that predicts the consequences of actions rather than merely generating video. Luo suggests that the future of embodied intelligence hinges on successfully integrating these elements into a cohesive system, with significant advancements expected in the next 12 to 18 months. He believes that the first team to effectively implement a "deployment-data-iteration" cycle in semi-structured environments like convenience stores will gain a substantial competitive edge.
Leave a comment