The rapid increase in robot model parameters has not been matched by the availability of high-quality physical interaction data. Current compliant data in China stands at only 500,000 hours, while commercial deployment requires tens of millions of hours, resulting in a gap exceeding 99%. The China Academy of Information and Communications Technology indicates that embodied intelligence models need at least tens of millions of hours of data to reach a 'ChatGPT moment', yet globally available high-quality data is still far from sufficient.
The scarcity of data is not due to a lack of collection efforts; approximately 100 embodied intelligence data collection centers have emerged in China over the past two years. However, the data collected is often of poor quality, incompatible formats, and not reusable across different projects. The high cost of collecting real machine data, estimated at around 275 yuan per hour for effective data, exacerbates the issue, making it a rare and expensive resource.
As over 70 training centers are operational and more than 40 are under construction, concerns about the quality of data collected persist. Some reports describe the business model of these centers as 'circular financing', where robot companies sell machines to government-built data centers and then funnel money back under the guise of data procurement. No further timeline was disclosed at the time of publication.
Editor's Note
The robotics industry is grappling with a significant data shortage, which poses challenges for the development and deployment of embodied intelligence technologies. As companies invest in data collection infrastructure, the focus must shift to ensuring data quality and compatibility to maximize the utility of collected data. This situation highlights the need for innovative solutions to bridge the gap between data availability and the requirements of advanced robotic systems.
Leave a comment