A single destination for timely, editor-curated robotics news from around the world.
AGIBOT announced that its WITA-Omni Preview multimodal foundation model has achieved the highest score on the Daily-Omni audio-visual reasoning benchmark, surpassing models from Alibaba, Google, and ByteDance. The model recorded an average accuracy of 85.21 percent, excelling in audio-visual alignment, event sequencing, and various evaluation metrics. This achievement is significant as it demonstrates WITA-Omni Preview's advanced capabilities in understanding and reasoning across audio and visual information, which are crucial for embodied AI systems in dynamic environments. The benchmark evaluates models on their ability to associate sounds with visible events and understand the progression of situations, highlighting the importance of these skills in real-world applications. Looking ahead, AGIBOT's WITA-Omni model is designed to enhance human-robot interactions by integrating movement and facial expressions with speech. The company has developed a human-centric multimodal interaction dataset to further improve the model's performance. No further timeline was disclosed at the time of publication.
RoboticsAndAutomationNews.com By David Edwards Aug 03, 2026 Computing Design Humanoids News agibot AI modelsRSF defines a common language for robot service capability, lifecycle operations, certification pathways, and service-provider networks.
Daily robotics news, in-depth analysis, conference highlights, and discussions with professionals worldwide.