A single destination for timely, editor-curated robotics news from around the world.
A comprehensive survey on vision-language-action models for embodied artificial intelligence has been published in the Journal of Field Robotics. This survey explores the integration of visual perception, language understanding, and action execution in AI systems, highlighting the advancements and challenges in this interdisciplinary field. The significance of this survey lies in its potential to enhance the development of more capable and intelligent robotic systems. By examining the interplay between vision, language, and action, researchers can better understand how to create AI that can interact with the world in a more human-like manner, which is crucial for applications in various sectors. Looking ahead, the survey may pave the way for future research initiatives aimed at improving embodied AI systems. No further timeline was disclosed at the time of publication.
JournalofFieldRobotics By Ning Xiong, Mingle Xu, Wei Chen, Jianming Liu, Chuanlei Zhang, Yuan Wang, Jucheng Yang Aug 26, 2026 SURVEY ARTICLE
Google has introduced a sign-language translation model capable of converting intricate body movements into text, utilizing over 100,000 hours of data from more than 50 sign languages. This technology, known as sign-language-to-text (SL2T), is being integrated into consumer devices, starting with American Sign Language (ASL) on Pixel 11. Users can now sign to search the web, compose messages, and engage with Google’s Gemini, enhancing accessibility for the Deaf community. The significance of this development lies in its ability to address the unique complexities of sign languages, which possess their own grammar and vocabulary. Unlike traditional speech transcription, sign-language translation involves interpreting simultaneous hand movements, facial expressions, and body posture. Google’s SL2T model employs computer vision techniques to convert the signer’s movements into a structured format before translating them into text, ensuring a more accurate representation of the signed language. Looking ahead, Google aims to refine the SL2T system further by addressing practical challenges such as latency and performance for various signing styles. The involvement of Deaf users and organizations throughout the development process underscores the commitment to creating a user-friendly and effective translation tool. No further timeline was disclosed at the time of publication.
InterestingEngineering.com By Neetika Walter Aug 12, 2026 AI and Robotics
On September 14, Digua Robotics and Giga Vision announced a strategic collaboration focused on integrating their respective technologies. Giga Vision will provide world models, embodied foundational models, and real-world application experience, while Digua Robotics will contribute an edge AI computing platform, algorithm toolchain, and robotics ecosystem. The initial integrated model chosen is GigaBrain-0.7, utilizing the Xuri S600 hardware base and GigaWorld's capabilities. This partnership aims to create an affordable, integrated solution for embodied intelligence at the edge. GigaWorld will handle scene generation, action consequence prediction, strategy evaluation, and retraining of failure samples, providing a virtual training environment for robots. GigaBrain will serve as the edge intelligence core, managing natural language tasks, spatial perception, task decomposition, skill routing, and decision-making, while the Xuri S600 will support multimodal reasoning and task scheduling. Looking ahead, both companies plan to accelerate the practical application of GigaBrain-0.7 across various robotic platforms, including industrial manufacturing and home services. The success of this collaboration will depend on the performance of GigaBrain-0.7 in real-world robotic applications, particularly in terms of task success rates and system stability.
leaderobot.com By Leaderobot Sep 14, 2026 Robotics AI Edge Computing Embodied Intelligence
AgiBot has introduced GE-Act 2.0, a native world-action model that has been pretrained from random initialization using embodied data. This new model has significantly scaled from 300 to 30,000 hours of training, enabling it to perform zero-shot skills, including towel folding, across two different robot embodiments. The release of GE-Act 2.0 is significant as it demonstrates AgiBot's commitment to advancing robotic capabilities through extensive data scaling. By increasing the training hours, the model can now execute complex tasks without prior specific training, showcasing the potential for greater versatility in robotic applications. Looking ahead, it will be important to monitor how GE-Act 2.0 performs in real-world scenarios and whether it can be adapted for additional tasks beyond towel folding. No further timeline was disclosed at the time of publication.
PanDaily.com By [email protected] (Pandaily) Sep 10, 2026
Beta Infinite has introduced its BetaWAM 0.1 world action model, showcasing a robot capable of autonomously completing complex, long-range tasks in dynamic environments. This model addresses the limitations of existing technologies that primarily focus on single-point actions or fixed operations. The significance of BetaWAM 0.1 lies in its ability to perform multiple objectives without human intervention, overcoming challenges such as obstructions and sudden changes in the environment. This advancement is crucial for the future of consumer robotics, as it aims to bridge the gap between action intelligence and task intelligence. Looking ahead, Beta Infinite's approach to integrating memory and spatial understanding into its robotic systems could redefine operational capabilities in the robotics sector. No further timeline was disclosed at the time of publication.
leaderobot.com By Leaderobot Aug 17, 2026 Embodied Intelligence Robotics Task Automation AI Consumer Robotics
Mimic Robotics has launched FLUX-mimic, an advanced Video-Action Model developed with Black Forest Labs, designed for industrial automation. This model allows robots to learn complex manipulation tasks in real-world settings, significantly reducing the time required for training and deployment. The introduction of FLUX-mimic is crucial as it addresses the limitations of traditional robot learning methods, which often rely on extensive demonstration data. By utilizing a generative video model, FLUX-mimic can fine-tune tasks with as little as 30 minutes of data, compared to the 30 hours typically needed, thereby streamlining the deployment process. Looking ahead, Mimic Robotics is collaborating with Audi to implement FLUX-mimic in their highly automated production network. This partnership aims to enhance the efficiency of industrial automation, paving the way for smart factories where AI and robots work alongside human employees to optimize production processes.
RoboticsAndAutomationNews.com By David Edwards Jul 29, 2026 Artificial Intelligence Computing Design News audi Black Forest Labs
NASA's Jet Propulsion Laboratory has successfully sent Google's Gemma 3 to space, marking the first in-orbit demonstration of a vision-language model analyzing satellite imagery. The NAVI-Orbital system utilized Gemma 3 to interpret images from Loft Orbital's YAM-9 satellite, showcasing a new method for scientists to interact with spacecraft through natural language prompts. This advancement is significant as it allows researchers to bypass traditional structured commands, enabling more intuitive communication with satellites. Juan M. Delfa from NASA highlighted that this shift could enhance how scientists engage with space missions, potentially streamlining operations and improving data analysis. Looking ahead, the implications of NAVI-Orbital extend beyond image analysis. The system could revolutionize satellite operations by enabling real-time data interpretation and reporting, which is crucial for applications like wildfire detection. No further timeline was disclosed at the time of publication.
IEEESpectrumAI By Matthew S. Smith Jul 23, 2026 Nasa Image-analysis Llms Satellite-imagery Google
On July 15, Stardust AI introduced its second-generation embodied base model, Lumo-2, which is the industry's first household latent world-action model. This launch includes the physical AI symbiotic agent, Agent Philia, enhancing their full-stack architecture of AI models, embodied operating systems, and rope-driven entities. The company will showcase its 'trinity' multi-scenario implementation solutions at the World Artificial Intelligence Conference in Shanghai from July 17 to 20. Lumo-2 autonomously performs 22 complex household tasks, demonstrating industry-leading capabilities in task range and complexity. This model addresses the challenges faced by robots in open environments, such as the inability to explain actions and the high costs of training complex skills. By predicting future scenarios before generating actions, Lumo-2 aims to overcome these bottlenecks and improve the practical execution of robotic tasks. Looking ahead, Stardust AI plans to enhance the scalability of Lumo-2 by expanding training data diversity and exploring efficient data engineering paradigms. The team is also focused on advancing real-world interactive learning to enable robots to adapt and evolve autonomously in dynamic environments. No further timeline was disclosed at the time of publication.
leaderobot.com By Leaderobot Jul 15, 2026 Household Robotics Physical AI AI Models Robotic Automation
Helix, an innovative Vision-Language-Action model, has been developed to enhance humanoid robotics by providing full upper-body control and facilitating collaboration among multiple robots. This cutting-edge technology enables robots to execute tasks involving new objects through natural language prompts, significantly improving their versatility and usability. Notably, Helix operates efficiently on low-power GPUs, positioning it for commercial applications. With its capabilities, Helix is set to revolutionize the field of robotics, making advanced robotic interactions more accessible and practical for various industries.
figure.ai By Figure AI Feb 20, 2025 robotics AI machine learning humanoid robots automation
Microsoft has unveiled a provisional code of conduct that establishes restrictions for its future artificial intelligence models. This decision follows a growing consensus among AI leaders, including those from Anthropic and OpenAI, advocating for a deceleration in AI development due to rising safety concerns. Mustafa Suleyman, CEO of Microsoft AI, emphasized the importance of AI serving humanity and promoting human autonomy. The initiative is significant as it reflects a broader industry shift towards responsible AI development, addressing public apprehensions about the rapid advancement of AI technologies. Recent events, including a resignation from an Anthropic researcher who criticized the race towards self-improving superintelligence, have intensified calls for more stringent AI safeguards. Microsoft aims to position itself as a responsible player in the AI landscape, particularly as it integrates models from leading AI labs into its products. Looking ahead, Microsoft plans to refine its guidelines further, with an update expected to influence AI model development starting in 2027. The company has engaged with experts across various fields to shape its code of conduct, indicating a commitment to ethical AI practices. No further timeline was disclosed at the time of publication.
CNBCTechnology 6 hours ago
Humanoid robots, designed to mimic human limbs and body structures, are being enhanced by an AI controller that translates virtual reality, video, and language commands into actionable movements. This advancement aims to simplify the teaching process for these robots, which often struggle with executing humanlike movements reliably. The significance of this development lies in its potential to streamline the deployment of humanoid robots across various environments, including homes and workplaces. By improving the efficiency of command translation, the AI controller could reduce the time and effort required to train these robots, making them more accessible for practical applications. Looking ahead, the focus will be on the effectiveness of the AI controller in real-world scenarios and its ability to adapt to diverse tasks. No further timeline was disclosed at the time of publication.
TechXplore:Robotics Sep 09, 2026 Robotics
RFID and machine vision technologies are increasingly being adopted as effective solutions for automation. These technologies provide a more practical and cost-efficient alternative to traditional large-scale robotics projects, which often necessitate extensive facility redesigns and substantial capital investments. The significance of RFID and machine vision lies in their ability to seamlessly integrate into existing warehouse and manufacturing operations. This integration allows businesses to enhance their operational efficiency without the need for disruptive changes, making automation more accessible to a wider range of industries. Looking ahead, the continued growth of RFID and machine vision technologies is expected as companies seek to optimize their processes. No further timeline was disclosed at the time of publication.
SupplyChainBrain Aug 24, 2026
Researchers from Nanyang Technological University, Peking University, HKUST (GZ), and Beijing Academy of Artificial Intelligence (BAAI) have introduced Omega-0, a latent-prediction world action model. This innovative model enables a humanoid robot to perform multiple actions simultaneously, including walking, looking, and working. The significance of Omega-0 lies in its impressive performance, achieving an 81.8 percent success rate across 11 real home tasks. This success surpasses that of previous models such as pi-0.5, EgoVLA, GR00T-N1.7, and psi-0, highlighting advancements in robotic capabilities for practical applications in domestic environments. Looking ahead, the development of Omega-0 suggests a growing trend in robotics focused on enhancing the functionality of humanoid robots in everyday tasks. No further timeline was disclosed at the time of publication.
PanDaily.com By [email protected] (Pandaily) Aug 11, 2026
Tencent has announced the open-source release of three embodied foundation models during the WAIC 2026 event. These models include a Visual Language Model (VLM) designed for scene understanding, the RxBrain cognitive model that facilitates planning with visual states, and the VLA model, which supports continuous action at frequencies between 500 to 1000Hz. This development is significant as it aims to enhance robot reaction speed and cognitive capabilities, addressing critical challenges in robotic performance. The introduction of these models is expected to advance the field of robotics by providing developers with powerful tools to improve the efficiency and effectiveness of robotic systems. Looking ahead, the industry will be keen to observe how these open-source models are adopted and integrated into various robotic applications. No further timeline was disclosed at the time of publication.
PanDaily.com By [email protected] (Pandaily) Jul 26, 2026 Technology
Ant LingBot, a subsidiary of Ant Group, has launched six open-source embodied AI models as part of its dual-track strategy focusing on Visual Language Agents (VLA) and world models. This initiative aims to enhance AI capabilities while addressing the growing demand for advanced AI solutions. The significance of this release lies in Ant LingBot's commitment to fostering an open-source ecosystem, which is crucial for collaboration and innovation in the AI field. However, the company is contending with challenges related to data scarcity and competition within the ecosystem, which could impact its development and deployment efforts. Looking ahead, it will be important to monitor how Ant LingBot navigates these challenges and whether it can successfully leverage its dual-track strategy to establish a strong presence in the AI landscape. No further timeline was disclosed at the time of publication.
PanDaily.com By [email protected] (Pandaily) Jul 23, 2026 Technology
Dexmal has launched its DM0.5 foundation model, DexOS operating system, and an embodied Mobility-as-a-Service (MaaS) platform during the Action developer conference. This event took place recently and aims to position Dexmal as a key player in the robotics industry by creating a versatile platform akin to Android for robotics applications. The introduction of the DM0.5 model and DexOS is significant as it seeks to address the challenges of scaling robotic models into practical, real-world scenarios. By providing a unified operating system and a robust foundation model, Dexmal aims to enhance interoperability and functionality across various robotic applications, potentially transforming how developers approach robotics solutions. Looking ahead, Dexmal's next steps involve further development of its MaaS platform and expanding the capabilities of the DM0.5 model. No further timeline was disclosed at the time of publication, but industry watchers will be keen to see how these innovations influence the robotics landscape and attract developer interest.
PanDaily.com By [email protected] (Pandaily) Jul 10, 2026 Robotics
Tesla's Optimus robots will not be used to repair Starmind satellites in orbit, as confirmed by recent statements from Elon Musk. Instead, these robots are intended to assist in the construction and operation of the Terafab chip manufacturing facility in Texas. The AI1 satellites, designed to disintegrate upon reentry, highlight the company's swap-and-replace strategy rather than traditional maintenance practices. This approach is significant as it reflects a broader trend in satellite management, where mass-produced satellites are replaced rather than repaired. The economics of servicing missions are prohibitive, with the cost of launching a replacement satellite being significantly lower than conducting a repair mission. This model aligns with SpaceX's operational history, where rapid replacement of satellites is more efficient than attempting to maintain them in orbit. Looking ahead, the focus will remain on the production capabilities of the Gigasat factory, which is expected to support the continuous replacement of satellites. No further timeline was disclosed at the time of publication, but the demand for rapid satellite turnover suggests a robust future for Optimus robots in terrestrial manufacturing rather than in-space servicing.
optimusk.blog By OptimusK Blog Jul 08, 2026
Reasoning about failures is crucial for building reliable and trustworthy robotic systems. Prior approaches either treat failure reasoning as a closed-set classification problem or assume access to ample human annotations. Failures in the real world are typically subtle, combinatorial, and difficult to enumerate, whereas rich reasoning labels are expensive to acquire. We address this problem by introducing
amazon.science By Amazon Science May 19, 2026 Automated reasoning
A team of former researchers from MIT and DeepMind has established Eka Robotics, unveiling a groundbreaking "Vision-Force-Action" (VFA) foundation model. This innovative technology harnesses the power of simulation and tactile sensing to enable robots to perform tasks with superhuman speed and remarkable physical intelligence. The launch marks a significant advancement in robotics, aiming to enhance the capabilities of machines in various applications. By integrating advanced sensory data with real-time decision-making, Eka Robotics seeks to revolutionize how robots interact with their environments. The initiative reflects a growing trend in the tech industry to develop more sophisticated and responsive robotic systems, addressing the increasing demand for automation across multiple sectors.
HumanoidsDaily By [email protected] (Humanoids Daily Staff) Apr 29, 2026 US Eka Robotics
NVIDIA has introduced the Nemotron 3 Nano Omni, an innovative open multimodal AI model designed to enhance the efficiency of AI agent systems. Announced today, this model integrates vision, speech, and language capabilities into a single framework, addressing the common issue of time and context loss that occurs when data is transferred between separate models. By streamlining these processes, the Nemotron 3 Nano Omni aims to improve the performance of AI applications across various domains. This advancement is particularly significant as it allows for more cohesive and contextually aware interactions, marking a notable step forward in the development of AI technologies.
NvidiaNews By NVIDIA Apr 28, 2026
In a significant development for the manufacturing sector, experts have highlighted the transformative potential of Variational Latent Models (VLMs) in enhancing quality assurance processes. While acknowledging that VLMs will not address every challenge faced in the realm of artificial intelligence within manufacturing, they emphasize that these models provide a unique capability that surpasses existing technologies, particularly in high-complexity production environments. This advancement comes at a time when industries are increasingly seeking innovative solutions to improve efficiency and accuracy in their operations. As manufacturers strive to meet rising demands and maintain high standards, the adoption of VLMs could represent a pivotal shift in how quality assurance is approached, ultimately leading to more reliable and efficient production outcomes.
roboticstomorrow-Robotics Apr 03, 2026
At the 2026 World Robot Conference, a clear trend emerged indicating that the embodied intelligence industry is shifting focus from creating general-purpose robots to developing specialized solutions that deliver real value in specific scenarios. Innovations in areas such as garment sewing, biomedical applications, and airport luggage handling highlight this transition. The significance of this shift lies in the ability of specialized robots to address industry-specific challenges, such as labor shortages and rising costs in garment production. For instance, Aitu's humanoid robot is designed solely for garment sewing, showcasing its capability to autonomously identify fabrics and perform precise stitching without the need for constant reprogramming. Looking ahead, the conference revealed advancements in collaborative robotics, with companies exploring multi-robot systems for enhanced efficiency. Notably, Xingyuan Intelligence demonstrated a novel multi-robot collaboration in a jumping rope task, indicating a trend towards more complex interactions among robots. No further timeline was disclosed at the time of publication.
leaderobot.com By Leaderobot Sep 09, 2026 Specialized Robots Industrial Automation Biomedical Robotics Robotic Collaboration
AGIBOT presented its newest embodied AI robotics innovations at IFA 2026, including the AGIBOT A3 humanoid and AGIBOT X2 Ultra. These robots are designed for various applications, such as retail, industrial operations, and commercial cleaning. The significance of this showcase lies in AGIBOT's commitment to enhancing robotic capabilities in practical environments. The announcement of TÜV Rheinland certifications for its robots further underscores the company's dedication to safety and quality standards in robotics. Looking ahead, AGIBOT's partnership with Tekpoint aims to strengthen its foothold in the European market. No further timeline was disclosed at the time of publication.
agibot.com By AgiBot Sep 04, 2026 AI Robotics Humanoid Robots Commercial Cleaning Retail Technology European Market Expansion
Researchers at Delft University of Technology (TU Delft) have created a system that allows passengers to influence the driving style of autonomous vehicles using natural language requests. This system utilizes a large language model (LLM) to interpret user instructions, such as adjusting speed or smoothness based on individual preferences, while maintaining safety through a motion-planning algorithm. The significance of this development lies in its potential to enhance user experience in self-driving cars by making them more adaptable to personal preferences. By allowing passengers to communicate their needs in everyday language, the system aims to bridge the gap between human driving styles and autonomous vehicle behavior, ensuring a smoother and more comfortable ride. Looking ahead, the researchers plan to further refine the system and explore its applications in real-world scenarios. The system's interactive nature, which allows for continuous adjustments based on passenger feedback, could pave the way for more personalized and user-friendly autonomous driving experiences. No further timeline was disclosed at the time of publication.
IEEESpectrumAI By Edd Gent Aug 24, 2026 Autonomous-vehicles Journal-watch Large-language-models
Chinese robotics firm Unitree has introduced UnifoLM-OminiA-0.3, a unified AI model designed for humanoid robots. This model facilitates real-time omni-modal interaction, reasoning, dialogue, and whole-body mobile manipulation, targeting home-care and wellness applications. The robots can autonomously perform tasks such as tidying rooms and assisting patients while responding to various inputs. The significance of UnifoLM-OminiA-0.3 lies in its ability to integrate multiple capabilities into a single system, allowing robots to process information from speech, vision, and environmental cues simultaneously. This unified architecture enables seamless task execution, as demonstrated by a humanoid robot that can adjust a hospital bed and respond to user commands mid-task, showcasing continuous human-robot interaction. Looking ahead, the trend towards embodied AI is expected to grow, with developers focusing on integrating vision-language models with robot control. This approach enhances flexibility in dynamic environments like homes and healthcare facilities, where tasks and interactions can vary significantly. No further timeline was disclosed at the time of publication.
InterestingEngineering.com By Jijo Malayil Jul 21, 2026 AI and Robotics
On July 10, Ant Group introduced LingBot-VA 2.0, the first embodied native action model in the industry. This release marks the culmination of their full-stack 2.0 model series, featuring advancements in spatial perception and real-time interaction. The launch has generated significant buzz across international platforms like Reddit and Hugging Face, indicating a strong interest in their technology. The significance of LingBot-VA 2.0 lies in its comprehensive technology stack, which includes LingBot-Depth 2.0 and LingBot-Vision. LingBot-Depth 2.0 enhances depth perception with a training dataset expanded from 3 million to 150 million, achieving top scores in 12 out of 16 benchmarks. Meanwhile, LingBot-Vision introduces a novel pre-training target, improving depth estimation accuracy despite using a smaller dataset compared to competitors. Looking ahead, the next steps for Ant Group involve further collaboration with industry partners, as LingBot-Depth 2.0 has already received professional certification from Orbbec. The company is also focusing on integrating their models into the industry, with no further timeline disclosed at the time of publication for upcoming releases or partnerships.
leaderobot.com By Leaderobot Jul 11, 2026 Robotics AI Machine Learning Computer Vision
On July 9, Yuanli Lingji introduced three key products, including the DM0.5 model and Apex robot, during the Action 2026 Developer Conference. This event highlighted the company's commitment to full-stack capabilities in embodied intelligence, a crucial factor in addressing industry fragmentation and enhancing model architecture with high-quality data. The significance of these product launches lies in their potential to drive commercialization within the embodied intelligence sector. By focusing on full-stack solutions, Yuanli Lingji aims to set itself apart in a competitive market, where the integration of robust data and model frameworks is essential for success. Looking ahead, industry observers will be keen to see how these products perform in the market and whether they can effectively address the challenges of fragmentation. No further timeline was disclosed at the time of publication regarding future developments or additional product releases.
leaderobot.com By Leaderobot Jul 09, 2026 Embodied Intelligence Robotics AI Machine Learning Automation
SpaceX has officially named its orbital AI infrastructure project 'Starmind,' which aims to deploy a constellation of up to 1 million satellites. This initiative, confirmed by Elon Musk on June 22, 2026, will enable AI inference directly in space, utilizing solar energy rather than terrestrial power sources. The first satellite, designated AI1, was unveiled on June 8, 2026, and is designed to operate in sun-synchronous orbits. The significance of Starmind lies in its potential to overcome the limitations faced by ground-based data centers, such as land, power, and water constraints. By running AI computations in orbit, Starmind can provide a more efficient solution to the growing demand for AI computing power. The project leverages the existing Starlink infrastructure for data transmission, distinguishing its function from Starlink's internet relay capabilities. Looking ahead, SpaceX plans to begin hardware deployment with the AI1 satellite, while full-scale production and deployment of the satellite constellation are targeted for 2028. As of now, no Starmind satellites have been launched, and further engineering challenges remain to be addressed, particularly regarding the scalability of the satellite design.
optimusk.blog By OptimusK Blog Jul 08, 2026
SpaceX's Starmind is designed to provide wholesale AI compute services to businesses, particularly AI labs and cloud customers, rather than individual consumers. The service operates similarly to AWS, where users benefit from applications running on Starmind without direct subscriptions. The compute capacity of a single AI1 satellite is comparable to one NVIDIA GB300 rack, emphasizing its enterprise-grade capabilities. The significance of Starmind lies in its positioning as a potential fourth hyperscaler, joining the ranks of AWS, Microsoft Azure, and Google Cloud. The Reflection AI contract, valued at $150 million per month, exemplifies the enterprise-focused model, with total payments potentially reaching $6.3 billion through 2029. This contract highlights the growing demand for AI compute resources, particularly from AI-native startups and labs. Looking ahead, the focus will remain on securing additional enterprise contracts as Starmind expands its offerings. No consumer-facing products or subscriptions have been announced, and the current strategy is to cater to businesses with substantial AI workloads. No further timeline was disclosed at the time of publication.
optimusk.blog By OptimusK Blog Jul 08, 2026
One morning in 2019, Adebayo Alonge was in a Cape Town hotel room, preparing to demonstrate his startup’s AI answer to a serious problem in African health care: counterfeit medication, which kills thousands of people across the continent every year.The RxScanner is a handheld spectrometer that scans a pill with infrared light, then sends the item’s molecular profile to an AI model equipped with a pharmaceutical database. In seconds, the AI identifies the medication from its molecular profile—or reports that it’s phony.Pharmacies were using the system in more than a dozen countries, including Ghana, Kenya, Myanmar, and Alonge’s native Nigeria. But that morning in South Africa, it didn’t work. “I was shocked,” Alonge says.The spectrometer connected to the AI model—but the data center was 14,000 kilometers away and bandwidth was limited. “Our server was in the United States, and just to get the result of a single scan was taking me over 5 minutes.”So Alonge immediately asked his engineers to shrink the AI model down to a smaller, low-power, unconnected version that could run entirely on his Android phone. They produced it 2 hours later, and that saved the demo.More importantly, the work birthed a new version of his device, which can authenticate a pill in places without broadband, computers, or even reliable electricity. It also turned Alonge into an advocate for this kind of “small AI.”Small AI for Global Health Care AccessSmall AI is a far cry from wealthy nations’ colossal large language models (LLMs), hyperscale data centers, multibillion-dollar investments, and debates about AI consciousness. But for millions of people around the world, the only AI that matters, and often the only kind available, is small. (According to a World Bank Report issued in November, only 0.7 percent of internet users in the world’s poorest countries have used ChatGPT, compared to a quarter of all internet users in the most developed nations.)“Most people are discussing AI from the LLM/generative side. But that needs a lot of computing power, electricity, massive data, and skilled people to manage it,” Ajay Banga, president of the World Bank, said last January at the World Economic Forum, in Davos. “Outside the developed world, other than maybe India and China, very few countries have that combination.”By contrast, small AI can deliver useful, even life-saving services to people in areas that have none of those things, Banga said. In India, where the government’s AI plans call for more development of small AI, many such systems are working for farmers.For example, a drone-based system developed by Bala Murugan and colleagues at the Vellore Institute of Technology, in India, takes photos of cashew plants and quickly identifies those with splotches that indicate disease. All the processing takes place on the drone itself, so there’s no need for a computer on-site, nor for a connection to a central server.Using small language models trained for a specific problem, and sometimes running on cheap, low-power devices, other small-AI implementations have been developed to identify ant infestations in a Uruguayan vineyard, detect the presence of malaria-carrying mosquitoes in a number of nations, and run electrocardiograms from an Arduino device in parts of Brazil that lack access to more complex equipment.“This is the most important area in AI nowadays,” says Marcelo José Rovai, a professor at the Institute of Engineering and Information Systems at the Federal University of Itajubá, in Brazil, who was involved in all three projects. “It’s growing very fast.”Low-Power, Small-AI Models on Devices Small AI models can run on a variety of low-power devices, including [from left to right] an Arduino Nano 33 BLE Sense, a Seeed Wio Terminal, and an Arduino Portenta.Moez AltayebFor Alonge, Rovai, and other advocates, small AI is not just “a promising trend,” as that November World Bank report calls it. It may be, in the long term, the form of AI that will touch the most lives and remain sustainable after some of the giant models become too costly for most users.“I think the future of AI is not like one giant model, at a center. I think it’s millions of small, precise models deployed at the edge, each one solving like a specific problem, a specific context,” Alonge says. This is partly because much of humanity—including people in parts of rich countries as well as the developing world—lives without access to cutting-edge frontier models. But, he says, it’s also because those models are not sustainable.“If someone is not subsidizing it, most people will not be able to afford those models. So those of us who are said to be small-AI developers are the ones who will have to build for the majority of the world,” Alonge says.There is no strict definition of “small AI,” but people often use the term for language models with at most a few billion parameters. (Compare that to cutting-edge models, which can include more than a trillion.) That’s small enough to run directly on a phone or a Raspberry Pi. That’s what allows these applications to run on devices without a connection to a data center and use only a few watts of power, often supplied by a battery or a solar panel.Despite their small footprint, these models aren’t fundamentally different technology from that of gigantic AI models, Rovai says. Many instances of small language models were created the same way the phone-based version of Alonge’s pharmaceuticals scanner was—by “pruning” large models, or removing the parameters that weren’t involved in the task. The result is a system that’s less capable generally but still very good at the specific job it was pruned for, Rovai says. A lighter version of RxAll’s RxScanner spectrometer sends its results to an AI model run locally on a phone to check that a drug’s molecular signature is genuine.RxAllOther small models are created by “distillation.” They are trained to mimic a large model, until their performance approaches that of their “teacher,” Rovai says. In other cases, a larger model’s precision is reduced, for example, so that a model run on 32-bit architecture can run on 8-bit designs. In situations where the machine learning application is being used to classify data or predict patterns (like an ant infestation), it’s trained from the beginning on a small device, not derived from a larger model at all. Running all these small, specialized systems is becoming easier, Rovai says, for two reasons.The first reason is that hardware is getting better and more capable while using less power, he says. This means more and more phones can run small AI—especially those equipped with neural processing units, which are specialized chips that handle AI tasks like facial recognition and changing the brightness, shadows, or contrast in a photo.In 2025, slightly more than a third of all smartphones shipped worldwide were capable of running generative AI, and that figure will reach 45 percent by the end of this year, according to the technology research firm Counterpoint. By the end of next year, slightly more than half of all smartphones will be able to run a small AI model.The second reason Rovai cites is the shrinking footprint of language models. Both Google DeepMind’s Gemma 4 (released in April) and Alibaba’s Qwen 3.5 are “fantastic” for small AI, Rovai says. Both models are “open weight,” meaning users can adjust the connections between parameters to suit their needs. This makes it easy, for example, “to take a lot of data from, say, the milk industry and retrain the model specifically on that,” Rovai says.Rovai illustrated these reasons on a Zoom call, using one of his most recent experiments. Holding up a device, he says, “This is the new Arduino UNO Q—a US $50 device with a Qualcomm chipset. I’m running a language model here, which collects data from sensors and analyzes that data to detect tiny pools of water where mosquitoes might be breeding. It takes 3 watts to run it.”Support for Small-AI DevelopmentConvinced that millions of people are already benefiting from these kinds of applications, the World Bank now actively promotes small AI with grants, mentorship programs, financing, technical advice, and models of government policies that are friendly for small-AI development. For example, in Rwanda, the World Bank is backing a government program to help low-income households get devices that can run AI.All that said, no one claims that large language models are going away entirely. To create a generative AI that can run on a phone or other small device requires the architectural insights, data processing, and results of a larger model, Rovai says. “We need the big models to create these smaller models.” And for all that small AI can benefit people without access to big AI, the technology can’t solve the larger problems of development and digital inequality, Alonge says. Implementing small AI won’t allow nations to escape the challenge of creating an ecosystem to support AI: reliable power, a supply chain that works, and an educational system that develops the talents needed to create AI tools.Though his drug-scanning system can run for days on a phone with no connection, “you still want to be able to enable periodic syncing for updates with new signatures for the medications and analytics,” Alonge says. “And even when you are using batteries, reliable power is important. That phone battery is not going to last forever.”In many parts of the world, the future of small AI isn’t assured, he says. “It works, and many places will eventually need to use it. The question is whether or not the political actors are wise enough to invest in infrastructure to support it long term.”
IEEESpectrumAI By David Berreby Jul 06, 2026 Small-language-models Artificial-intelligence Llms
Researchers have introduced the LA4VLA framework, a new approach that enhances the capabilities of robots in understanding language commands and executing actions. This framework distinguishes language-action supervision from visual input, enabling robots to learn the relationship between commands and actions independently of visual cues. The study, which highlights the limitations of traditional Vision-Language-Action models, was conducted to address the tendency of these models to rely on visual inputs when confronted with conflicting information. By focusing on a more robust language-action learning process, the LA4VLA framework aims to improve the overall understanding of how language influences robotic actions.
leaderobot.com By Leaderobot Jul 03, 2026 Vision-Language-Action Robotics Machine Learning AI Training
Liquid AI, a company founded by former MIT computer scientists, has unveiled its latest AI language model, LFM2.5-230M, which is designed for efficient data extraction and local deployment on devices such as smartphones and laptops. Released today, this 230-million-parameter model is noted for its ability to run on various hardware platforms, outperforming larger models like Alibaba's Qwen3.5 and Google's Gemma 3 in specific benchmarks. Targeting developers and engineers, LFM2.5-230M operates under a dual-use commercial license, allowing free access for individuals and companies with annual revenues below $10 million, while larger enterprises must secure a paid agreement. The model distinguishes itself by utilizing the LFM2 architecture, enabling high inference speeds with a minimal memory footprint, making it suitable for edge computing. Liquid AI's launch reflects a broader industry shift towards architectural efficiency rather than sheer parameter counts, as major AI firms focus on models with hundreds of billions of parameters. The LFM2.5-230M is specifically tailored for lightweight data extraction tasks, allowing businesses to automate processes without relying on costly cloud services. In practical applications, the model has been successfully deployed in a humanoid robot, demonstrating its capability to process complex commands efficiently. Available immediately on platforms like Hugging Face, LFM2.5-230M aims to revolutionize how enterprises manage data extraction, moving away from traditional, rigid systems to more adaptable AI-driven solutions.
Venturebeat.com By [email protected] (Carl Franzen) Jun 25, 2026 Technology
Large language models (LLMs) have transitioned from research labs to everyday use in engineering, significantly altering how digital infrastructures are developed and maintained. As technical professionals increasingly rely on LLMs for complex tasks—such as identifying vulnerabilities in source code and converting fragmented discussions into detailed specifications—the demand for expertise in this technology is surging. According to MarketsandMarkets, the LLM technology market is projected to grow by approximately 33% annually through 2030. To effectively utilize LLMs, engineers must move beyond basic interactions and understand the underlying transformer architecture that enables these models to process vast datasets simultaneously. This knowledge is crucial to mitigate risks associated with inaccuracies, often referred to as "hallucinations," and to ensure reliable performance in coding and data handling. Key advancements include integrating LLMs with application programming interfaces (APIs) for direct database connections, addressing hallucination issues through retrieval-augmented generation (RAG), and prioritizing data security by establishing private model instances. Additionally, LLMs automate repetitive tasks, allowing engineers to focus on higher-level design and problem-solving. To bridge the growing knowledge gap, IEEE has launched an online program titled "Large Language Models Demystified," designed to equip technical professionals with a deeper understanding of LLMs. The curriculum covers the evolution of AI technology, transformer architectures, and practical model-building exercises. Participants will earn professional development credits and a digital badge upon completion, enhancing their credentials in this rapidly evolving field. Organizations interested in training their teams can consult with IEEE for tailored enrollment options.
IEEESpectrumAI By Angelique Parashis Jun 19, 2026 Ai Type-ti Education Ieee-educational-activities Large-language-models Ieee-products-and-services
At the 8th Beijing Zhiyuan Conference, Xingyuan unveiled its innovative ω-EVA model, marking a significant advancement in the field of embodied intelligence. This model represents a shift from traditional world models, which have typically acted as passive observers, to a more dynamic role in robotic decision-making. By integrating real-time feedback into action generation, the ω-EVA model emphasizes the necessity of predicting outcomes prior to executing movements. This development highlights a broader industry trend towards the practical application of artificial intelligence capabilities, showcasing how robotics can evolve to become more responsive and effective in various tasks.
leaderobot.com By Leaderobot Jun 17, 2026 Embodied Intelligence Robotic Decision-Making AI Models Real-Time Feedback Technology Innovation
Researchers at Los Alamos National Laboratory have unveiled a groundbreaking tool named Prelim Attention, designed to enhance the analysis of complex data sets. This innovative tool, which leverages advanced machine learning techniques, aims to streamline the process of identifying significant patterns and insights within large volumes of information. The development was announced in October 2023, highlighting the laboratory's commitment to advancing data science and its applications in various fields. The motivation behind creating Prelim Attention stems from the increasing demand for efficient data analysis solutions in scientific research, national security, and other sectors that rely heavily on data interpretation. By improving the capability to focus on critical data points, the tool is expected to facilitate more informed decision-making and accelerate research outcomes. The researchers employed a combination of algorithms and user-friendly interfaces to ensure that Prelim Attention can be utilized effectively by both experts and non-experts alike. This approach not only enhances accessibility but also broadens the potential user base, allowing a wider range of professionals to benefit from its capabilities. The introduction of Prelim Attention marks a significant advancement in the field of data analysis, promising to transform how researchers and analysts approach complex data challenges in the future.
InterestingEngineering.com By Atharva Gosavi Jun 15, 2026 AI and Robotics
A recent study led by Seung Chan Hong at the University of Melbourne explores the emotional capabilities of collaborative robots as they increasingly work alongside humans. Published on May 18 in IEEE Robotics and Automation Letters, the research investigates how robots can better understand human emotions through contextual cues, beyond just facial expressions. Involving 40 volunteers, the study trained a vision language model (VLM) to interpret emotions based on video interactions where robots handed objects to humans. The VLM outperformed traditional AI systems, scoring 0.86 in emotional accuracy compared to 0.77 for conventional methods. This improvement is attributed to the VLM's ability to consider the entire context of interactions rather than isolated facial expressions. In a follow-up experiment, participants interacted with a robot that was programmed to make an error, receiving either an emotionally adaptive apology or a standard one. The majority preferred the adaptive response, but trust in the robot diminished after it failed to complete its task, highlighting that emotional responses cannot compensate for a lack of functionality. While the VLM effectively recognized emotions from a third-party perspective, its accuracy dropped when compared to participants' self-reported feelings, indicating that robots still struggle to fully understand human emotions. The findings suggest that while emotional adaptivity is valuable, the primary concern for users remains the robot's competence in performing tasks.
Spectrum.ieee.orgAutomaton By Michelle Hampson Jun 13, 2026 Robotics Journal-watch Ai-models Emotion-recognition
On June 3, 2026, ZhiYuan unveiled the second phase of the AGIBOT WORLD 2026 dataset, which centers on the theme of 'Rich Interaction.' This innovative open-source dataset is pioneering in its focus on physical interactions, meticulously documenting both successful and unsuccessful scenarios between robots and their environments. By offering a comprehensive range of data, the initiative seeks to improve world model training, thereby advancing the capabilities of robotic understanding and physical intelligence. This development marks a significant step forward in the field of robotics, as it aims to better equip machines to navigate complex real-world situations.
leaderobot.com By Leaderobot Jun 03, 2026 World Models Robotic Interaction Physical Intelligence Open-Source Datasets
Financial institutions have invested significant time and resources in developing artificial intelligence applications, including fraud detection models, credit assessment tools, recommendation systems, and risk management frameworks. However, despite the effectiveness of these specialized models, the industry faces challenges due to the prevalence of siloed systems. These isolated systems hinder the ability to share data and insights across different departments, limiting the overall potential of AI in enhancing operational efficiency and decision-making. As the financial sector seeks to leverage AI more effectively, there is a growing need to integrate these disparate systems to foster collaboration and innovation. This shift is crucial for maximizing the benefits of AI technologies and addressing the evolving demands of the market.
NvidiaNews By NVIDIA Jun 01, 2026
In recent decades, roboticists around the globe have developed sophisticated robots capable of interpreting human commands and navigating their environments to perform basic manual tasks. Despite these advancements, many of these robots continue to face challenges in accurately translating user instructions into specific, executable actions necessary for successfully completing desired tasks. This ongoing struggle highlights the complexities involved in human-robot interaction and the need for further innovation in robotic technology to enhance their effectiveness in practical applications.
TechXplore:Robotics May 22, 2026 Robotics
Carnegie Mellon University’s Robotics Institute is set to host the latest phase of the Vision-Language-Navigation (VLN) Challenge, aimed at advancing the ability of robots to comprehend and execute human instructions in real-world environments. This new iteration of the challenge, which takes place this year, marks a significant evolution from previous versions by eliminating certain constraints, thereby enhancing the complexity and applicability of the tasks involved. The initiative seeks to unite researchers in tackling one of the most challenging aspects of robotics, ultimately striving to improve the interaction between humans and machines.
ri.cmu.edu By Mallory Lindahl May 01, 2026 Research
Shengshu Technology has achieved a significant milestone with its Motubrain model, which has secured the top position on both the WorldArena and RoboTwin 2.0 benchmarks. This accomplishment highlights the model's innovative unified world-action approach, allowing humanoid robots to effectively execute long-horizon tasks across various embodiments. The recognition from these prestigious benchmarks underscores the advancements in robotics technology and the potential for enhanced performance in diverse applications.
PanDaily.com By [email protected] (Pandaily) Apr 30, 2026 News
Galbot, China's leading unlisted embodied AI company valued at over RMB 20 billion (approximately $2.8 billion), has unveiled its latest innovation, the LDA-1B. This advanced model, featuring 1.6 billion parameters, integrates world and action learning, showcasing its ability to scale effectively with diverse data sets. The LDA-1B has been open-sourced and is set to be presented at the upcoming Robotics: Science and Systems (RSS) conference in 2026. This release marks a significant step in the field of AI, reflecting Galbot's commitment to advancing technology and contributing to the global AI community.
PanDaily.com By [email protected] (Pandaily) Apr 30, 2026 News
ShengShu Technology has unveiled Motubrain, an innovative robotic brain designed to integrate perception and action seamlessly. This new technology aims to surpass traditional Variable-Length Architecture (VLA) models in performance on global benchmarks. The announcement was made recently, showcasing ShengShu's commitment to advancing robotics and artificial intelligence. By creating a hardware-agnostic solution, Motubrain allows for greater flexibility and efficiency in robotic applications, potentially transforming various industries that rely on automation and intelligent systems. The development of Motubrain reflects the growing demand for more sophisticated robotic technologies that can adapt to diverse environments and tasks.
HumanoidsDaily By [email protected] (Humanoids Daily Staff) Apr 29, 2026 WAM China ShengShu Technology world-model MotuBrain
Motubrain has emerged as a leader in the field of embodied artificial intelligence, achieving top rankings in two global benchmarks that assess its capabilities in understanding and generating information about the world. This significant milestone was reached in October 2023, highlighting the company's innovative approach to AI technology. By redefining the standards of embodied AI, Motubrain aims to enhance the interaction between humans and machines, making AI systems more intuitive and responsive. The advancements made by Motubrain are expected to have a profound impact on various industries, paving the way for more sophisticated applications of AI in everyday life.
RoboticsTomorrow.com Apr 29, 2026
Mimic Robotics, a company based in Zurich, has unveiled its innovative "pixel-to-action" architecture, which is designed to transform the current landscape of artificial intelligence by moving away from traditional static vision-language models. This release, which includes both the code and accompanying research, marks a significant shift towards utilizing dynamic video-based foundations. The initiative aims to enhance the capabilities of AI systems, enabling them to better interpret and respond to visual information in real-time. By sharing this technology, Mimic Robotics seeks to foster advancements in the field and encourage further exploration of video-based AI applications.
HumanoidsDaily By [email protected] (Humanoids Daily Staff) Apr 14, 2026 Mimic Robotics Europe open-source ETH Zurich
AGIBOT has unveiled its latest innovation, Genie Envisioner 2.0, a significant advancement in embodied artificial intelligence. This new platform transforms traditional world models into scalable and interactive simulators, enabling robots to learn and optimize their performance within environments generated by these models. The launch, which took place recently, signifies a pivotal shift from merely understanding the world to actively engaging with it, enhancing the robots' training capabilities and facilitating real-time interactions. This development aims to improve the efficiency and effectiveness of robotic learning processes, positioning AGIBOT at the forefront of AI technology.
agibot.com By AgiBot Apr 10, 2026 Embodied AI World Models Robotics Simulation Technology Artificial Intelligence
In May 2026, the Journal of Field Robotics published a significant study exploring advancements in robotic technology. Researchers from various institutions collaborated to examine the latest innovations in field robotics, focusing on their applications in agriculture, search and rescue operations, and environmental monitoring. The study highlights how these robotic systems are designed to enhance efficiency and safety in challenging environments, addressing the growing demand for automation in various sectors. By employing cutting-edge artificial intelligence and machine learning techniques, the researchers demonstrated how robots can perform complex tasks with increased precision and reliability. This research aims to provide insights into the future of robotics, emphasizing the importance of continued development in this field to meet societal needs and improve operational capabilities.
JournalofFieldRobotics By Wenhao Sun, Sai Hou, Zixuan Wang, Bo Yu, Shaoshan Liu, Xu Yang, Shuai Liang, Yiming Gan, Yinhe Han Apr 08, 2026 RESEARCH ARTICLE
Recent advancements in large language models (LLMs) have led to significant improvements in various domains, particularly in coding. However, a notable limitation remains: LLMs struggle to play video games effectively. Despite some successes, such as Gemini 2.5 Pro defeating Pokémon Blue in May 2025, these models often perform poorly compared to human players, making frequent mistakes and requiring specialized software to assist them. Julian Togelius, director of New York University’s Game Innovation Lab and co-founder of AI game-testing firm Modl.ai, discussed these challenges in a recent interview with IEEE Spectrum. He highlighted that while coding resembles a well-structured game with clear tasks and immediate feedback, video games present a more complex landscape that LLMs have yet to navigate successfully. Unlike games like chess or Go, which have been mastered by AI through retraining, video games vary significantly in mechanics and input requirements, complicating the development of a general game AI. Togelius pointed out that the lack of comprehensive benchmarks for video games further hinders LLMs' performance. While benchmarks have driven improvements in coding, the diverse nature of video games makes it difficult to establish similar metrics. He noted that current LLMs perform poorly even compared to basic algorithms in gaming contexts, primarily due to insufficient training data and challenges in spatial reasoning. Despite their coding capabilities, LLMs cannot engage in the iterative process of game development, which involves testing and refining gameplay. This disparity raises questions about the future of AI in mastering video games and its implications for broader AI applications.
IEEESpectrumAI By Matthew S. Smith Mar 29, 2026 Llms Artificial-intelligence Video-games
Rhoda AI, a technology company based in Palo Alto, has emerged from stealth mode with the announcement of a $450 million Series B funding round. This significant investment will support the development of its innovative "Direct Video-Action" framework, which leverages hundreds of millions of internet videos to educate robots on the principles of physics. The funding aims to enhance the company's capabilities in artificial intelligence and robotics, positioning Rhoda AI at the forefront of technological advancements in these fields. The announcement marks a pivotal moment for the company as it seeks to revolutionize how machines learn and interact with the physical world.
HumanoidsDaily By [email protected] (Humanoids Daily Staff) Mar 10, 2026 DVA US rhoda-ai
The recent deployment of foundation models across various industries marks a significant transformation in the application of artificial intelligence. This shift is not merely about introducing a singular, powerful solution to replace existing workflows; rather, it emphasizes a complex and iterative process of aligning advanced technology with engineering practices and market demands. On Thursday, the founders of two prominent Chinese AI unicorns discussed these developments, highlighting the importance of adapting AI agents to meet specific industry needs. Their insights reflect a broader trend in the tech landscape, where the focus is on integrating AI solutions that enhance rather than disrupt current operational frameworks. This evolution is expected to drive innovation and efficiency, as businesses seek to leverage AI in a way that complements their existing systems.
TechNode.com By TechNode Staff Sep 06, 2024 Events AI Ant GroupRSF defines a common language for robot service capability, lifecycle operations, certification pathways, and service-provider networks.
Daily robotics news, in-depth analysis, conference highlights, and discussions with professionals worldwide.