Besides its two headline prizes, IROS 2026 named finalists for eight category awards. The 22 papers are closer to products than to theory. Here is what each one does, and what it means for industry.
The Best Paper Award gets the headlines, but IROS 2026 in Pittsburgh handed out eight more paper awards, each with its own shortlist and, in most cases, its own industry sponsor. Twenty-two papers made those shortlists.
They read like a map of where robotics is heading into products. Three finalists ran on Unitree's G1 humanoid: one carries a tray without spilling, one plays tennis rallies with people, and one turns high-level goals into whole-body action. Others target car factories, smartphone assembly lines, tomato greenhouses, the seafloor and the human gut.
By first author, mainland China placed eight finalists and the United States seven. Germany had two, and Hong Kong, Japan, the UK, Australia and the Netherlands one each. South Korea, which led the Best Paper shortlist with four finalists, had none here.
Below, each finalist in about 90 words, grouped by award. Winners confirmed by their teams are marked; IROS has not yet posted an official winners list.
1 Best Application Paper (ICROS) and Best Paper on Mobile Manipulation (OMRON SINIC X)
Two awards shared one session and one shortlist of four: ICROS's prize for work closest to real-world use, and OMRON SINIC X's prize for robots that move and manipulate at once.
1.1 SteadyTray: a humanoid that carries a tray without spilling — Winner, Mobile Manipulation (announced by the team)
Anlun Huang, Michael C. Yip et al. · UC San Diego · USA
Every step a humanoid takes sends a jolt to its hands. The UCSD team's ReST-RL splits the job in two: a base policy handles walking, and a residual module actively cancels gait-induced wobble at the end-effector. In simulation it reached a 96.9% success rate in variable-speed tracking and 74.5% under external pushes, beating end-to-end baselines. It transferred zero-shot to a Unitree G1 carrying various objects.
Industry angle: Serving, hospitality and logistics use cases all depend on humanoids carrying things without spilling.
1.2 ULTRA: from motion replay to goal-driven humanoid control
Xialin He, Liang-Yan Gui et al. · University of Illinois Urbana-Champaign, with Shanghai Jiao Tong and Tsinghua · USA
Paper: ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation
Most humanoid controllers replay predefined motion clips. ULTRA generates whole-body behavior from perception and a high-level goal instead. Physics-based retargeting turns large motion-capture libraries into usable humanoid data. One multimodal controller then accepts either detailed motion references or sparse task commands, from clean motion-capture state down to noisy onboard camera input. On a real Unitree G1, it performed goal-directed loco-manipulation from egocentric vision.
Industry angle: Going from "replay this clip" to "do this task" is the step humanoid makers need for general-purpose work.
1.3 Designing a capsule robot that ultrasound can track
Yizhao Qian (CUHK), Max Q.-H. Meng (SUSTech), Li Liu (Great Bay University) et al. · Hong Kong / mainland China
Paper: Acoustic-Aware Texture Optimization for Ultrasound-Based Capsule Pose Estimation
Wireless capsule robots for the gut need to know their position and orientation, and ultrasound can see them without radiation. But a smooth capsule looks the same from many angles. This team solved the problem at its source by designing the capsule's surface texture, optimized through a differentiable model of ultrasound imaging. Error fell to about 1.2 mm and 1.4°, versus 15–21° for plain or helical surfaces. In an ex-vivo pig stomach, tracking error stayed under 0.33 mm.
Industry angle: It moves closed-loop control of capsule endoscopes a step closer to the clinic.
1.4 PGFlowNav: real-time navigation for a robotic fish
Lianyi Yu, Min Tan et al. · Institute of Automation, Chinese Academy of Sciences, with Peking University · China
A bionic robotic fish must steer in real time through turbulent water. Diffusion-based navigation policies model uncertainty well but are slow, because they sample step by step. PGFlowNav uses conditional flow matching, which produces an action in a single pass. A variational autoencoder adds task-aligned priors that keep actions within what the fish can physically do. Real-world tests showed stable, goal-directed swimming under real-time control.
Industry angle: Fast generative policies on small onboard computers suit underwater inspection and environmental monitoring.
2 Best Paper on Safety, Security, and Rescue Robotics
Named in memory of Motohiro Kisoi, this award covers robots for hazardous, disaster and security work. Two of its three finalists are drones.
2.1 Using touch to navigate on the seafloor
Michele Grimaldi, Yvan R. Petillot et al. · Heriot-Watt University, with JAMSTEC · UK / Japan
Paper: Contact-Aided Factor-Graph Localization for Underwater Sampling
An underwater vehicle sampling the seafloor flies low over flat, featureless sand, where cameras struggle and navigation drifts. This team turns physical contact into a navigation aid. Each time the suction manipulator touches an object, the event becomes a high-confidence constraint in a factor graph, acting like a loop closure without visual place recognition. Visual odometry is weighted by its own reliability. Tank, harbor and simulation tests showed less drift and more accurate returns to targets.
Industry angle: Seabed sampling for science, mineral surveys and subsea inspection all need precise relocation.
2.2 AUG: drone routes that keep LiDAR navigation healthy
Enguang Feng, Xiwang Dong et al. · Beihang University, with Harbin Institute of Technology · China
Drones without GPS often rely on LiDAR-inertial odometry, which can fail over open, feature-poor terrain. AUG plans routes that keep the sensor pointed at informative structure. It combines a lightweight topological graph for long-range routing, a local trajectory optimizer and a perception-aware yaw planner. It succeeded in every run across three 1.2 km × 1.2 km simulated maps using under 1 GB of memory, and averaged 2.24% drift in ten 200-meter outdoor flights.
Industry angle: GPS-denied flight is routine in disaster response, mines and contested airspace.
2.3 NEUROSYMLAND: drone landing decisions you can audit
Weixian Qian, Xi Zheng et al. · Macquarie University, with UC Santa Barbara and Anitron · Australia
Paper: NEUROSYMLAND: Neuro-Symbolic Landing-Site Assessment for Robust and Edge-Deployable UAV Autonomy
Picking a safe landing spot is a common failure point for autonomous drones. NEUROSYMLAND builds a probabilistic scene graph from the onboard camera, then checks candidate sites against explicit rules for flatness, clearance and consistency, so each decision is explainable. It succeeded in 61 of 72 simulated scenarios, against 37 to 57 for four baselines. Hardware-in-the-loop tests showed the rule-checking adds little latency on edge hardware. Full onboard closed-loop landing is left to future work.
Industry angle: Regulators and operators increasingly want safety decisions they can inspect.
3 Best Entertainment and Amusement Paper
This award recognizes robots for play, performance and sport. Its four finalists were the most visual of the conference: piano, monkey bars, juggling and tennis.
3.1 PianoFingering-1.5K: teaching robot hands from piano videos
Ruoyu Wang, Zhen Kan et al. · University of Science and Technology of China · China
Paper: PianoFingering-1.5K: A Large-Scale Expert Piano Fingering Dataset for Dexterous Robot Learning
Robot hands learning piano need to know which finger plays which key, and such labels are scarce. The USTC team built an automated pipeline that extracts expert fingering from online performance videos, using a probabilistic model to handle uncertain key contacts. The dataset covers more than 1,500 pieces, 1.1 million annotated frames and 2.2 million notes, at over 98% accuracy. Policies trained on it played more accurately, with more natural hand poses, than those using computed fingering.
Industry angle: Mining human video is a scalable way to train dexterous hands, well beyond music.
3.2 A life-sized robot that swings on monkey bars
Ayumu Iwata, Kei Okada et al. · The University of Tokyo · Japan
Paper: Robust Brachiation on a Life-Sized Dual-Arm Robot Using Waypoint-Guided Reinforcement Learning
Brachiation is moving hand over hand along bars, as gibbons do. The Tokyo team taught it to a life-sized dual-arm robot with Waypoint-Guided Reinforcement Learning: a few sparse waypoints for the hand trajectory guide learning, and the whole-body motion is learned. Rewards for task success and mechanical energy, plus training designed for sim-to-real transfer, delivered forward progress with stability. Tests on varied monkey bars, in simulation and hardware, showed robust swinging, including recovery from failures.
Industry angle: Arm-based locomotion lets robots cross spaces with no footholds, such as overhead structures.
3.3 Catch, Throw, Repeat: a robot that juggles with people
Jonathan Rainer Lippert, Jan Peters, Alap Kshirsagar et al. · TU Darmstadt, with Justus Liebig University Giessen · Germany
Paper: Catch, Throw, Repeat: Planning for Human–Robot Partner Juggling
Juggling with a partner tests timing, prediction and contact all at once. This team built a planning and control system that lets a robot catch and throw balls in shared multi-ball patterns with a person, combining predictive ball tracking, online trajectory optimization and a coordination state machine. In a study with eight people, from beginners to experts, all beat previously reported results within ten minutes. One reached 20 consecutive robot catches in a shared three-ball cascade, five times the previous record.
Industry angle: Fast, reliable hand-offs between people and robots apply well beyond juggling.
3.4 LATENT: humanoid tennis from imperfect human data — Winner (announced by co-author Galbot)
Zhikai Zhang, Li Yi et al. · Tsinghua University, with Galbot, Peking University and Shanghai AI Lab · China
Paper: Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data
LATENT teaches a humanoid to play tennis from imperfect human data: short fragments of basic strokes rather than full match recordings, which are far easier to collect. The team corrects and combines those fragments into a policy that hits incoming balls across a wide range of conditions and returns them to chosen targets, with natural, human-like motion. Deployed on a Unitree G1, it sustained multi-shot rallies with human players.
Industry angle: Humanoids can learn fast, athletic skills without costly, perfect motion capture.
4 Best Paper on Cognitive Robotics (KROS)
Sponsored by the Korea Robotics Society, this award targets robots that reason. All three finalists combine vision-language models with structured 3D information.
4.1 VL-Nav: neuro-symbolic navigation from abstract instructions — Winner (announced by the team)
Yi Du, Chen Wang et al. · University at Buffalo · USA
Paper: VL-Nav: A Neuro-Symbolic Approach for Reasoning-Based Vision-Language Navigation
VL-Nav helps a mobile robot follow complex, abstract instructions in large, unseen spaces. It pairs a vision-language model with symbolic structure: a 3D scene graph and image memory help the model break tasks down and replan, while a symbolic heuristic steers exploration away from aimless wandering and repeat travel. On DARPA TIAMAT Challenge tasks it reached 83.4% success indoors and 75% outdoors, and 86.3% in real-world tests that included a 483-meter run.
Industry angle: Neuro-symbolic designs offer a practical middle path between end-to-end models and hand-built pipelines for service and security robots.
4.2 SoftNav: feeding 3D scenes straight into a VLM
Yi Wu, Guang Li et al. · Zhejiang University, with Shandong University · China
Paper: SoftNav: Injecting 3D Scene Tokens into VLMs for Embodied Navigation
Navigation systems usually describe a 3D scene to a vision-language model in text. SoftNav argues that loses information and instead injects one learned token per detected object or frontier directly into the model's hidden layers. With the 3D encoder and the VLM frozen, it needs only about 1,200 training samples and 17 million trainable parameters. It beat prior methods on the HM3D-OVON object-navigation benchmark and transferred zero-shot to other benchmarks and a real robot.
Industry angle: Cheap adaptation of off-the-shelf VLMs lowers the cost of building navigation for new robots.
4.3 GeoVLA: giving vision-language-action models a sense of 3D
Lin Sun, Jiale Cao et al. · Tianjin University, with Dexmal, Tsinghua and MEGVII · China
Paper: GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
Most vision-language-action models see only flat 2D images. GeoVLA adds 3D geometry: depth maps become point clouds, which a dedicated encoder turns into features fused with the vision-language stream in a 3D-enhanced action expert. It set state-of-the-art results on the LIBERO and ManiSkill2 simulation benchmarks and proved more robust than 2D baselines in real tasks requiring height adaptation, scale awareness and viewpoint changes.
Industry angle: Poor spatial awareness is a common failure point for VLA-driven manipulation in real workplaces.
5 Best Paper on Robot Mechanisms and Design (Cornerstone Robotics)
This award honors hardware design. The three finalists cover a new tactile sensor, a wearable robot and a finger joint.
5.1 GelSphere: a tactile sensor that rolls in any direction
Seoyeon Lee, Wenzhen Yuan et al. · University of Illinois Urbana-Champaign · USA
Vision-based tactile sensors capture fine surface detail but are usually fixed pads that wear out when dragged. GelSphere is a spherical sensor that rolls freely. Steel balls between the gel and the rigid housing form a bearing layer, so the sphere can roll across a surface in any direction while its internal camera images the contact. It streams tactile images over Wi-Fi and reconstructs large surfaces online, keeping geometric accuracy under multi-directional rolling.
Industry angle: Rolling tactile scanning could speed up surface inspection, from manufactured parts to infrastructure.
5.2 A robotic zipper for clothing that puts itself on
Amanda Weckerly, Allison M. Okamura, Cynthia Sung et al. · Stanford University, with the University of Pennsylvania · USA
Paper: Design, Modeling, and Performance of a Robotic Zipper for Self-Donning Clothing
Dressing is hard for people with limited mobility or dexterity. This team designed clothing that closes itself: a small snap-on robot with a motor-driven gear runs along the zipper teeth. They found the required torque scales with garment mass, from 0.045 to 0.11 N·m per kilogram depending on orientation, and that the robot can traverse curves as tight as 10 mm. Those limits guide seam design. Demonstrations used one or several robots on garments of varying complexity.
Industry angle: Aging populations are driving demand for adaptive garments and assistive wearables.
5.3 PDS Joint: a compliant finger joint for dexterous hands
Haoyang Li, Yufeng Yue et al. · Beijing Institute of Technology · China
Paper: PDS Joint: A Parametric Double-Spiral Joint Tailored for Dexterous Hands
Compliant joints make robot hands safer, but they must bend through large, human-like ranges while resisting motion in the wrong directions. The PDS joint uses a parametric double-spiral shape whose geometry sets stiffness separately for bending, side-to-side and twisting motion. Embedded inductive sensors, calibrated with a learned model, track joint state; for the hardest motion, error fell 59.2% versus conventional curve fitting. In an open-source hand, it succeeded in every adaptive grasping test.
Industry angle: Dexterous-hand makers need joints that are compliant, durable and self-sensing at once.
6 Best Paper for Industrial Robotics Research for Applications
The industrial award had the shortest list: two finalists, both built in or for production settings, one in automotive and one in electronics.
6.1 Picking Bins Empty: bin picking that finishes the job
Florian Töper et al. · Mercedes-Benz AG, with DLR and Fraunhofer IAO / Reutlingen University · Germany
Factory bin-picking robots often stall before a bin is empty, when parts are occluded or a planned grasp fails. This team combines both main approaches in a four-tier hierarchy. A model-based pipeline does most of the work; a model-free agent steps in to break deadlocks and discover new grasp points; and online self-learning ranks grasps from gripper feedback, cutting manual tuning. On three automotive parts, bin clearance rose from 50.9% to 100%.
Industry angle: Full clearance without human help is what production lines need before trusting bin picking.
6.2 MCFR: precision connector assembly for smartphones
Guanghui Shen, Dan Wu et al. · Tsinghua University · China
Inserting board-to-board connectors inside phones demands extreme precision across many connector variants. MCFR uses the Segment Anything Model to mask out distracting background, a vision-transformer network for coarse pose estimation and a photometric refinement step for pixel-level alignment. Tested on a multi-variant dataset, a batch-insertion testbed and a real smartphone assembly task, it reached a 99.25% physical insertion success rate.
Industry angle: Connector insertion is still one of the most manual steps in electronics assembly.
7 Best Paper on Agri-Robotics (YANMAR)
Sponsored by agricultural machinery maker YANMAR, this award covers farm robotics. The finalists tackle 3D plant modeling, night work and pest scouting.
7.1 SSL-Semantic-NBV: a robot that learns where to look at plants
Jianchao Ci, Gert Kootstra et al. · Wageningen University & Research, with China Agricultural University · Netherlands
To measure crops, a robot often needs 3D models of specific plant parts rather than the whole scene. SSL-Semantic-NBV chooses the next camera view that best reveals targets such as tomato fruits and nodes. It trains itself through robotic self-supervision, with synthetic plant models for pretraining, so it needs little human labeling. It improved reconstruction efficiency by 47–114% over non-NBV methods in simulation, and by 8–16% in real-world tests.
Industry angle: Automated phenotyping and harvest planning both rely on targeted, occlusion-aware 3D views.
7.2 AgriNight: teaching farm robots to see at night
Robel Mamo, Taeyeong Choi et al. · Kennesaw State University, with the University of Lincoln · USA / UK
Farm robots could monitor, harvest and detect pests around the clock, but labeled nighttime images are scarce. This team converts daytime field images into near-infrared night images without paired examples, using CLIP to keep the meaning of each region consistent and a mask for the limited reach of night lighting. Daytime labels can then train nighttime perception. The team also released AgriNight, the first benchmark for nighttime agricultural navigation, and ran a physical robot autonomously at night.
Industry angle: More operating hours is a direct productivity gain for farm automation.
7.3 STEMbot: a climbing robot that scouts under the canopy
Zachary Charlick, Dmitry Berenson et al. · University of Michigan · USA
Paper: STEMbot: A Compliant Robot for Under-Canopy Plant Navigation
Many crop pests hide under leaves and on stems, out of view of drones and ground rovers. STEMbot is a small climbing robot that navigates beneath the canopy. It combines PIN-SLAM localization with a semantic map and a branch-aware planner that can reach hidden spots. It climbed stems 7 to 33 mm thick and navigated four different plants autonomously, with 3D reconstructions within 1 cm of an offline reference.
Industry angle: Earlier pest detection could cut crop losses and labor costs, especially in organic farming.
Leave a comment