Out of 1,933 papers presented in Pittsburgh, ten made the final round for IROS's top two prizes. Four came from South Korea, a country with about 6% of the program. Here is what each one does, and what it means for industry.
IROS 2026 in Pittsburgh put 1,933 papers on its program. Only ten reached the final round for the conference's two headline prizes, the Best Paper Award and the Best Student Paper Award.
The country with the most finalists was not the United States or China. South Korea placed four of the ten, from DGIST, POSTECH, Korea University of Technology and Education (KOREATECH), and a team spanning Hyundai's software unit 42dot, Seoul National University and Samsung Research. That is 40% of the shortlist from a country with about 6% of the program, according to a RobotToday analysis of the official paper index.
Mainland China placed three finalists, the U.S. two, and Germany one. According to reports from the Sept. 30 awards lunch, DGIST's LT-Mem took Best Paper and the University of Pennsylvania's quadrotor formation work took Best Student Paper.
The topics run from compressing vision-language-action models for on-robot chips to handling living tissue with sound. Below, each finalist in about 150 words.
How the IROS 2026 paper awards work
IROS does not crown a single "best" paper. This year it gave 10 paper awards, drawn from 32 finalists who presented in dedicated award sessions on Monday and Tuesday, Sept. 28–29.
The headline pair. Best Paper and Best Student Paper share one shortlist of 10 finalists, presented across three sessions. That shortlist is the subject of this article.
Eight category awards. Separate shortlists of two to four papers each cover application, mobile manipulation, safety and rescue, entertainment, cognitive robotics, mechanisms and design, industrial robotics, and agricultural robotics. Most are sponsored by a company or society, such as YANMAR, OMRON SINIC X and KROS.
When winners are named. All awards are presented at the Awards Lunch that closes the technical program, held Wednesday, Sept. 30.
Being a finalist is itself a distinction: these ten were picked from roughly 1,585 accepted conference papers.
The ten finalists
The two reported winners come first, then the other eight in program order.
1. LT-Mem: a long-term memory for robots — Best Paper (reported)
Yumin Lee, Hyoseok Ju, Giseop Kim · DGIST · South Korea
Paper: LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding
Robots that work in the same building for months face a memory problem. Their maps are either overwritten at each visit, erasing history, or stored as snapshots that lose track of which object is which. The authors call this temporal amnesia. LT-Mem gives every object a persistent identity across visits, built on a multi-session SLAM backbone. For each object it then decides whether to overwrite, hold, or keep several hypotheses, depending on how often that object tends to move. Memory is split into three layers: live state, recent changes, and long-term history. A robot can then answer questions such as where a chair has been over several weeks. The team also released LT-VQA, a multi-session dataset with temporal questions. LT-Mem beat baselines on every metric while using about a tenth of the tokens.
Industry angle: Security, facility and service robots patrol the same sites daily. This is a structured alternative to stuffing their history into a language model's context window.
2. Tight quadrotor formations with physics-aware learning — Best Student Paper (reported)
Pei-An Hsieh, Fengjun Yang, Nikolai Matni, M. Ani Hsieh · University of Pennsylvania · USA
Paper: Flatness-Preserving Residual Learning for Real-Time Tight Quadrotor Formation Flight
When quadrotors fly close together, downwash from one drone can push another off course and cause a collision. The Penn team learns a correction to the physics model that captures these aerodynamic interactions. The key constraint is that the correction keeps the whole multi-drone system differentially flat, a property that allows a simple, easily tuned feedback controller. That controller cancels the disturbance before it builds up. In hardware tests, average tracking error fell 31% against the baseline controller. Performance matched nonlinear model predictive control at roughly a tenth of the computation. The team reports stable tight-formation flight with less than 30 seconds of training data and a 5-millisecond control loop.
Industry angle: Close formations are central to drone light shows, cooperative payload carrying and inspection swarms. This method fits the small flight computers those fleets already use.
3. Shallow-π: making vision-language-action models fit on robot chips
Boseong Jeon (42dot), Yunho Choi (Seoul National University), Taehan Kim (Samsung Research) · South Korea
Paper: Shallow-π: Knowledge Distillation for Flow-Based VLAs
Vision-language-action (VLA) models are powerful but slow to run on a robot's own computer. Most speed-up work trims the number of visual tokens. This team instead removes whole transformer layers, using knowledge distillation on both the vision-language backbone and the flow-based action head of a π-style model. Depth drops from 18 layers to 6. Inference runs more than twice as fast, with less than one percentage point lost in task success on standard manipulation benchmarks. That is the best result reported among compressed VLA models. The team validated it in real-world manipulation on NVIDIA Jetson Orin and Jetson Thor across several robot platforms, including humanoids.
Industry angle: This is the finalist closest to a product roadmap. Humanoid and mobile-manipulator makers need VLAs running on embedded chips, not cloud GPUs. 42dot is Hyundai Motor Group's software subsidiary.
4. Handling bioprinted tissue without touching it
SeungTaek Hong, Jaewon Byun, Daekeun Kim, Jinah Jang, Keehoon Kim · POSTECH · South Korea
Bioprinted tissue modules are mostly hydrogel: wet, soft and fragile. Conventional tools deform them, stick to them and risk contamination. The POSTECH team built a gripper that holds the modules without contact, using inverted near-field acoustic levitation. Sound pressure creates a thin air gap that lifts centimeter-scale hydrogel blocks. A camera-based localization system and an autonomous pipeline then place each module. Cell viability tests showed that acoustic handling kept cells intact. Compared with human operators, the robot was more consistent and caused no damage, though it worked more slowly. The team demonstrated autonomous assembly of several 1-cm phantom tissue cubes with different stiffness.
Industry angle: It is an early building block for automated tissue transplantation and organ fabrication. The same non-contact approach could apply to other delicate goods.
5. Streaming Gaussian encoding for 4D occupancy tracking
Maximilian Luz, Abhinav Valada (University of Freiburg); Thomas Nürnberg, Yakov Miron (Bosch) · Germany
Paper: Streaming Gaussian Encoding for 4D Panoptic Occupancy Tracking
Camera-only 4D panoptic occupancy tracking works out, from a car's cameras, which 3D space is occupied, by what, and which object it is over time. Current methods rebuild the 3D scene at every frame, so geometry flickers when objects are hidden. This team keeps a persistent scene memory instead. A fixed set of latent Gaussian queries is carried forward with the vehicle's motion and refreshed under a confidence budget. Depth supervision trains each Gaussian's opacity to stand in for visibility, so confidence accumulates over time. The method set a new state of the art on extended nuScenes and Waymo occupancy benchmarks. It adds negligible compute and plugs into existing pipelines.
Industry angle: Camera-only perception is the main cost lever in driver assistance and autonomous driving. Bosch's co-authorship points to a short path from paper to Tier-1 product.
6. Decoding Mandarin tones from brain and muscle signals
Yongjie Zou, Jiawei Ju, Chengyu Li · Lingang Laboratory, Shanghai · China
Mandarin is a tonal language, so a speech brain-computer interface must decode tone, not just syllables. This team fuses noninvasive brain signals (EEG) with muscle signals (EMG), for both spoken and silent speech. The obstacles are differences between people, mismatches between the two signal types and between spoken and silent modes, easily confused tones, and heavy electrode counts. Their model, GEM-FOCUS, uses learnable gates to balance the two signals, alignment methods to generalize across people, and a module that selects the most useful channels. On a 10-person dataset it performed competitively with only 20 EEG and 5 EMG channels. The team also built a calibration-free, real-time online decoder.
Industry angle: Too many electrodes and per-user calibration are what keep wearable speech interfaces out of daily use. Both are addressed here, with uses from assistive communication to silent robot commands.
7. Helping robots see glass
Jiamin Zheng, Guangcheng Chen, Hong Zhang (SUSTech); Jingwen Yu (HKUST and SUSTech) · China
Paper: Enhancing Glass Surface Reconstruction via Depth Prior for Robot Navigation
Glass walls and doors corrupt depth sensors. Light passes through or reflects, so a robot sees holes or phantom surfaces and may drive into them. Depth foundation models such as Depth Anything 3 infer plausible structure from images, but they lack true metric scale. This team uses the foundation model as a structural guide and aligns it to the raw sensor depth with a robust local fitting step. The fit ignores the corrupted glass readings and recovers real-world scale, with no training required. The team also released GlassRecon, an RGB-D dataset with geometric ground truth for glass regions. The method beat leading baselines, especially when sensor depth was badly corrupted.
Industry angle: Offices, malls, airports and hospitals, where service robots work, are full of glass. A training-free fix can be added to existing navigation stacks.
8. DAWN: quadruped parkour with noisy depth cameras
Yohan Choi, Min-Jun Kim, Jin-Sung Kim, Yong-Jae Kim, Youn-Hee Han · Korea University of Technology and Education · South Korea
Paper: DAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models
Legged robots that see with depth cameras are usually trained on clean simulated depth. Real cameras are noisy, so teams fall back on hand-tuned filters that are rarely published. DAWN builds noise robustness into a world model instead. During training, the encoder receives noisy depth but must reconstruct clean depth, and contrastive learning aligns the noisy and clean internal states. Nothing needs tuning at deployment, and inference costs no more than existing world-model methods. On a Unitree Go1, with no filter calibration, the robot climbed 18-cm stairs, cleared 70-cm gaps and mounted 45-cm steps. Code and videos are public.
Industry angle: Sensor noise is a daily obstacle for quadruped and humanoid teams using low-cost depth cameras. Removing the filter-tuning step shortens field deployment.
9. VTAP Gripper: dexterity from sensing, not finger count
Yuhao Zhou, Sheeraz Athar, Zhixian Hu, Juan Wachs, Yu She (Purdue University); Binghao Huang, Yunzhu Li (Columbia University) · USA
Instead of a many-jointed humanoid hand, this team built a three-finger gripper. Its compliant, reconfigurable fingers carry tactile sensor arrays. The actuated palm combines a camera, for locating objects at a distance, with tactile sensing once contact is made. For teleoperation, a staged, gesture-based retargeting method maps human hand motion onto the three fingers. The gripper handled reactive grasping of standard and fragile objects. It reoriented a syringe in hand and pressed its plunger, separated clustered objects as small as 3 mm, and completed peg-in-hole insertion guided by vision and touch.
Industry angle: It shows that dexterity can come from sensing and finger-palm coordination rather than joint count. Gripper makers weighing cost against capability should take note, and the design doubles as a data-collection rig for robot learning.
10. Automatic gait switching for a magnetic soft robot
Bin Wang, Shengming Luo, Jiansheng Du, Yuanbiao Ma, Xuanyu An, Qianqian Wang · Southeast University · China
Paper: Autonomous Multimodal Locomotion Control of a Magnetic Strip Soft Robot
Magnetic soft robots can roll, crawl and climb through confined spaces. Switching between those modes usually needs custom electromagnetic coils, which limit the workspace, or a person turning a magnet by hand, which is imprecise. This team automates it for a 3D-printed magnetic strip robot. A camera tracks the robot's outline, and its aspect ratio serves as a live signal of which mode it is in. Different magnet trajectories then drive different gaits. In closed-loop tests the robot passed through narrow gaps, pushed boxes, climbed stairs and delivered items over obstacles. For delivery, a state machine switched between rolling and crawling on its own.
Industry angle: Reliable, repeatable control is a prerequisite for using magnetic soft robots in confined-space delivery and inspection.
By the numbers
South Korea's 117 program papers were 6% of the total, yet the country produced 4 of the 10 finalists.

RobotToday analysis of the official IROS 2026 Paper & Author Index (1,933 papers) and Awards page · a paper counts for every country with at least one author there
The U.S. and mainland China together had a hand in about two-thirds of all program papers but placed five finalists between them. All four of South Korea's award-session papers at IROS 2026 were in this headline shortlist.
Leave a comment