Market and Business News Research and Academia

CASIA's Peiyan Li Brings Two VLA Papers to IROS 2026

CASIA Ph.D. candidate Peiyan Li co-authored two IROS 2026 papers on VLA models for long robot tasks. We also cover his BridgeVLA and SpatialVAM work.

Share
CASIA's Peiyan Li Brings Two VLA Papers to IROS 2026
Share

The Ph.D. candidate co-authored two papers on training vision-language-action models for long, multi-step tasks. His broader research looks at how robots can learn manipulation from a small number of demonstrations.

Peiyan Li, a Ph.D. candidate at the Institute of Automation, Chinese Academy of Sciences (CASIA), attended IROS 2026 in Pittsburgh as a co-author of two papers. Both deal with vision-language-action (VLA) models, which take camera images and a text instruction and output robot actions. Both also target the same weakness: these models still fail often on long tasks with many steps.

Most of Li's own work addresses a related cost. Training a manipulation policy usually takes many human demonstrations, and collecting them is slow and expensive. His papers look for ways to get good results from far fewer.

The IROS 2026 Papers

The two papers came from Li's group at CASIA, with Prof. Yan Huang and Prof. Liang Wang as senior authors. Both were presented in VLA lightning-talk sessions.

HCoT-VLA (Paper #2265, session "Extending Vision-Language-Action Model Capabilities") adds a reasoning step before the model acts. Some earlier approaches have the model write its reasoning out as text. HCoT-VLA instead encodes subtask goals, grasp points and inverse kinematics as numerical vectors inside the model, which carries more information than text does. The team also built a semi-automated pipeline to label this extra data. The model beat its baselines on the LIBERO-Plus long-horizon benchmark and on real-robot tasks, with the largest margins when the scene was disturbed or differed from the training data.

StaKe (Paper #2610, session "Architectural Advances in Vision-Language-Action Models") changes how a VLA model is fine-tuned. Standard training weights every timestep equally, and errors tend to pile up around the moments when the gripper opens or closes. StaKe reads the gripper states in the demonstration data and derives two extra training targets: which stage of the task the robot is in, and what the arm should do at the next gripper change. No manual labeling is needed, and the model used at run time is unchanged. Success rates rose by a relative 14% in bimanual simulation and 56% on a real Franka arm, and longer tasks gained the most.

Background

Li has been at CASIA's New Laboratory of Pattern Recognition (NLPR) since 2023, advised by Prof. Tieniu Tan. Before that he completed a B.Eng. (2019–2023) in the Artificial Intelligence honors program at Xi'an Jiaotong University, which was set up by Prof. Nanning Zheng. He works on general-purpose manipulation through data, model architecture and training methods, and is currently focused on pretraining robot foundation models at scale.

He has also spent time in industry. He interned at Tencent Robotics X in 2023, working on language-conditioned motion planning, and at ByteDance Seed Robotics from January 2024 to August 2025, working on manipulation and VLA models. Since January 2026 he has been in the Top Internship Talent Program at Xiaomi Robotics, where he was a core contributor to pretraining and simulation for Xiaomi-Robotics-1, a VLA model trained on more than 100,000 hours of real-robot data.

As first or co-first author, he has published at NeurIPS 2025, ECCV 2026 and IEEE RA-L, and has three papers accepted to NeurIPS 2026. He is also a co-author of a Nature Machine Intelligence study on VLA design. His BridgeVLA model won the COLOSSEUM Challenge at the CVPR 2025 GRAIL Workshop and the ARNOLD Challenge at the CVPR 2026 EAI Workshop. He was named an ICML Gold Reviewer in 2026. As an undergraduate, he was selected by People's Daily as one of its Outstanding National Scholarship Student Representatives, the only undergraduate from Xi'an Jiaotong University chosen that year. He also runs marathons and has raced in Xi'an, Shenzhen, Tianjin, Zhengzhou and Lanzhou.

BridgeVLA

BridgeVLA, Li's first-author NeurIPS 2025 paper, deals with a mismatch in many 3D robot models. The vision-language models they build on are trained on 2D images, while the robot works in 3D. BridgeVLA projects the 3D point cloud into several 2D views and outputs actions as 2D heatmaps, so the input and output are both images. It is also pretrained to locate named objects as heatmaps before it learns any robot task.

On the RLBench benchmark, average success went from 81.4% to 88.2%. On COLOSSEUM, which tests robustness to visual and physical changes, it went from 56.7% to 64.0%. On a real robot it reached 96.8% success across more than 10 tasks with three demonstrations per task.

image.png

BridgeVLA projects 3D input into several 2D views and predicts actions as heatmaps. Source: BridgeVLA project page.

SpatialVAM

SpatialVAM, accepted to NeurIPS 2026, adds prediction over time. It is built on a video generation model and predicts what the scene will look like next, both as multi-view video and as heatmap video, then turns those predictions into actions. With a small number of demonstrations and no task-specific pretraining, it outperformed the strongest baselines reported in the paper by about 22% on Meta-World, 15% on RoboCasa and 16% on real-robot tasks.

image.png

SpatialVAM predicts future multi-view video and heatmaps, then converts them into actions. Source: SpatialVAM project page.

What It Means for Industry

For companies putting robots into factories, warehouses or homes, collecting demonstrations for each new task is a large part of the deployment cost. Methods that cut the number needed from hundreds to a handful lower that cost directly. The two IROS papers he co-authored work on a different problem, reliability on long tasks, which is often what keeps a robot that works in the lab from doing a real job. His internships at Tencent, ByteDance and Xiaomi have put him on the industrial side of both questions as well as the research side.

Paper Information

Paper

Venue

Links

HCoT-VLA (Hybrid Chain-of-Thought Reasoning for VLA Models)

IROS 2026, Paper #2265

2026.ieee-iros.org/program/paper-index

StaKe (Improving VLA Fine-Tuning with Structured Stage and Keyframe Supervision)

IROS 2026, Paper #2610

2026.ieee-iros.org/program/paper-index

BridgeVLA

NeurIPS 2025

bridgevla.github.io  |   arxiv.org/abs/2506.07961

SpatialVAM

NeurIPS 2026

arxiv.org/abs/2604.03181

BridgeVLA++

arXiv 2026

bridgevla-plus.github.io

About the researcher: Peiyan Li, Ph.D. candidate, NLPR, Institute of Automation, Chinese Academy of Sciences. Homepage: lpy1219.github.io  | Email: [email protected]

RobotToday Initiative

Robotics needs a service framework.

RSF defines a common language for robot service capability, lifecycle operations, certification pathways, and service-provider networks.

Share
Written by
Kelly Stone - Associtae Editor

Kelly Stone is an Associate Editor focused on industrial technology, covering robotics, automation systems, and AI applications. Her reporting emphasizes company funding, market structure, and emerging industry trends. She has three years of experience in technology media.

inJoin the RobotToday community on LinkedIn

Daily robotics news, in-depth analysis, conference highlights, and discussions with professionals worldwide.