The Ph.D. candidate co-authored two papers on training vision-language-action models for long, multi-step tasks. His broader research looks at how robots can learn manipulation from a small number of demonstrations.
Peiyan Li, a Ph.D. candidate at the Institute of Automation, Chinese Academy of Sciences (CASIA), attended IROS 2026 in Pittsburgh as a co-author of two papers. Both deal with vision-language-action (VLA) models, which take camera images and a text instruction and output robot actions. Both also target the same weakness: these models still fail often on long tasks with many steps.
Most of Li's own work addresses a related cost. Training a manipulation policy usually takes many human demonstrations, and collecting them is slow and expensive. His papers look for ways to get good results from far fewer.
The IROS 2026 Papers
The two papers came from Li's group at CASIA, with Prof. Yan Huang and Prof. Liang Wang as senior authors. Both were presented in VLA lightning-talk sessions.
HCoT-VLA (Paper #2265, session "Extending Vision-Language-Action Model Capabilities") adds a reasoning step before the model acts. Some earlier approaches have the model write its reasoning out as text. HCoT-VLA instead encodes subtask goals, grasp points and inverse kinematics as numerical vectors inside the model, which carries more information than text does. The team also built a semi-automated pipeline to label this extra data. The model beat its baselines on the LIBERO-Plus long-horizon benchmark and on real-robot tasks, with the largest margins when the scene was disturbed or differed from the training data.
StaKe (Paper #2610, session "Architectural Advances in Vision-Language-Action Models") changes how a VLA model is fine-tuned. Standard training weights every timestep equally, and errors tend to pile up around the moments when the gripper opens or closes. StaKe reads the gripper states in the demonstration data and derives two extra training targets: which stage of the task the robot is in, and what the arm should do at the next gripper change. No manual labeling is needed, and the model used at run time is unchanged. Success rates rose by a relative 14% in bimanual simulation and 56% on a real Franka arm, and longer tasks gained the most.
Background
Li has been at CASIA's New Laboratory of Pattern Recognition (NLPR) since 2023, advised by Prof. Tieniu Tan. Before that he completed a B.Eng. (2019–2023) in the Artificial Intelligence honors program at Xi'an Jiaotong University, which was set up by Prof. Nanning Zheng. He works on general-purpose manipulation through data, model architecture and training methods, and is currently focused on pretraining robot foundation models at scale.
He has also spent time in industry. He interned at Tencent Robotics X in 2023, working on language-conditioned motion planning, and at ByteDance Seed Robotics from January 2024 to August 2025, working on manipulation and VLA models. Since January 2026 he has been in the Top Internship Talent Program at Xiaomi Robotics, where he was a core contributor to pretraining and simulation for Xiaomi-Robotics-1, a VLA model trained on more than 100,000 hours of real-robot data.
As first or co-first author, he has published at NeurIPS 2025, ECCV 2026 and IEEE RA-L, and has three papers accepted to NeurIPS 2026. He is also a co-author of a Nature Machine Intelligence study on VLA design. His BridgeVLA model won the COLOSSEUM Challenge at the CVPR 2025 GRAIL Workshop and the ARNOLD Challenge at the CVPR 2026 EAI Workshop. He was named an ICML Gold Reviewer in 2026. As an undergraduate, he was selected by People's Daily as one of its Outstanding National Scholarship Student Representatives, the only undergraduate from Xi'an Jiaotong University chosen that year. He also runs marathons and has raced in Xi'an, Shenzhen, Tianjin, Zhengzhou and Lanzhou.
BridgeVLA
BridgeVLA, Li's first-author NeurIPS 2025 paper, deals with a mismatch in many 3D robot models. The vision-language models they build on are trained on 2D images, while the robot works in 3D. BridgeVLA projects the 3D point cloud into several 2D views and outputs actions as 2D heatmaps, so the input and output are both images. It is also pretrained to locate named objects as heatmaps before it learns any robot task.
On the RLBench benchmark, average success went from 81.4% to 88.2%. On COLOSSEUM, which tests robustness to visual and physical changes, it went from 56.7% to 64.0%. On a real robot it reached 96.8% success across more than 10 tasks with three demonstrations per task.

BridgeVLA projects 3D input into several 2D views and predicts actions as heatmaps. Source: BridgeVLA project page.
SpatialVAM
SpatialVAM, accepted to NeurIPS 2026, adds prediction over time. It is built on a video generation model and predicts what the scene will look like next, both as multi-view video and as heatmap video, then turns those predictions into actions. With a small number of demonstrations and no task-specific pretraining, it outperformed the strongest baselines reported in the paper by about 22% on Meta-World, 15% on RoboCasa and 16% on real-robot tasks.

SpatialVAM predicts future multi-view video and heatmaps, then converts them into actions. Source: SpatialVAM project page.
What It Means for Industry
For companies putting robots into factories, warehouses or homes, collecting demonstrations for each new task is a large part of the deployment cost. Methods that cut the number needed from hundreds to a handful lower that cost directly. The two IROS papers he co-authored work on a different problem, reliability on long tasks, which is often what keeps a robot that works in the lab from doing a real job. His internships at Tencent, ByteDance and Xiaomi have put him on the industrial side of both questions as well as the research side.
Paper Information
Paper | Venue | Links |
HCoT-VLA (Hybrid Chain-of-Thought Reasoning for VLA Models) | IROS 2026, Paper #2265 | 2026.ieee-iros.org/program/paper-index |
StaKe (Improving VLA Fine-Tuning with Structured Stage and Keyframe Supervision) | IROS 2026, Paper #2610 | 2026.ieee-iros.org/program/paper-index |
BridgeVLA | NeurIPS 2025 | bridgevla.github.io | arxiv.org/abs/2506.07961 |
SpatialVAM | NeurIPS 2026 | arxiv.org/abs/2604.03181 |
BridgeVLA++ | arXiv 2026 | bridgevla-plus.github.io |
About the researcher: Peiyan Li, Ph.D. candidate, NLPR, Institute of Automation, Chinese Academy of Sciences. Homepage: lpy1219.github.io | Email: [email protected]
Leave a comment