World Models vs. VLA: WAIC 2026 Exposes a Technical Fork in Robotics, With 'ChatGPT Moment' Predictions Ranging From 2 to 5 Years
Executive Summary
At the 2026 World AI Conference (WAIC) in Shanghai in July, the debate in embodied intelligence shifted from whether generalist robot models are achievable to which architecture gets there first. Vision-language-action (VLA) models and world models collided directly across multiple forums, and companies on both sides offered specific predictions for robotics' "ChatGPT moment" — a generalist model that works reliably out of the box — ranging from under three years to five years. Georgia Tech assistant professor Dan Fei Xu presented what she called the first verified data scaling law in robotics, showing task success rate rising linearly with the log of data volume across 20,000 hours of human demonstration data. Daxiao Robotics launched its Kairos 3.1 world model, claiming 125-millisecond inference latency for an 8-billion-parameter model running on-device. NVIDIA released an open-source multimodal world model, Cosmos 3. Morgan Stanley data cited at the conference projects humanoid shipments rising from 50,000 units in 2026 to 445,000 in 2030, a 106% compound annual growth rate.
Industry Context
The Zhiyuan x Mifeng "Embodied Intelligence 2026" forum, the Daxiao Robotics-hosted "World Model Six Little Dragons Summit," the JD-hosted "AI Enters the Physical World" forum, and the RISC-V and Physical AI forum formed the technical core of this year's embodied-intelligence coverage at WAIC.
SenseTime chairman and CEO Xu Li cited Morgan Stanley's forecast at the World Model Summit: global humanoid shipments will grow from 50,000 units in 2026 to 176,000 in 2027, 296,000 in 2028, and 445,000 in 2030 — a 106% compound annual growth rate from 2025 to 2030. He attributed the growth to robots gaining "a brand new brain that can adapt to the physical world and understand the physical interactions," rather than to hardware iteration in motor torque or gearbox precision.
Tang Wenkan, director of the Shanghai Municipal Commission of Economy and Information Technology, disclosed that the city's nearly 20 robot platforms ranked first among cities worldwide in 2025 shipments. JD.com separately announced a plan to collect 10 million hours of embodied data within two years, with weekly capacity already near 100,000 hours.
Technology: Two Parallel Pretraining Paths
Embodied intelligence follows two main technical paths. The first is VLA, which maps visual input and language instructions directly to robot actions; leading companies include Physical Intelligence (PI), Dyna Robotics, and Sunday Robotics. The second is world models, which first build an internal representation of physical laws and environment state, then use it to predict future states and plan actions; leading companies include Daxiao Robotics, X Square Robot, and NVIDIA.
Yao Maoqing, partner, senior vice president and president of the embodied intelligence unit at Zhiyuan (AgiBot), argued at the "Embodied Intelligence 2026" forum that physical AI must break through "three walls" before generalist intelligence can emerge: a data wall (real interaction data is scarce and expensive, unlike scraping web pages), a representation wall (no unified physical representation exists across tasks, environments, and embodiments), and a closed-loop wall (real-world trial and error is costly). "The cost of trial and error in the real world is still extremely high," he said. "In the physical world, one failure may mean a broken component, or more likely the loss of an entire robot." Zhiyuan said it is pursuing both a VLA line and a "World Action Model" (WAM) line in parallel, aiming eventually for a unified architecture.
Allen Ren, a research scientist at Physical Intelligence, described the evolution of the company's pi-series models from pi-0 through pi-0.7. Pi-0.7 borrows techniques from large language and video-generation models: an efficient video encoder compressing historical context, dense multimodal instructions, and episode-level metadata scoring data quality and flagging mistakes. Ren said a single checkpoint, without task-specific fine-tuning, can match or exceed previous single-task policies, and exhibits emergent cross-embodiment transfer — a clothes-folding skill transferred from a low-cost dual-arm robot to a UR5 arm that had never collected training data. He acknowledged VLA has not been superseded: "So far I haven't seen anyone produce a more impressive demo than PI's, so VLA is probably not dead yet."
Wang Xiaogang, chairman of Daxiao Robotics, unveiled Kairos 3.1, describing it as an integrated world model combining generative, physical, and cognitive intelligence on a unified mixture-of-experts architecture with hybrid shared attention, forming an "understand-predict-simulate-evaluate-execute-reflect" loop. He stressed that prediction errors in the physical world carry irreversible costs: "The misunderstandings or the faulty predictions from the emerging must trigger wrong actions, incurring real and irreversible costs." NVIDIA vice president of engineering and solutions Lai Junjie announced Cosmos 3, an open-source model natively supporting five modalities — text, vision, image/video, audio, and action — using a dual-tower "reasoning plus generation" architecture in which shared multimodal attention constrains generation to obey physical laws.
Wang Hao, co-founder and CTO of X Square Robot, offered a more cautious view: a world model does not need to precisely understand physics, only physical "plausibility." "It only needs to know, say, that iron falls faster than cotton. As for exactly how fast, it can observe — it doesn't need to be that precise," he said, adding roughly 90% confidence that building physical understanding purely from video is fundamentally flawed.
Engineering Analysis: Data Scaling Laws and the Latency Constraint
Dan Fei Xu, assistant professor at Georgia Tech's School of Interactive Computing, presented what she described as the first verified data scaling law in robotics. Her team trained robot policies on first-person human demonstration data — captured via smart glasses developed with Meta — and observed task success rate rising linearly against the log of data volume, up to 20,000 hours. "We got a very clear scaling law in pretraining," she said. "This may be the first time this has been verified in robotics." The team also found that 100 human demonstrations plus one robot demonstration is enough for a robot to learn a new task, and that greater embodiment diversity in the base model improves how well it learns from human data — "the pyramid may be upside down."
Yang Chao, a professor at Peking University's School of Mathematical Sciences, offered theoretical support along with a hard engineering constraint at JD's forum. He said he believes "scaling law will definitely be established in the era of physical AI," but physical AI inference must run on-device in real time: actuation frequencies of tens to hundreds of hertz require latency below tens of milliseconds, ruling out cloud inference. He cited results showing two NVIDIA B200 GPUs achieve only 6.7 Hz inference frequency — far below what robots require. He also described a "three crosses" generalization problem — across scenarios, tasks, and embodiments — noting current technology handles one cross reasonably well, two only with great difficulty, and three not at all.
Data recipes remain contested. The founder and CEO of Xinyan Robotics (surname Liu) said the company's training mix is roughly 90% non-embodiment data for base pretraining and 10% real-robot data, deemed critical because deploying to a specific robot requires recollecting data on that embodiment; he does not expect "pretrain and deploy with no post-training" within five years. Wu Wei, founder and CEO of Liuxin Space, argued AI development proceeds architecture-first, then algorithm scaling, then data scaling — mass data collection before an architecture converges risks becoming a sunk cost: "Data is defined by the model architecture." His team has abandoned pure simulation data, arguing mainstream simulators carry only about 20,000 hyperparameters, too few to responsibly train models with billions of parameters.
Commercial Progress: From Coffee to Napkins, Testing for Reliability
Commercial validation is shifting from demo spectacle to repeatable reliability metrics. Ren described a continuous coffee-making demo by PI's pi-star-0.6 model running from 5:30 a.m. to 11:30 a.m. with zero failures, fully autonomous. Dyna Robotics co-founder and chief scientist Ma Yecheng described napkin-folding progress: an early VLA model achieved roughly 80% success, meaning 100 consecutive folds would almost certainly include a failure. After iterating with an in-house reward model — predicting task-completion percentage frame-by-frame to locate where the model errs, then targeting recovery-data collection there — the model folded more than 800 napkins over 24 continuous hours, doubling initial throughput. "If the success rate is not high, there is not much commercial value," he said.
Zhiyuan released GenieSim 2.0, a closed-loop world simulator completing a 25-frame rollout in 2.3 seconds on an H100 cluster and ranking first in the World Arena challenge, with weights and code fully open-sourced. The company also introduced a distributed real-robot online reinforcement learning system spanning more than 100 robots, distributing tens-of-billions-of-parameters weights to all robots within two seconds and compressing training cycles from days to minutes. Yao said a 3C factory in Nanchang, Jiangxi ran robots live for six days, more than 10 hours daily, completing roughly 65,000 operations at 99.99% success.
Daxiao Robotics claims Kairos 3.1 achieves on-device world-model deployment: the 8B-parameter model runs at 125 milliseconds average latency in NVIDIA BF16 precision, versus 139 milliseconds for Pi-0.5 (3.3B parameters) and roughly 6,486 milliseconds for a comparable world model the company called "Smooth 3" — figures self-reported and not independently verified. It also open-sourced an "L5-level" dataset, ACE Data 0, spanning 24 camera views, six modalities, and 75,000 interaction sequences.
JD disclosed that leading embodied-model companies currently train on datasets of at most roughly 20,000 samples. Guo Yucheng, JD Group vice president and head of JD Cloud's foundational cloud business, said the company judges 1 million hours as the baseline for a major capability leap and 10 million hours as the scale current technology can meaningfully exploit. JD plans to reach 10 million hours within two years, with weekly capacity already near 100,000 hours. Duan Nan, deputy head of the JD Explore Academy, framed the target against the language-model timeline: "Our ultimate goal is to reach a GPT-3 or ChatGPT-style moment of capability emergence in the embodied domain."
Market Perspective: Compute Requirements and Shipment Forecasts
Chip vendors quantified the compute demand embodied intelligence is expected to generate. Li Huaqing, co-founder and vice president at Lingrui Zhixin, said at the RISC-V and Physical AI forum that robot compute requirements "can range anywhere from 200 TOPS to 2,000 TOPS," with tactile-sensing data volume driving the higher end. Sun Yanbang, co-founder and president of SpacemiT, offered a more aggressive figure: "To truly develop the brain for embodied intelligence, it will likely take over a thousand TOPS of AI compute."
Dai Weimin, chairman of the Shanghai Open-Source Processor Innovation Center and chairman, CEO and president of VeriSilicon, mapped current robot capability onto autonomous-driving grading conventions: "If robots could also achieve L1, L2, L3, L4 level of autonomous driving type of skills... we are still at the L2 level" — and, he added, robotics is harder because generalization requirements are more complex. He also rejected transplanting large language models directly into robot brains: "Today's Transformer language models have only taken the first of three steps — they are limited, and you cannot use such a large language model to do robots."
The Morgan Stanley shipment data implies a 106% compound annual growth rate from 2025 to 2030, which Xu Li attributed to the new "brain" world models provide rather than hardware iteration. Shanghai's 2025 AI industry reached 640 billion RMB, up 39.5% year over year.
Challenges: Missing Standards, Zero-Shot Skepticism, and the Value of Failure Data
Several researchers directly challenged current promotional claims. Stephen Redmond, a professor at University College Dublin, said he is skeptical of unqualified "zero-shot generalization" claims: "They just mention they can do zero-shot generalization... I usually don't believe such claims." Wolfram Burgard, founding director of the computer science and AI discipline at the Technical University of Nuremberg, said current VLA models are insufficient for complex tasks and what is actually needed is a multimodal action model incorporating vision, language, touch, and force, noting roughly 20% of factory production work remains beyond robots' current capability. He stressed the need to tolerate failure: "Like humans, AI will not be perfect... we have to allow robots to fail."
Zhu Chunhua, founder and CEO of Jupiter Robotics, criticized how loosely the term "world model" is now used: "What exactly is a world model? Frankly, the market is a mess — anyone doing anything can claim a connection to world models." Citing reinforcement-learning pioneer Richard Sutton's "bitter lesson," he argued that the ceiling on algorithmic paradigms, not raw data volume, determines how much intelligence can be extracted from data: "Data is crucial, but algorithmic capabilities have decided the ceiling of this."
Missing industry standards are also slowing progress. Zhu Zheng, co-founder and chief scientist of Jiajia World Vision, said the world-model industry prematurely explored commercialization scenarios in early 2026 that mostly did not pan out, prompting his team to refocus on core model capability: "The most painful lesson we learned is that the fundamental reason we can't deploy today is that model capability is still insufficient." He compared the industry's stage to coding models awaiting their "Codex moment": "For world models, what exactly is our Codex product?" In response, the Shanghai AI Industry Association, Shenzhen Hetao Institute, and CAICT's East China branch, with dozens of companies, launched the Physical IQ third-party benchmark to address the awkwardness of being "both referee and athlete." Tao Dacheng, chief scientist at Daxiao Robotics, summarized while moderating: "The industry will not pay for the intelligence; the industry will pay for constant intelligence," and "the yardstick of a demo is how good your best run is; the yardstick of deployment is how bad your worst run is."
Nobel laureate economist Thomas Sargent, a professor at New York University, provided theoretical context: learning algorithms do not converge to a rational-expectations equilibrium, only to a "self-confirming equilibrium," in which agents are exactly right about frequently repeated events but systematically biased about rare ones — a boundary that must inform world-model design. "Will learning algorithms converge to a rational expectations equilibrium? No — they only converge to a self-confirming equilibrium," he said.
RobotToday Analysis
Further reading: The VLA-versus-world-model contest documented here maps directly onto the "VLA vs. Wolrd Model" architecture debate mapped out in RobotToday's companion framework piece, Physical AI Landscape: From Digital Intelligence to the Embodied Physical World.
The technical split visible at WAIC 2026 reflects a deeper disagreement over where generalist intelligence should emerge. The VLA camp is betting that end-to-end policy learning plus massive demonstration data can approximate generality. The world-model camp is betting that a controllable, predictive representation of physics must come first, with action capability derived from it. Several speakers — including PI's Ren and Dyna Robotics' Ma — argued the two paths are converging rather than competing, which suggests today's rivalry is more a question of resource allocation than an architectural endgame.
The spread of "ChatGPT moment" predictions — from under three years to five — is itself informative. The most aggressive forecasters, such as Sunday Robotics ("within three years") and Zhiyuan (two years), tend also to be the ones making the most aggressive commitments on data collection or deployment scale. The more conservative forecasters — Xu and Ma both said five years — placed more emphasis on validating scaling laws and building data infrastructure over the long term. Readers should weigh speakers' commercial incentives alongside their timelines.
Three engineering bottlenecks deserve scrutiny: on-device inference latency (two B200 GPUs manage only 6.7 Hz, an order of magnitude short of the tens-of-milliseconds response robots require, ruling out a cloud-brain model near-term), missing data standards (speakers noted the industry has not converged on modalities, formats, or annotation protocols, constraining reuse across institutions), and benchmark credibility (platforms like Physical IQ signal the industry recognizes self-graded demos are unsustainable, though adoption by major vendors remains to be seen). The order-of-magnitude gap between Yao's estimate of "100 million hours" and JD's target of "10 million hours" indicates the industry has not agreed on what the scaling law actually looks like in practice. Robotics' "ChatGPT moment" is more likely to arrive as the gradual convergence of several infrastructure layers than as a single, discrete model release.
Quotes and figures in this article are drawn from conference public remarks at WAIC 2026 and have not been individually verified with the speakers or their organizations. Where a speaker's name could not be confirmed from the source transcript, this is noted in the text. Corrections are welcome.
WAIC Review 2/4: World Models vs. VLA: Robotics' Next 'ChatGPT Moment'
Leave a comment