WAIC 2026 panelists argued that Chinese world model startups may be moving beyond U.S. The central wager is that world models and embodied robots reinforce each other: models help robots predict actions, while deployed robots generate the interaction data needed to improve models.
Research answer

Create a landscape editorial hero image for this Studio Global article: How did WAIC 2026 panelists from Muka Robotics, Shengshu Technology, EvoPhys.ai, and Chengwei Capital argue that Chinese world-model startup. Article summary: The panel’s argument was not that China has matched U.S. frontier-AI resources, but that world models change the source of advantage: progress depends less exclusively on raw compute and more on efficient model design, a. Topic tags: general, general web, user generated, academic. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, char
World models are becoming a central theme in physical AI: systems intended not only to generate plausible video, but also to represent how environments change and how actions affect them. At WAIC 2026, a panel involving Muka Robotics, Shengshu Technology, EvoPhys.ai and Chengwei Capital argued that this shift could let Chinese startups compete on a different basis than the familiar race to match U.S. frontier-model scale.
Their claim was not that China has equal access to frontier compute. Instead, the speakers’ case was that world-model progress may depend increasingly on learning efficiency, real-world interaction data and fast robot-based testing—areas where a dense robotics and hardware ecosystem could matter as much as raw GPU totals. Coverage of WAIC supports the broader context: the conference featured substantial embodied-AI activity, including more than 200 participating embodied-AI and robotics companies in one report. 8
The panel framed the traditional comparison with U.S. AI labs as less useful for a field that is still defining its own technical and commercial standards. Their thesis was that a team with less compute could still be competitive if it can collect useful physical-world data cheaply, train architectures that use heterogeneous data efficiently, and test failures quickly on real machines.
Muka Robotics’ reported WorldArena result—described in coverage as a high ranking achieved with 32 GPUs—was presented as an illustration of this efficiency-first approach. That is a company-reported benchmark result, however, rather than an independently audited measure of overall technical leadership. 3
The important distinction is practical: a video model can be judged by visual plausibility, whereas an embodied world model can be tested by whether a robot’s predicted action succeeds under physical constraints. A model that anticipates contact, motion or object state incorrectly can be exposed when it controls or informs a real robot.
The panel’s argument rests on a feedback loop between world models and embodied intelligence:
WAIC’s exhibition footprint illustrates why the panel sees this as relevant in China. Shanghai’s official WAIC guide described more than 300 robots in dynamic application scenarios, while other coverage reported more than 200 embodied-AI and robotics companies at the event. 10
8
This does not establish that any country has a decisive data advantage. It does explain the panelists’ strategic logic: a large and varied robotics base could provide both more opportunities to collect interaction data and more settings in which to check whether a model’s physical predictions are useful.
The panel also pointed to Shenzhen-style hardware coordination: component providers, integrators, robot manufacturers and software teams can be geographically and commercially close enough to shorten the loop between a model failure and a new experiment.
In this view, the advantage is not merely cheaper hardware. It is the ability to alter a sensor setup, end effector, data-collection routine or robot configuration, then bring the resulting data back into training quickly. That may matter in physical AI, where model quality is inseparable from the machines, sensors and environments used to collect and validate data.
It remains a strategic proposition rather than a demonstrated guarantee of technical leadership. Supply-chain responsiveness can accelerate experimentation, but it does not by itself solve core questions about generalization, safety, reliability or commercial viability.
Shengshu Technology’s world-action work provides a concrete example of the architecture-led approach described at the panel. Its MotuBrain research presents video and action in a shared generative formulation, treating policy modeling, world modeling, video generation, inverse dynamics and joint video–action prediction as different inference modes of one model. 32
Shengshu describes MotuBrain as jointly modeling understanding, prediction and action, with the stated goals of improving use of heterogeneous data and strengthening transfer across tasks. 33 In practical terms, that is an attempt to reduce the separation between a system that observes the world, predicts its evolution and chooses an action.
The significance is not that a single published architecture has settled the field. Rather, it illustrates why the panel believes model design may create leverage: if one framework can reuse information across video, action and physical prediction tasks, useful capability may not scale in a simple one-to-one relationship with compute.
EvoPhys.ai was presented as approaching the same bottleneck from another direction: learning physical regularities from primitive interactions instead of relying primarily on vast curated video corpora. The intended benefit is a model that is more grounded in forces, contact, motion and action consequences.
Related activity was visible at WAIC. TrendForce reported that Moore Threads demonstrated full training of Peking University’s EvoPhys-World, described as a 5D world model, on domestic computing infrastructure. 9 That confirms active work in this technical area, but it does not independently validate the panel’s broader claims about comparative performance, data efficiency or national leadership.
The metaphor captures two sides of the same argument.
If world models become a core layer for robots, autonomous systems and simulation, early teams could help define important standards: what data to collect, how to evaluate physical reasoning, which robot embodiments to support and which applications create real value. The absence of a settled global leader or established product playbook could create room for first movers.
The same lack of a template raises execution risk. Startups must make expensive choices about data pipelines, hardware platforms, model architectures, evaluation methods, customers and business models before the market has shown which choices will endure. Physical-world validation can also be slower and costlier than evaluating a software-only model.
That is why the panel’s argument should not be read as a declaration that China has already won the world-model race. It is a claim about the shape of the competition: in an emerging field where models must learn from and act in the physical world, iteration speed, data quality and architecture may become strategic variables alongside compute.
The thesis will be tested less by conference narratives than by evidence in four areas:
For now, the available sources substantiate active robotics participation at WAIC and Shengshu’s published world-action formulation. They do not independently establish the panel’s specific compute comparisons, the claimed number of potential data-supplying robotics firms, or an industry-wide Chinese lead. Those claims are best understood as an ambitious strategic thesis—one whose credibility will depend on repeatable, real-world results. 8
10
32
33
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
WAIC 2026 panelists argued that Chinese world model startups may be moving beyond U.S.
WAIC 2026 panelists argued that Chinese world model startups may be moving beyond U.S. The central wager is that world models and embodied robots reinforce each other: models help robots predict actions, while deployed robots generate the interaction data needed to improve models.
The opportunity also creates a burden: without a proven global playbook, founders must discover the right architectures, data pipelines, evaluations and commercial use cases themselves.