After working on AutoML and autonomous driving perception at Alibaba DAMO Academy, Yu Kaicheng joined Westlake University and built AutoLab around world models. His team’s late 2024 pivot abandoned pixel level generation as the primary predictive target in favor of a latent space representing objects, states, relati...
Research answer

Create a landscape editorial hero image for this Studio Global article: How did Westlake University researcher and Awomo founder Yu Kaicheng evolve from an Alibaba DAMO Academy AutoML PhD and Alibaba Star—whose B. Article summary: Yu’s shift was not a rejection of perception or large scale data; it was a conclusion that pixel level generation is an inefficient intermediate target for control once a system can already synthesize plausible observati. Topic tags: general web, chatgpt, agents, ai, workflow. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts w
Yu Kaicheng’s change of direction was not a rejection of perception, large datasets or the scaling of AI systems. It was a more specific conclusion: once a model can generate visually plausible observations, predicting every pixel may be an expensive detour from the information an autonomous system actually needs for control.
His alternative is to make the predictive state itself structured. Instead of repeatedly rendering a video of the future, the model should represent the entities in a scene, their states and relationships, the available actions, and the causal changes that follow. An agent could then recombine those concepts and plan in latent space. 1
10
Yu first worked on automatic machine learning, or AutoML, before joining Alibaba DAMO Academy through the Alibaba Star talent programme. At Alibaba, he moved into autonomous-driving perception and helped develop BEVFusion, a framework that combines camera and LiDAR data in a shared bird’s-eye-view representation. The work became a widely cited contribution to multisensor fusion. 10
He joined Westlake University in 2023 and formed the Autonomous Intelligence Lab, or AutoLab. The lab initially explored generative world models for autonomous driving. Its subsequent results included a first-place finish in an ECCV autonomous-driving track in 2024 and the top position in the Fail2Drive ranking in 2026, according to contemporary reporting. 10
That success made the eventual pivot more striking. Yu’s team had demonstrated the value of data-driven perception and generative modelling, yet its work on real-world applications raised a practical question: if the system already understands enough of the environment to predict what will happen, why spend the computational budget reconstructing every visual detail?
The answer, in Yu’s account, is that the useful question for an autonomous agent is not simply “what will the next frame look like?” It is “what changes if I take this action?” Rendering pixels can help create realistic observations, but it may obscure the task-relevant state and causal structure needed for planning. 1
10
In late 2024, the team reportedly moved from pixel generation to an end-to-end learned mathematical latent space. The aim is to decompose a scene into a finite set of reusable physical concepts and relations, then compose them into a predicted next state—including combinations that did not appear as complete scenarios in the training data. 1
“Implicit” in this context does not mean unstructured. Awomo’s proposal is that the latent representation should capture concepts such as:
The intended advantage is compositional generalization. If a system learns the relevant concepts separately, it may be able to reason about a new combination of them without needing demonstrations of every complete situation. That is the basis for Yu’s argument that simply adding more vision-language-action, or VLA, data and demonstrations will eventually deliver diminishing returns in physical intelligence. 1
The approach therefore shifts the modelling target from visual appearance to the evolution of task-relevant physical state. This broader latent-space direction is already associated with model-based reinforcement learning and predictive world models, which use compact internal states to imagine possible futures rather than reconstructing full-resolution observations at every step. 13
Awomo’s concept world model belongs to the same broad family as Yann LeCun’s Joint Embedding Predictive Architecture, or JEPA: both emphasize predicting a representation of a future state instead of rebuilding all the pixels in an observation. 1
The distinction Awomo claims is what the latent variables are meant to contain. In Yu’s characterization, JEPA leaves the content of the latent space relatively open, while the concept-world-model proposal explicitly seeks a composable vocabulary of physical concepts. 1
That distinction is important, but it remains a research hypothesis rather than an independently established advantage. A structured latent space could make reasoning more efficient and interpretable; it could also impose assumptions that fail in environments with unexpected or poorly modelled dynamics. Public results will need to show whether the proposed structure consistently improves generalization.
Awomo has reported internal 2025 experiments in which concept decomposition and association improved success on complex tasks compared with imitation-learning and reinforcement-learning baselines, while also reducing prediction cost. 1
The public material available here does not provide the full experimental protocol, model sizes, task distribution, numerical success rates, confidence intervals, ablation studies, released code or independent replication. The scale of the claimed advantage is therefore not yet sufficiently evidenced publicly. 1
Yu’s example of catching a falling cup while closing a door illustrates the intended benefit. A capable physical agent would need to track several concurrent state transitions and their causal constraints at once. It should not merely select a familiar response from a demonstration corpus. But this example is best understood as a motivating illustration of the research direction, not as a verified benchmark result. 1
The research was commercialized through Westlake Digital Intelligence Technology (Hangzhou) Co., Ltd., branded Awomo. The company was registered in Hangzhou’s Xihu District on January 8, 2026. It says it is developing physical-AI capabilities for autonomous driving, general-purpose robotics, industrial equipment, intelligent hardware and simulation-data generation. 11
Reporting describes the company’s initial core as Yu and chief scientist Tong, later expanded with members of the Westlake AutoLab team. Descriptions of the group as “consistently self-validated geniuses” are promotional language, not an independently measurable assessment of the team. 1
Awomo announced cumulative seed and angel-stage financing of more than RMB 100 million. The named investors include InnoFund, Dongfang Jiafu Fund, Zhengxuan Investment, Tianqi Capital, Westlake Innovation Fund and Jinma Investment. 13
The company’s branding reflects its ambition: to build a general-purpose model for the physical world rather than another system focused only on language, images or video. Its stated technical route is “latent-first”—maintaining an internal physical representation that can support prediction, planning, action and self-correction. 11
August 2026 coverage presented Awomo-v0.1 as the first public version of the company’s concept-world-model effort and referenced evaluation on RoboTwin 2.0. RoboTwin 2.0 is a simulated framework and benchmark for generating data and evaluating robust bimanual robotic manipulation. 1
3
That public appearance provides an initial evaluation venue, but it does not by itself demonstrate reliable transfer to physical robots, safety in open-ended environments or superiority over leading alternatives. A stronger assessment would require disclosed benchmark scores, clearly defined baselines, ablation results and real-world transfer experiments.
Yu’s trajectory is therefore less a rejection of his earlier work than an attempt to move one level deeper. BEVFusion helped unite different sensors in a common physical representation. The concept-world-model programme asks whether an AI system can go further: maintain a compact, composable account of objects and causal change, then use that account to decide what to do before it acts.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
After working on AutoML and autonomous driving perception at Alibaba DAMO Academy, Yu Kaicheng joined Westlake University and built AutoLab around world models.
After working on AutoML and autonomous driving perception at Alibaba DAMO Academy, Yu Kaicheng joined Westlake University and built AutoLab around world models. His team’s late 2024 pivot abandoned pixel level generation as the primary predictive target in favor of a latent space representing objects, states, relations, actions and causal changes.
Awomo says the approach could help AI generalize to unfamiliar combinations of events, although the public evidence currently lacks detailed protocols, benchmark scores and independent replication.