On September 22, 2026, Li Auto presented ME Brain 1.0 as an embodied system, ME VLM as a vision language model and ME U0 as an understanding and action model. ME U0’s reported 99.1% LIBERO and 81.8% LIBERO Plus success rates follow benchmark specific adaptation; they are not zero shot real world success rates.
Published byEdited with GPT-6 SolImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Li Auto’s Foundation Model team release in its September 22, 2026, MachEmbodied embodied-AI trilogy, and how do ME-Brain-1.0, ME-VL. Article summary: Li Auto’s Foundation Model team presented three complementary pieces on September 22: ME-Brain-1.0 as an embodied-agent system, ME-VLM as its vision-language cognitive core, and ME-U0 as a model for understanding scenes . Topic tags: general, academic, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, char
Li Auto’s Foundation Model team introduced three related embodied-AI releases on September 22, 2026: the ME-Brain-1.0 system, the MachEmbodied-VLM (ME-VLM) vision-language model and the MachEmbodied-U0 (ME-U0) understanding-and-generation model. The releases outline a route from observing a scene to reasoning about it and producing actions. They do not, by themselves, demonstrate that the entire stack operates as one independently validated robot on edge hardware. 3
5
ME-Brain-1.0 is the system-level concept. Li Auto describes it in terms of memory, cognition and action. Reporting on the release says its memory-related modules can run on the company’s M100 system-on-chip (SoC) for on-device memory generation and retrieval. That is a narrower claim than running every part of ME-Brain locally. 3
5
ME-VLM supplies vision-language cognition and agent coordination. Its technical report describes training across embodied and multimodal agent tasks, including observations and feedback intended to help assess outcomes and refine decisions. Reporting identifies a 35B-A3B mixture-of-experts version and a smaller 4B version designed for edge deployment. 1
5
ME-U0 links understanding to generation. Its reported Mixture-of-Transformers design connects experts for interpreting tasks and scenes with experts for generating actions. The distinction matters: a model that can identify what to do still has to turn that understanding into useful robot behavior. These releases are presented as complementary parts of a stack, not proof that all three have been tested together in a single autonomous deployment. 2
5
Li Auto’s comparison tables reportedly give the large ME-VLM an average of about 70.9 on embodied suites and 72.5 on agent suites, with the 4B variant ranking second in those company comparisons. Those figures summarize the reported evaluations; without the full test conditions and independent reproduction, they should not be read as general measures of robotic reliability. 5
For ME-U0, Li Auto reports a 17.66 score on RoboDojo and average success rates of 99.1% on LIBERO and 81.8% on LIBERO-Plus. The reported LIBERO results followed adaptation using supervision supplied by each benchmark. They therefore should not be mistaken for zero-shot performance on unfamiliar physical tasks. 3
Li Auto says it selected approximately 4,200 hours of ME-U0 pretraining data from raw pools containing roughly 5,700 hours of robot data and 3,920 hours of first-person video. Combining those sources could expose a model to visual situations beyond recorded robot demonstrations, but the data totals alone cannot show how well it transfers to a new robot or environment. 3
For edge inference, the ME-VLM authors report visual-token compression, W4A8 quantization and hardware–software co-optimization for the 4B model on Li Auto’s M100 SoC. They report a reduction in prefill latency from 400 ms to 188 ms. Prefill measures one stage of inference—not the time from a robot seeing a change to completing a physical response. The cited material does not establish that the 35B-A3B model or the complete ME-Brain and ME-U0 stack runs together on one M100. 1
5
Coverage of the launch describes an hour-long grasping demonstration and a roughly four-minute coffee-making sequence with 15 steps. Those demonstrations illustrate the tasks Li Auto chose to show, but they do not establish success rates under independently selected objects, lighting, interruptions or repeated long-term use. The available cited material does not substantiate a more specific third-party finding about failures in those demos. 3
The clearest next test is an independently specified, reproducible evaluation of the integrated system: task success and failure recovery on unfamiliar setups, alongside full perception-to-action latency measured on the stated hardware. Until then, the benchmark scores and M100 prefill result are promising but distinct pieces of evidence—not a single verdict on general-purpose robotics. 1
3
5
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
On September 22, 2026, Li Auto presented ME Brain 1.0 as an embodied system, ME VLM as a vision language model and ME U0 as an understanding and action model.
On September 22, 2026, Li Auto presented ME Brain 1.0 as an embodied system, ME VLM as a vision language model and ME U0 as an understanding and action model. ME U0’s reported 99.1% LIBERO and 81.8% LIBERO Plus success rates follow benchmark specific adaptation; they are not zero shot real world success rates.