Xiaomi’s Xuanjie O100 prototype reached a reported 295 tokens per second running the company’s MiMo model, but it is an engineering demonstrator—not a shipping phone. The AI Cube prototype takes the same local AI strategy to a stationary mini PC: three Xuanjie chips, 80GB of unified memory and 150W sustained perform...
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Xiaomi president Lu Weibing reveal on August 25, 2026, about the Xuanjie O100 prototype—an ultra-high-speed on-device AI terminal t. Article summary: Lu Weibing presented the Xuanjie O100 prototype as a purpose-built demonstration of phone-class, on-device large-model inference—not a conventional consumer handset. Its design sacrifices the rear-camera system for dual-. Topic tags: general, education, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, cha
Xiaomi president Lu Weibing’s August 25, 2026 reveal was less a phone launch than a glimpse at what a local-AI-first handset might require. The Xuanjie O100 prototype combines the Xuanjie O3 smartphone processor with the O100 AI accelerator, removes the traditional rear-camera module, uses a rebuilt motherboard and adds active cooling rated for up to 10 watts. Running Xiaomi’s MiMo on-device model, it reportedly reached 295 tokens per second in testing. 3
7
That figure comes with an important qualification: the device is a technology-validation prototype, not an announced consumer product. 3 Its value is therefore not whether buyers can purchase it today, but what it demonstrates about the hardware trade-offs involved in running large models locally.
The O100 prototype is designed around one objective: sustained, high-speed inference at the edge. Rather than treating AI as another feature inside a conventional camera-focused smartphone, Xiaomi paired two chips for collaborative computing:
To accommodate that design, Xiaomi removed the conventional imaging module and redesigned the motherboard. It also added active air cooling capable of supporting up to 10 watts of cooling power. 3
4
The result is a phone-derived form factor built to showcase local large-model performance rather than normal handset priorities such as cameras, silent operation or maximum thinness. Xiaomi’s reported real-world result was 295 tokens per second while running its MiMo edge model. 3
5
That number should be read as a Xiaomi-reported result for a specific prototype and model, not as a universal benchmark for every AI model or future O100 device. The available reports do not establish how third-party models would perform under the same conditions.
The two prototypes demonstrate different endpoints for Xiaomi’s local-AI strategy. The O100 is a mobile proof of concept; AI Cube is a stationary personal AI-computing terminal.
| Feature | Xuanjie O100 prototype | Xiaomi AI Cube prototype |
|---|---|---|
| Form factor | Phone-derived AI demonstration terminal | Mini-PC and personal AI workstation |
| Chip configuration | Xuanjie O3 + O100 | Xuanjie O3 + O100 + D100 |
| Cooling or sustained output | Active cooling rated up to 10W | Claimed 150W sustained performance |
| Memory | Not specified in the cited reports | 80GB unified memory |
| Demonstrated model capability | Xiaomi MiMo at a reported 295 tokens per second | 120B model plus 3B on-device model, with fast/slow system switching |
| Product status | Prototype | Prototype |
AI Cube has substantially more room for power delivery, memory and thermal management. Its aerospace-aluminum unibody includes 33,874 CNC-machined ventilation holes, and Xiaomi says the system can sustain up to 150 watts. Inside are all three Xuanjie chips—O3, O100 and D100—along with 80GB of unified memory. 4
6
10
Xiaomi says AI Cube can deploy a 120-billion-parameter model alongside a 3-billion-parameter on-device model and switch between faster and slower systems depending on the task. The demonstrations were positioned around demanding workloads such as front-end development and complex programming. 3
7
In practical terms, AI Cube is closer to a compact local-AI workstation. The O100 prototype is more interesting as a test of whether that kind of inference acceleration can eventually be adapted to a consumer phone’s much tighter power and thermal envelope.
The prototypes fit into a broader three-chip product architecture Xiaomi presented at its August 24 Xuanjie technology briefing:
This division suggests that Xiaomi is building an AI-computing stack across phones, personal computers and vehicles rather than relying only on cloud processing. That is an interpretation of the hardware direction, not a confirmed promise that the two prototypes will become retail products.
The commercial timeline also matters. Xiaomi says the O3 has entered mass production and is scheduled to debut in the Xiaomi 18 Fold in September. The O100 and D100 have been developed but are planned for commercialization later, with Xiaomi’s community announcement describing that timing as the following year. 18
29
The Xuanjie O100 prototype makes Xiaomi’s local-AI ambitions tangible, but it also exposes the cost of chasing sustained inference speed. A conventional camera module had to go, the motherboard had to be rebuilt and active cooling had to be added. That is a very different design priority from a mainstream flagship phone.
AI Cube shows the other side of the equation: give the system a larger chassis, far more power and more memory, and Xiaomi can demonstrate much larger local models and heavier workloads. Together, the devices show a company testing multiple form factors for self-developed silicon and on-device AI.
For now, the most concrete product takeaway is the O3-powered Xiaomi 18 Fold. The O100 prototype is better understood as a research and engineering statement: Xiaomi believes fast local large-model inference may eventually belong in phones, but the path to a practical consumer device still requires solving cooling, power, software compatibility and product-design constraints.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Xiaomi’s Xuanjie O100 prototype reached a reported 295 tokens per second running the company’s MiMo model, but it is an engineering demonstrator—not a shipping phone.
Xiaomi’s Xuanjie O100 prototype reached a reported 295 tokens per second running the company’s MiMo model, but it is an engineering demonstrator—not a shipping phone. The AI Cube prototype takes the same local AI strategy to a stationary mini PC: three Xuanjie chips, 80GB of unified memory and 150W sustained performance, with support for 120B and 3B models.
Only the Xuanjie O3 has a near term retail device announced: Xiaomi says it will debut in the Xiaomi 18 Fold in September, while O100 remains a prototype.