China’s first July 17, 2026 AI terminal test results named 11 mobile devices—nine smartphones and two tablets—as L3. L3 is the highest level with defined, testable requirements today, but the framework also names a future L4 collaboration level.
Research answer

Create a landscape editorial hero image for this Studio Global article: What is China’s MIIT national AI-terminal intelligence grading standard, what does the first July 2026 batch of 11 certified devices—includi. Article summary: China’s AI-terminal grading system is a voluntary capability benchmark, not a phone-approval regime. The July 2026 L3 results show that some specific devices can act as bounded, user-supervised agents—not that they are a. Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
China’s new AI-terminal grading system is best understood as a capability benchmark for complete devices, not as a “self-driving” scale for phones. The first results, released on July 17, 2026, listed 11 mobile terminals at L3: nine smartphones and two tablets. They show that particular submitted configurations can provide bounded, user-supervised assistance—not that the devices can act without oversight across every app or service. 2
5
The framework is called GB/Z 177—2026, Intelligence Grading of Artificial Intelligence Terminals. The “GB/Z” designation identifies a national standardization guidance document, so the system provides a common way to evaluate AI capability rather than a compulsory condition for selling a phone. 1
5
6
It assesses the terminal as a system. That includes the model, operating system, app and tool interfaces, memory, permissions, and the relationship between on-device and cloud processing—not simply the size or conversational quality of an underlying language model. 5
19
The framework organizes capability around five dimensions:
Those dimensions are further divided into 14 sub-capabilities, including user perception, task planning, tool invocation, short- and long-term memory, and contextual adaptation. 19
The four labels describe an increasing ability to move from reacting to instructions toward coordinating work across systems. The first three levels have defined capability requirements; L4 is named in the framework but remains to be fully specified. 5
21
25
| Level | What it means in practice |
|---|---|
| L1 — Response | Understands a clear instruction and performs a straightforward, usually single-step action, such as opening a function or setting a reminder. |
| L2 — Tool | Understands a simpler intent, performs limited reasoning and follows predefined tools or workflows for constrained multi-step tasks. It is closer to a capable chatbot or in-app assistant than an open-ended agent. |
| L3 — Assistance | Understands a broader goal, asks for missing information, breaks work into steps, selects and orchestrates tools, handles multimodal interaction, and uses short- and long-term memory. |
| L4 — Collaboration | Intended to coordinate among agents, devices and services, but its detailed requirements and test methods are still being developed. |
The important distinction is not that an L3 phone “thinks like a person.” It is that the terminal can connect understanding, planning, tool use, memory and action into a more complete task workflow. 5
19
The July results included nine smartphones associated with Huawei, Motorola, Honor, vivo, OPPO, Xiaomi and Stepfun, along with two tablets. 2
3
5
The narrow conclusion is that the submitted device configurations met the current L3 test requirements. The result does not mean:
The first batch came from vendor-submitted products rather than a random survey of the market. An L3 result is therefore a product-and-configuration rating, not a brand-wide intelligence ranking. 5
6
Calling the result a “national L3 certification” can obscure what the label does and does not do. Because GB/Z 177—2026 is a guiding technical document, an L3 result is not a regulatory authorization required before a phone can be sold. It is an industry benchmark for capability and evaluation. 1
6
The label also should not be read as proof of universal autonomy. L3 evaluates whether a terminal can provide complex assistance under defined conditions; it does not grant the device unrestricted access to personal data, accounts, payments or every third-party app. The user remains part of the control loop, especially for consequential actions. 5
19
An L2 assistant might answer a question about the weather in Hangzhou, search for information or generate a suggested itinerary through a preset workflow. The conversation and task usually remain within a limited tool path or current session. 5
19
An L3 phone is designed to handle the broader goal behind a request. For example, after being asked to plan a weekend trip, it could:
That final step defines the boundary. L3 is supervised delegation, not unrestricted authority. The agent may prepare an action, but final control should return to the user when money, sensitive data or an irreversible commitment is involved. 5
L4 is not simply “a more powerful L3.” Its defining challenge is coordination among independent agents, devices and services. A workable L4 test would need clear rules for:
Those questions involve the phone maker, operating-system provider, app, service provider, model provider and user. Until the industry can express those relationships as repeatable and safe test cases, a detailed L4 certification method is difficult to establish. The next meaningful step may therefore depend as much on interoperability, permissions and accountability as on faster chips or larger models. 5
7
The shared L1–L4 labels are easy to misread. Automotive autonomy grades describe responsibility for the dynamic driving task, including who monitors the vehicle and when a human must take over. The AI-terminal framework instead grades capabilities such as perception, intent understanding, reasoning, tool use, memory and learning. 5
An L3 phone is therefore not equivalent to a conditionally autonomous car. Likewise, a future L4 phone would not automatically mean a device that can safely make every consequential decision without human confirmation. 5
The Huawei Pura X Max and Lenovo’s Motorola devices are useful examples of products associated with the first L3 results, but the label should not be the only buying criterion. 2
5
A phone’s practical agent capabilities can depend on operating-system updates, assistant and model availability, app integrations, permissions and cloud connectivity. That means useful L3-style functions do not automatically require a new hardware purchase, although newer hardware may improve factors such as speed, privacy or battery performance. 5
Before treating an L3 badge as a buying decision, check:
The clearest takeaway from the July batch is that AI phones are becoming measurable as task-performing systems. But L3 is a supervised assistance benchmark for specific configurations—not a blanket promise of autonomous behavior, not a mandatory market-entry approval, and not the final level of terminal intelligence. 1
5
7
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
China’s first July 17, 2026 AI terminal test results named 11 mobile devices—nine smartphones and two tablets—as L3.
China’s first July 17, 2026 AI terminal test results named 11 mobile devices—nine smartphones and two tablets—as L3. L3 is the highest level with defined, testable requirements today, but the framework also names a future L4 collaboration level.