Announced at Apsara 2026, Qwen Intelligence is a phone-agent software stack that Alibaba is offering to handset makers, rather than a smartphone Alibaba intends to manufacture. It combines mobile-optimized Qwen models, an agent platform and task-specific agents designed to carry out multi-step requests across apps. Honor is the first announced formal handset adopter.
13
19
45
How the three mobile agents divide the work
Mobile Planner Agent turns a request into a sequence of tasks, selects tools and adjusts the plan as a task progresses. In Alibaba’s description, the model handles understanding and decisions, while a harness manages task state, memory and available skills. That distinction matters for requests that cannot be completed in a single interaction.
12
43
Mobile-Use Agent performs the phone operations. Its reported approach is API first, graphical interface as a fallback: it can use an available interface or deep link, then rely on visual understanding and screen interaction where no suitable interface exists. This is the part of the stack intended to act across apps, not merely explain which buttons a person should press.
12
45
Mobile Creative Agent addresses a different need: image generation and editing using a lightweight model optimized for phones. It is a creative capability, rather than the component that plans or executes a cross-app workflow.
7
39
Where the OS Harness fits
A harness is the runtime around an agent: it keeps track of context and connects the model’s decisions to tools and execution. Reports describe a harness within Alibaba’s modular Qwen Intelligence platform. Separately, Honor says it has built a system-level Agent Harness in MagicOS 11 to coordinate perception, planning, tool calls and execution. These are related layers in the proposed integration, but they should not be presented as one identically owned product or as Qwen replacing MagicOS.
43
48
51
What phone makers can customize—and test
The offering is modular rather than a fixed Alibaba-branded phone experience. Reported options include deploying models independently, adapting the harness, choosing device–cloud arrangements, adding a manufacturer’s own tools and skills, and using platform operations and automated evaluation. Those choices let a handset maker fit the agents to its operating system and device constraints.
35
51
Safety becomes especially important when an agent moves from answering questions to operating apps. Reports say the Mobile-Use approach includes security and privacy boundaries, with sensitive actions such as payments or deletions subject to restrictions that leave the final decision with the user. Qwen Intelligence’s team has also described four evaluation sets covering complex planning, cross-app execution, real-phone tasks and safe action in risky scenarios. The available reporting does not show how every safeguard will be configured or perform on each shipping handset.
35
39
45
Honor’s role and what the numbers prove
Honor plans to integrate Qwen Intelligence into its Magic9 series, scheduled for a September 28, 2026 launch, and to support it in the Honor Robot Phone. The companies say they are co-developing mobile-specific capabilities around Qwen Intelligence and MagicOS: Alibaba supplies foundation-model capabilities, while Honor contributes system integration and device use cases. They are also defining model roles around phone limits such as power use, memory and latency.
1
42
48
Launch-period reports cite 91.8% overall task accuracy, 3.6 seconds for a GUI operation, support for tasks exceeding 100 steps, and about 90% end-to-end task completion for the joint solution. These are reported figures, not independently verified results from finished Magic9 phones; the cited accounts do not establish a common test setup for all four measures. One account also labels the 91.8% figure as an intent-understanding measure, underscoring the need to check metric definitions before comparing claims.
9
35
46
A separate Qwen-UI-Agent technical report gives a 92.2% success rate on MobileWorld-Real, its real-device benchmark. That is a result for the evaluated agent and benchmark—not a measured success rate for Honor’s forthcoming phones.
33