Microsoft's Build 2026 conference focused on turning Windows into the premier local AI development platform, headlined by the new Aion 1.0 on device AI models, a significant expansion of Windows AI APIs to run across... Aion 1.0 Instruct and Aion 1.0 Plan bring small language model capabilities directly into Windows...

Create a landscape editorial hero image for this Studio Global article: What did Microsoft announce at Build 2026 regarding its Aion 1.0 on-device AI models, the new Windows AI APIs that run across CPUs, GPUs, an. Article summary: At Build 2026 on June 2–3, Microsoft made three key announcements covering on-device AI models, expanded Windows AI APIs, and a dedicated developer workstation.. Topic tags: general, documentation, general web, user generated. Reference image context from search candidates: Reference image 1: visual subject "Microsoft expected to unveil new AI models and windows enhancements at build 2026: Report "Microsoft expected to unveil new AI models and windows enhancements at build 2026: Report" source context "Microsoft Expected to Unveil New AI Models and Windows ..." Reference image 2: visual subject "Sign up with your email below to instantly access member features,
At its annual Build developer conference on June 2-3 in San Francisco, Microsoft laid out a comprehensive vision for making Windows the go-to platform for local AI development. The strategy rests on three pillars: new on-device AI models to power features directly on your PC, a hardware-agnostic API layer that works across silicon from every major vendor, and a purpose-built developer workstation that brings serious compute power to the desktop.
Microsoft introduced two new small language models (SLMs) under the Aion 1.0 brand, signaling a clear move to embed AI capabilities directly into the operating system .
Aion 1.0 Instruct is the everyday workhorse. It's a next-generation SLM that is smaller, faster, and more efficient than the previous Phi-4-mini model. Crucially, it runs on a much broader range of hardware, including PCs with less-capable GPUs and CPUs, not just those with dedicated Neural Processing Units (NPUs) . It is available now in developer preview, with an open-source release planned for Hugging Face in July
. The model handles common tasks like text summarization, rewriting, and intent detection
.
Aion 1.0 Plan is the more ambitious release. This is a reasoning and tool-calling model designed to enable fully local, on-device agentic capabilities. For the first time, a reasoning model capable of orchestrating sub-agents will be built directly into Windows, allowing complex AI workflows to run without a cloud connection . One source reports the model has 14 billion parameters
. Both Aion models are expected to be broadly available "in the coming months"
.
Until now, many of Windows' advanced AI features were locked to Copilot+ PCs with dedicated NPUs. Microsoft shattered that limitation at Build 2026 by expanding its Windows AI APIs to run on CPUs and GPUs, opening up on-device AI to a massive new install base across hardware from AMD, Intel, Qualcomm, and Nvidia .
The expansion is immediate and tangible. Phi Silica, Microsoft's on-device language model, is now available on GPUs. Video Super Resolution (VSR) and live captions can now run on CPUs . A new Speech Recognition API, available in preview, delivers real-time, offline speech-to-text on the NPU or CPU, enabling dictation, captioning, and accessibility tools without an internet connection
. This API is initially launching in English
.
Underpinning this cross-platform strategy is Windows ML, the built-in AI inferencing runtime that handles model execution across any available processor. At Build 2026, Microsoft integrated Windows ML with Microsoft Foundry on Windows, creating a full lifecycle platform for model selection, fine-tuning, optimization, and deployment . To reinforce the hardware ecosystem, AMD announced updated NPU and GPU Execution Providers for Windows ML, delivering up to a 1.5x improvement in time-to-first-token on NPUs and a 3.5x increase in sustained token generation
.
The hardware hero of Build 2026 was the Surface RTX Spark Dev Box, a compact desktop workstation designed for serious local AI workloads . Its design draws comparisons to the Xbox Series X—compact, aluminum-bodied, and fanless, with the chassis itself acting as a 100W heatsink
.
The core of the machine is Nvidia's new RTX Spark ARM-based superchip, which delivers up to 1 petaflop of AI compute paired with 128 GB of unified memory shared between CPU and GPU . This gives developers enough horsepower to run large language models with up to 120 billion parameters locally, with support for a 1-million-token context window
.
It ships with a purpose-built, developer-optimized Windows 11 Pro configuration. Visual Studio Code, GitHub Copilot in Windows Terminal, Windows Subsystem for Linux (WSL), and PowerShell 7 all come preconfigured and ready out of the box . Microsoft says the machine is "designed for sustained workloads: long-running training jobs, agentic AI pipelines and local model fine-tuning"
. The Surface RTX Spark Dev Box is expected to ship in the US later this year; pricing has not yet been announced
.
Studio Global AI
Use this topic as a starting point for a fresh source-backed answer, then compare citations before you share it.
Microsoft's Build 2026 conference focused on turning Windows into the premier local AI development platform, headlined by the new Aion 1.0 on device AI models, a significant expansion of Windows AI APIs to run across...
Microsoft's Build 2026 conference focused on turning Windows into the premier local AI development platform, headlined by the new Aion 1.0 on device AI models, a significant expansion of Windows AI APIs to run across... Aion 1.0 Instruct and Aion 1.0 Plan bring small language model capabilities directly into Windows, with the latter enabling fully local agentic reasoning and tool calling for the first time.
The Surface RTX Spark Dev Box, powered by Nvidia's new ARM based superchip, puts up to 1 petaflop of AI compute and 128GB of unified memory on a developer's desk, capable of running 120 billion parameter models locally.