Both systems use AMD’s Ryzen AI Max+ PRO 495 and shared LPDDR5X memory to make large local models possible without a discrete GPU. Do not treat their 131 TOPS platform figure as an LLM speed result: available memory determines whether a model can fit, but practical tokens per second also depends on quantization, con...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: How do Lenovo’s ThinkCentre X Ultra and Acemagic’s F9A mini workstations, unveiled at IFA 2026, use AMD’s Ryzen AI Max+ Pro 495 “Gorgon Halo. Article summary: Both machines use AMD’s Ryzen AI Max+ PRO 495 “Gorgon Halo” APU to avoid the usual split between system RAM and GPU VRAM: CPU, Radeon 8065S iGPU, and AI resources share one fast LPDDR5X pool. That makes high-quantization. Topic tags: general, general web, news. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers
AMD’s Ryzen AI Max+ PRO 495 gives these compact workstations an unusual advantage for local AI: the CPU, Radeon 8065S integrated graphics and AI resources draw from the same LPDDR5X memory pool. Instead of being constrained by a separate, fixed amount of discrete GPU VRAM, a local-model runtime can access a much larger shared-memory budget. 4
42
That is the central appeal of Lenovo’s ThinkCentre X Ultra and Acemagic’s F9A. They are not interchangeable, however. Lenovo is positioning its 1.6-liter system as a managed business workstation with a stated scale-out option, while Acemagic is pursuing the biggest announced single-machine memory configuration.
| Lenovo ThinkCentre X Ultra | Acemagic F9A PRO 495 | |
|---|---|---|
| Processor | Up to Ryzen AI Max+ PRO 495 with Radeon 8065S graphics |
Ryzen AI Max+ PRO 495 with Radeon 8065S graphics |
| Maximum unified memory | 128GB LPDDR5X |
192GB LPDDR5X-8533 |
| Maximum graphics-memory allocation | Up to 96GB allocated from unified memory |
Up to 160GB allocated to graphics, according to reported specifications |
| Form factor | 1.6 liters |
About 2 liters; 158.5 × 158.5 × 81.5 mm |
| Storage | Up to two 4TB M.2 2280 PCIe Gen5 SSDs, or 8TB total |
One M.2 2280 PCIe 4.0 x4 slot supporting up to 4TB has been reported |
| Networking and expansion | Enterprise-oriented configuration with AMD PRO and Lenovo security/manageability features; detailed I/O varies by configuration |
Dual 2.5GbE, Wi‑Fi 7, two USB4 ports, HDMI 2.1, DisplayPort 2.1 and OCuLink |
| Software position | Windows 11 and Linux compatibility, plus Lenovo’s AI Developer Center and business-management stack |
Local-AI-focused hardware announcement; equivalent enterprise fleet-management commitments were not specified in the launch material |
| Availability and price | Expected from November 2026, starting at €3,100 |
Price, regions and retail date had not been announced at the IFA showcase |
A conventional desktop with a discrete GPU separates system memory from GPU VRAM. For local inference, model weights and the growing KV cache for a long conversation must fit into the memory accessible to the accelerator running the model. That makes VRAM capacity a common limit.
The PRO 495 design replaces that hard division with shared memory. Acemagic explicitly says its CPU, integrated Radeon graphics and AI resources share the same pool, which it positions for local AI inference and other memory-intensive workloads. 4 Lenovo similarly specifies up to 128GB of unified memory, with up to 96GB allocated as graphics memory.
42
This does not turn shared LPDDR5X into the equivalent of high-end discrete GPU memory. Its value is primarily capacity and flexibility: it can enable larger quantized models, longer contexts or more KV-cache headroom than a small discrete-VRAM configuration. Actual generation speed remains dependent on the model format, inference engine, prompt/context size, memory bandwidth and sustained power behavior.
Both machines pair the PRO 495 with Radeon 8065S integrated graphics. Acemagic lists the top configuration with 16 CPU cores and 32 threads, up to 5.2GHz boost, and a 55-TOPS NPU; AMD’s quoted total platform figure is up to 131 TOPS. 4
7 Lenovo’s announced system also reaches a 55-TOPS NPU and uses the same Radeon 8065S graphics.
33
42
Those shared headline specifications mean the biggest difference is not the chip—it is the amount of memory each vendor makes available.
The F9A’s maximum 192GB LPDDR5X-8533 configuration is its defining feature. Reported specifications say up to 160GB can be assigned to the integrated GPU. 2
3 That extra 64GB over Lenovo’s maximum system memory can matter substantially when a model, its context window and its KV cache compete for the same pool.
Acemagic has promoted the system for models up to 120B in Q4 quantization and 300B mixture-of-experts models. 2 These should be read as vendor fit claims, not as independent evidence of interactive performance. A model that loads successfully may still generate too slowly for a particular use case, especially with a long context or a demanding runtime.
The F9A also has enthusiast-friendly expansion: two USB4 ports, dual 2.5GbE, Wi‑Fi 7, display outputs and OCuLink. 4 OCuLink is particularly notable because it offers a route to external PCIe expansion, including an external GPU, although the usefulness of that setup depends on software and the attached hardware.
Lenovo’s X Ultra tops out at 128GB of unified memory and up to 96GB of allocated graphics memory. 42 That is less headroom than Acemagic’s maximum configuration, but it remains a large pool for a desktop of this size.
Its differentiator is a stated cluster option: Lenovo says an X Ultra can connect to as many as three additional units for larger models and multi-agent workloads. 44 Four fully configured systems would provide up to 512GB of aggregate unified memory.
That aggregate figure should not be mistaken for one coherent 512GB GPU-memory space. Multi-node inference requires software that can divide work or model layers across systems, and network communication adds overhead. The cluster may expand what can be run locally, but it does not guarantee the latency or simplicity of a single accelerator with the same amount of directly attached memory.
Lenovo also supplies a clearer business platform around the hardware, including AMD PRO support, ThinkShield security and DASH manageability. 36 It lists Windows 11 and Linux compatibility, and its expected November 2026 launch and €3,100 starting price give prospective buyers a more concrete procurement target.
42
45
The 55-TOPS NPU rating and up-to-131-TOPS platform figure are useful indicators of theoretical supported AI capability, but they are not a benchmark for chatbot response speed or model quality. Neither announcement supplies independent, comparable tokens-per-second results across common LLMs.
For a buying decision, separate three questions:
For one large quantized model in a compact box, the F9A’s 192GB maximum is the more ambitious specification. For an organization that values a known launch timetable, Linux/Windows options, management features and a vendor-described cluster setup, Lenovo’s X Ultra is the more structured proposition.
The two AMD systems arrive as Nvidia is expanding RTX Spark into Windows laptops and compact desktops. Nvidia said RTX Spark systems from ASUS, Dell, HP, Lenovo, Microsoft Surface and MSI would be available in fall 2026, with Acer and Gigabyte systems to follow. 32 Reuters likewise reported that the platform would appear in laptops and compact desktops from major OEMs, including Dell, HP and Lenovo.
17
RTX Spark is therefore a meaningful competitive direction, especially for buyers who prioritize Nvidia’s GPU software ecosystem. But it is not yet a direct, model-for-model substitute for either machine: final OEM configurations, pricing and independent performance data vary by device and still need validation. 23
The broader takeaway is that compact AI workstations are beginning to compete on memory capacity as much as raw accelerator throughput. In this segment, the ability to keep a model local—and keep enough memory free for useful context—can be more important than a headline compute number.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Both systems use AMD’s Ryzen AI Max+ PRO 495 and shared LPDDR5X memory to make large local models possible without a discrete GPU.
Both systems use AMD’s Ryzen AI Max+ PRO 495 and shared LPDDR5X memory to make large local models possible without a discrete GPU. Do not treat their 131 TOPS platform figure as an LLM speed result: available memory determines whether a model can fit, but practical tokens per second also depends on quantization, context length, runtime and therma...
Lenovo is the more clearly specified business option; Acemagic is the capacity first choice, but its final pricing, regions and shipping schedule were not announced at launch.