GEEKOM says the four machines communicate over USB4 and use a software stack consisting of Ubuntu, AMD’s ROCm platform and DwarfStar. The optimized DeepSeek V4 Flash model is partitioned across the nodes, which then serve requests through an OpenAI-compatible endpoint.
That API compatibility matters for deployment. It can allow existing internal applications, agents and developer tools that already support OpenAI-style endpoints to connect to the local model without redesigning their entire integration layer. The intended use case is therefore less about running a chatbot on one desktop and more about creating a private AI service inside an organization’s own environment.
The demonstration does not mean that the cluster has the same characteristics as a conventional data-center server. The supplied announcement does not specify the topology, effective bandwidth, latency, sharding strategy, model precision, quantization format, orchestration settings or failure-recovery behavior. Those details can materially affect inference performance, especially when model-parallel work requires frequent communication between nodes.
Each A9 Mega has two PCIe 4.0 M.2 slots and is specified for up to 8TB of SSD storage. That gives a fully populated four-node installation a theoretical maximum of 32TB, assuming every machine is equipped to its stated capacity. The storage figure is separate from the 512GB of system unified memory and does not indicate how much disk space the DeepSeek deployment itself requires.
The reported inter-node connection is USB4. GEEKOM’s product material also lists USB4 connectivity on the A9 Mega, with ports rated at up to 40Gbps. Those interface figures are not equivalent to measured cluster performance: protocol overhead, topology and the software’s communication pattern determine the bandwidth and latency available to distributed inference.
GEEKOM’s positioning is an on-premises private-AI appliance. In principle, an organization could keep prompts, internal documents and generated outputs within its own network while allowing business applications and agents to use the model through the local API. ROCm provides the AMD software layer for the graphics hardware, avoiding a requirement for an NVIDIA CUDA-based system in this particular design.
That may appeal to teams with data-residency, privacy or network-isolation requirements. It also shifts responsibility to the operator: the organization would need to manage hardware availability, updates, model files, access controls, monitoring, cooling, power and software reliability itself.
GEEKOM lists the A9 Mega at $3,999. Four nodes therefore total approximately $15,996 before adding storage beyond the included configuration, cabling, network equipment, installation, electricity, maintenance or support.
That is a substantial investment for a system whose production characteristics have not yet been independently established. The relevant comparison is not simply the purchase price of four mini PCs, but the complete cost of operating a distributed inference service and the workload it can reliably support.
GEEKOM’s announcement establishes that the company deployed DeepSeek V4 Flash across four A9 Mega systems, but it does not provide independent, standardized results for this exact configuration. There are no supplied figures for sustained tokens per second, time to first token, concurrent users, long-context latency, power consumption or failure behavior.
As a result, the demonstration should be treated as a proof of deployment rather than a validated production benchmark. The model’s one-million-token context capability is an important specification, but it does not show that the four-node A9 Mega cluster can process million-token requests quickly or economically. Real performance would depend on the model build, precision, memory allocation, prompt length, number of simultaneous requests and inter-node communication overhead.
The Ryzen AI Max+ 395 platform is not exclusive to GEEKOM. Other vendors sell mini PCs with the same processor and, in some cases, 128GB of unified memory. Minisforum’s N5 Max is another comparable design, although ServeTheHome reported that its retail version shipped with 64GB of soldered LPDDR5X unified memory rather than the 128GB configuration seen in an earlier prototype.
That makes GEEKOM’s more meaningful differentiator the claimed four-node software demonstration—not the processor alone. Whether the approach is useful beyond a technical showcase will depend on details that remain unpublished: reproducible setup instructions, standardized benchmarks, network measurements, power data and evidence of stable operation under realistic workloads.
GEEKOM’s setup shows a compact way to assemble a large-memory AMD inference platform from four mini PCs: 64 CPU cores, 128 threads, 160 Radeon 8060S compute units and 512GB of unified memory, connected over USB4 and managed with Ubuntu, ROCm and DwarfStar.
The practical promise is local, private access to a very large mixture-of-experts model through a familiar API. But the evidence currently supports the existence of the deployment, not a conclusion about its speed, efficiency or readiness for production. Until GEEKOM publishes reproducible benchmarks and implementation details, the cluster is best understood as an intriguing on-premises demonstration rather than a proven replacement for dedicated AI infrastructure.