GEEKOM is pushing compact computing into territory traditionally associated with AI servers, deploying DeepSeek V4 Flash across four GEEKOM A9 Mega Mini PCs connected through USB4. The configuration uses AMD’s Ryzen AI Max+ 395 processors, local memory and distributed inference software to demonstrate how organizations could run advanced AI workloads closer to their data without relying entirely on cloud infrastructure.
The next phase of enterprise AI infrastructure may not always require a rack of servers.
GEEKOM, a manufacturer of high-performance Mini PCs, has demonstrated a distributed AI cluster built from four GEEKOM A9 Mega systems running DeepSeek V4 Flash. Rather than connecting conventional server hardware through a specialized high-speed networking fabric, the configuration links the compact computers through USB4 and distributes model inference across them.
The experiment is an example of a broader shift in AI infrastructure: moving inference closer to the data and experimenting with smaller, modular systems rather than depending exclusively on centralized cloud GPUs.
Each A9 Mega uses AMD’s Ryzen AI Max+ 395, with 16 Zen 5 CPU cores and Radeon 8060S graphics alongside unified memory. The systems run Ubuntu and AMD’s ROCm software stack, while DwarfStar handles distribution of the optimized DeepSeek V4 Flash model across the four machines.
An OpenAI-compatible API provides a familiar interface for applications and AI agents.
The hardware configuration matters because it addresses one of the biggest challenges in enterprise AI adoption: organizations increasingly want access to capable models without necessarily sending proprietary information to external cloud services.
Local AI moves beyond the workstation
Running AI locally is not a new idea. Workstations and specialized edge computers have been performing inference for years.
What is changing is the scale of models organizations expect to run.
Large language models can process internal documentation, source code, customer records and operational information, but sending that data to a public AI service can create security, privacy and compliance concerns. Local inference offers another architecture: keep the model and data within infrastructure controlled by the organization.
GEEKOM’s cluster is aimed at that use case.
The company says businesses can use the system for private knowledge assistants, document analysis, source-code review, retrieval-augmented generation (RAG), multi-source research and controlled workflow automation.
For enterprises, the attraction is not necessarily that a Mini PC replaces a conventional AI server. Instead, it could provide a modular inference layer that can begin small and expand as workload requirements increase.
An organization could deploy one system for experimentation, add another as demand grows and eventually distribute inference across four machines.
That approach differs from the traditional cloud model, where capacity is generally purchased as a service, and from enterprise AI clusters that can require significant capital investment in networking, GPUs, power and cooling.
AMD’s integrated architecture is important
The A9 Mega’s processor is central to the demonstration.
The Ryzen AI Max+ 395 combines CPU and GPU resources with unified memory, giving AI workloads access to a shared memory architecture rather than relying exclusively on discrete accelerators.
That architecture is increasingly relevant as AI workloads move beyond simple chatbot interactions.
Long-context inference, code analysis and RAG can require substantial memory capacity because the system must maintain large amounts of prompt and retrieved information during generation.
GEEKOM says its configuration has operated with contexts of up to 250,000 tokens, although actual usable context depends on the model, software configuration and workload.
The company’s reported single-concurrency results were approximately 14.61 tokens per second, with a P95 time to first token of about 0.42 seconds in its 32- and 128-token tests.
Those numbers should be viewed as vendor-reported results rather than a universal benchmark. Performance can change significantly depending on quantization, prompt length, model configuration, concurrency, software versions and the specific workload.
Still, the demonstration points to an important trend: AI performance is increasingly about more than raw token-generation speed.
For enterprise applications, memory capacity, latency, data locality and total cost of ownership can matter as much as peak throughput.
USB4 turns networking into part of the experiment
Perhaps the most interesting technical aspect is how the systems communicate.
GEEKOM says the four Mini PCs use USB4 to create the distributed platform without a proprietary high-speed switch or conventional server rack.
That simplifies the physical architecture, but it also raises questions about how such configurations compare with dedicated AI networking technologies used in large clusters.
Modern AI data centers rely on high-bandwidth, low-latency interconnects to move model and activation data between accelerators. Distributed inference across smaller machines has different requirements, and workload partitioning becomes critical.
The value of GEEKOM’s approach is therefore likely to be greatest for organizations whose workloads can tolerate the characteristics of a smaller distributed cluster.
It could be particularly relevant to edge AI, private AI laboratories, education, development teams and small enterprises that need more capability than a single workstation but do not require a traditional data-center deployment.
AI agents create another use case
The cluster is also designed to support AI agents.
GEEKOM says systems such as Hermes Agent can process tools, policies, memory, logs, code and retrieved information locally before taking action.
That is an important distinction from basic text generation.
An agent can interact with software systems, retrieve information and execute tasks. Keeping those operations within local infrastructure could reduce the exposure of sensitive prompts, credentials, source code and intermediate results.
It does not eliminate security risks. Local AI infrastructure still requires identity management, access controls, model security, software updates and monitoring.
But the architecture illustrates why private AI infrastructure is becoming an important enterprise category.
A different path to AI infrastructure
The broader market remains dominated by hyperscale cloud platforms and specialized AI hardware from companies such as NVIDIA, AMD, Intel, Microsoft, Google and Amazon.
Large organizations with demanding workloads will continue to rely heavily on data-center infrastructure. NVIDIA’s GPU platforms, for example, are designed around precisely the high-bandwidth networking and acceleration requirements that large-scale AI training and inference demand.
GEEKOM’s proposition is different.
It is demonstrating that several relatively compact, independently useful systems can be assembled into a distributed AI platform. That modularity could appeal to organizations that prioritize physical control, incremental deployment and data locality.
The bigger story is not that Mini PCs are replacing AI data centers.
It is that AI infrastructure is becoming more heterogeneous.
Cloud GPUs, enterprise servers, edge accelerators, workstations and compact distributed systems can coexist, with organizations choosing where inference happens according to latency, privacy, cost and workload requirements.
GEEKOM’s DeepSeek deployment is a small but revealing example of that shift. As capable open and locally deployable models become more accessible, the question for businesses is increasingly not whether they can run AI—but where they want to run it, what data they are willing to move, and how much infrastructure they actually need.
Market Landscape
AI infrastructure is increasingly dividing into three broad layers: hyperscale cloud computing, enterprise/on-premises AI infrastructure and edge/local AI.
Cloud platforms from Microsoft Azure, Amazon Web Services and Google Cloud offer virtually unlimited access to accelerated computing but require organizations to manage data, networking and cloud costs within an external environment.
At the other end, local AI systems provide greater physical control and can reduce data movement. The trade-off is that organizations assume responsibility for hardware procurement, software configuration, security and capacity planning.
GEEKOM’s four-system cluster occupies an emerging middle ground: modular local inference built from relatively compact hardware.
The approach could become more attractive as AI models become smaller and more efficient through quantization, distillation and architecture improvements. Meanwhile, advances in integrated CPU/GPU designs from AMD and others are making capable AI processing possible outside traditional GPU servers.
The main enterprise question will be workload economics. A compact cluster makes the most sense where privacy, predictable workloads, low-latency access or incremental capacity outweigh the convenience of cloud AI APIs.
Top Insights
- GEEKOM is distributing DeepSeek V4 Flash across four A9 Mega Mini PCs, demonstrating a modular approach to local AI inference outside conventional data-center infrastructure.
- The cluster combines AMD Ryzen AI Max+ 395 processors, unified memory, ROCm and USB4 to create distributed computing for enterprise and edge AI workloads.
- Local inference can help organizations keep documents, source code, credentials and AI-agent workflows inside controlled infrastructure rather than sending sensitive data to public clouds.
- Reported performance reached approximately 14.61 tokens per second at single concurrency, with context windows reaching up to 250,000 tokens in the company’s configuration.
- The deployment highlights growing demand for heterogeneous AI infrastructure spanning cloud platforms, enterprise servers, workstations and compact edge computing systems.
Power Tomorrow’s Intelligence — Build It with TechEdgeAI









