Home / News EN / First AMD Gorgon Halo PCs Offer 192 GB Unified RAM for AI

First AMD Gorgon Halo PCs Offer 192 GB Unified RAM for AI

AMD Gorgon Halo PCs. The AMD Gorgon Halo platform introduces 192 GB of unified memory in mini PCs. The GMKtec EVO-X5 Pro is capable of running large language models (LLMs) with up to 320 billion parameters locally, but the cost is considerable.

AMD Gorgon Halo PCs

In the race for local artificial intelligence, the amount of available memory plays a role as crucial as processor power. The new GMKtec EVO-X5 Pro demonstrates this; it is one of the first systems based on the AMD Gorgon Halo platform and features 192 GB of unified LPDDR5X-8533 memory.

AMD Gorgon Halo PCs: why it matters

The manufacturer initially designed the machine to handle models with 300 billion parameters locally, but with recent optimizations claims to have reached up to 320 billion parameters completely offline. This capability radically changes the type of models executable on a desktop computer: with 4-bit quantization, such a configuration can host models that until recently required server infrastructure or professional accelerators with much larger amounts of memory.

GMKtec explicitly cites local execution of LLMs up to 320 billion parameters, although compatibility and speed depend on the model used, the level of quantization, and software optimization.

The main hardware component is the Ryzen AI Max+ PRO 495, a solution that AMD places at the top of its Gorgon Halo platform. The chip integrates 16 Zen 5 cores and 32 threads, reaches a boost frequency of 5.2 GHz, and pairs the CPU with a Radeon 8065S featuring 40 Compute Units based on RDNA 3.5 architecture. Completing the picture is an XDNA 2 NPU capable of reaching 55 TOPS.

The strength of Gorgon Halo is not a revolutionary increase in CPU performance, but memory: the 192 GB shared between CPU and GPU can be used more flexibly than the classic configuration of a PC with separate system RAM and VRAM. In the case of the EVO-X5 Pro, up to 160 GB can be assigned to the GPU, leaving the rest for the operating system and applications.

This architecture is particularly suitable for LLM inference. When running a language model locally, a huge portion of the required resources serves simply to keep the model weights in memory. At equal quantization levels, increasing memory means being able to load much larger models without resorting to a professional graphics card with a high amount of VRAM.

What changes and what are the effects

The LPDDR5X reaches 8,533 MT/s, with a declared bandwidth of approximately 273 GB/s. This figure is higher than the 256 GB/s of the previous Strix Halo platform, but not sufficient to transform Gorgon Halo into a drastically faster system for inference. The main advantage therefore remains capacity, while the increase in bandwidth may translate into a more contained improvement in token generation speed.

The EVO-X5 Pro also features three M.2 PCIe 4.0 slots for a total declared capacity of up to 24 TB, two 10 GbE Ethernet ports, USB4, and support for external GPUs and multi-system configurations. Cooling utilizes a vapor chamber paired with three fans, offering three operating modes to manage prolonged loads. The system also integrates AMD DASH for remote management and a dedicated TPM 2.0 module, features that make it more interesting even for small professional environments.

All this memory comes at a high cost, however. The EVO-X5 Pro model with 192 GB of RAM and a 2 TB SSD costs €6,599, while the version with 4 TB reaches €6,899. The issue is linked primarily to the availability of LPDDR5X memory, whose demand has been enormously driven by the artificial intelligence boom.

Gorgon Halo therefore brings to the desktop a memory capacity sufficient to tackle language models with hundreds of billions of parameters, but it does so at a price that clearly distances it from the regular PC market. NVIDIA’s competition is moving in the same direction with systems based on GB10, but even in that case, rising prices for AI workstations are making it increasingly expensive to bring advanced inference out of data centers.

Source and further reading on AMD Gorgon Halo PCs: original article.

* Content created with the assistance of artificial intelligence systems.