AMD announced Gorgon Halo GPU pricing at $6,799 for the 192GB variant, highlighting the cost premium of high-memory accelerators for enterprise inference.
The first systems powered by AMD's Gorgon Halo system-on-chip (SoC) platform have arrived with up to 192 GB of unified memory on board—enough capacity to run DeepSeek V4 Flash locally, provided the cost doesn't prove prohibitive. GMKtec's EVO-X5 Pro is among the first to feature the Ryzen AI Max+ 495 SoC, with pricing starting at $6,799. For local AI enthusiasts and privacy-conscious small businesses, the capability to run larger, more capable models from their own offices and homes may justify the expense.
At 4-bit precision, Gorgon Halo can support models up to around 345 billion parameters on Linux or 320 billion parameters on Windows—a difference driven by memory partitioning. Windows requires allocating 160 GB to the GPU, whereas Linux and AMD's GPU drivers allow users to leverage nearly all available memory. This puts frontier models like DeepSeek V4 Flash (284 billion parameters) and Z.AI's GLM-5.3-Flash (320 billion parameters) within practical reach.
Announced earlier this year, Gorgon Halo is essentially a factory-overclocked version of the Strix Halo APU that powers AMD's AI Halo workstation. Both chips feature up to 16 Zen 5 cores, a 40-compute-unit integrated GPU, and an XDNA 2-based NPU. The key difference lies in clock speeds—Gorgon Halo is clocked approximately 100 MHz higher—and memory capacity. Gorgon Halo supports 192 GB of LPDDR5x memory at 8,533 MT/s, compared to Strix Halo's 128 GB at 8,000 MT/s. This clock bump alone is unlikely to yield meaningful performance gains, though faster memory could modestly improve LLM inference. Because LLM inference is memory-bound, faster memory enables quicker token generation. Gorgon Halo achieves 273 GB/s of memory bandwidth, roughly 6.5 percent faster than Strix Halo's 256 GB/s—enough to add a few tokens per second on well-optimized models, but nothing dramatic.
The platform's primary appeal is higher memory capacity, but the semiconductor supply squeeze has inflated costs across the market. A year ago, a 128 GB Strix Halo system cost between $2,000 and $3,000; by summer, when AMD launched its AI Halo workstation, prices had climbed to $4,000. GMKtec's 192 GB unit is now 64 percent more expensive, and the company is typically priced at the affordable end of the spectrum. Premium brands like Framework and notebooks such as HP's ZBook Ultra G3a will likely command even steeper premiums.
AMD has little room to maneuver. LPDDR5x is in severe short supply due to the AI boom. Each NVIDIA NVL72 rack consumes between 36 and 54 TB of this memory for CPU alone, and the situation will worsen as AMD deploys LPDDR5x in its AI-optimized Verano Epyc processors. Samsung warned in early 2026 that memory supplies will likely remain constrained through 2028—a problem affecting NVIDIA as well. NVIDIA has seen prices for its GB10 SoC-based AI workstations roughly double over the past year, rising from $3,000–$4,000 to $6,000–$8,000. Announced at Computex this spring, NVIDIA's RTX Spark family offers up to 128 GB of memory in both mini PCs and notebooks, and these systems are expected to command a premium.
AMD's higher memory capacity at a lower price point offers a theoretical advantage, yet the practical benefit evaporates when both products are priced beyond reach.