125B model on 12GB GPU. An open source project has enabled the execution of the Qwen3.8-Flash-Next model, which features 125 billion parameters, on graphics cards equipped with 12 GB of VRAM.

The system distributes the workload across the GPU, system memory, and SSD, achieving a speed of 94 tokens per second measured on a GeForce RTX 5070. The project, available on GitHub under an MIT license, is based on components from llama.cpp and ggml.
125B model on 12GB GPU: why it matters
The minimum requirements for execution include a graphics card with at least 12 GB of VRAM, 32 GB of system RAM, and approximately 80 GB of free disk space, preferably on an SSD. The model to be downloaded is approximately 70 GB in size.
Tests were conducted on two different configurations. On a 12 GB GeForce RTX 5070 paired with a Ryzen 5 7600 and 64 GB of RAM, the Q2_0 quantization offers 94 tokens/s in response and a prompt ingestion speed of 2,650 tokens/s. With IQ2_XS, it drops to 79 tokens/s; with IQ3_XXS, to 62 tokens/s; and with IQ3_S, considered the most accurate but also the slowest variant, to 53 tokens/s.
A 16 GB Radeon RX 9070 XT, paired with a Ryzen 9 3900X and 47 GB of RAM, reaches 60 tokens/s with Q2_0, 52 with IQ2_XS, and 44 with the Coder build. Those with 24 GB of video memory, such as on the GeForce RTX 3090, might achieve a speed between 100 and 140 tokens/s, but this is an estimate.
The model utilizes an expert architecture composed of 24,576 specialized modules, of which only 10 are activated for each token. The project keeps the few thousand most frequently used experts on the graphics card, parks the entire set in RAM, assigns residual work to the processor, and offloads a lookup table to the SSD.
Speculative decoding is also employed: an auxiliary model proposes subsequent words that are verified in bulk by the main model, with a declared gain between 1.6 and 1.8 times at equal response quality. Long text reading occurs in batches of up to 8,192 tokens at a time, exceeding one thousand tokens per second.
What Changes and What Are the Effects
The choice of quantization depends on the amount of installed memory. With 32 GB, the Coder variant is proposed, obtained by removing half of the experts; it reaches 91% of the SWE-bench Verified score of the full model but performs less well with CJK texts and code. With 64 GB, all sizes are available; Unsloth’s UD-Q4_K_XL, which is closer to the original, is read primarily from SSD and drops to 7-8.5 tokens/s.
Access occurs via browser at 127.0.0.1:8080, with chat and resource monitoring, or via API through OpenAI-compatible (/v1), Anthropic (/v1/messages), and Responses API (/v1/responses) endpoints. The system processes one request at a time, and image processing on AMD cards works only under Linux.
At startup, the system loads between 35 and 55 GB into memory and may appear frozen for one to three minutes. Processing the first message of a conversation requires approximately one minute per 30,000 tokens.
Source and further reading on 125B model on 12GB GPU: original article.
* Content created with the assistance of artificial intelligence systems.
Hardware Ready Ready to Bench?