Running GLM-5.3-Flash on CMP 170HX — a sizing guide

NVIDIA CMP 170HX 64 GB: a full-length passive card with a silver metal shroud, black vented end bracket and gold PCIe edge connector — photographed on our bench

Read this before the numbers. Every tokens-per-second figure, every RAM minimum and every "fits / does not fit" on this page is calculated from memory bandwidth, not measured. We publish it because the arithmetic is useful for choosing parts, and because nobody else seems to have written it down. We will replace it with real figures once we have run the model — the benchmark is planned on an 8-channel, 512 GB machine and has not happened yet. Plan for the lower end of every range.

GLM-5.3-Flash runs on one or two unlocked CMP 170HX 64 GB cards when the rest of the model sits in system RAM. With two cards and 128 GB of RAM it should run at good quality; the amount and speed of system memory decide how fast it answers. From four cards on, the good-quality model fits entirely on the cards and system RAM stops mattering.

What we have actually measured

These are ours, from our own bench and from a four-card system we commissioned in September 2026. They are the foundation the estimates are built on.

Measured on our hardware Result
Unlocked capacity per card 65,536 MiB (64 GB), confirmed across many cards
Peak VRAM bandwidth 931–970 GB/s (memtest_vulkan; four-card system read 931.3 / 944.2 / 947.3 / 970.0)
PCIe link state with our x16 modification Gen2 ×16
Sustained power draw 244–245 W held over a 60-minute soak (250 W board limit)
ECC None. Fused off — the card neither corrects nor reports memory errors

Everything below this line is modelling.

How it works

Each word uses only 18B of the 320B parameters: the card holds the always-used layers and whatever experts fit, system RAM holds the rest

GLM-5.3-Flash is a mixture-of-experts model: 320 billion parameters in total, but only 18 billion are used for each word. The card holds the parts every word needs; the rest stays in system RAM, and only the few experts each word picks are read from there.

Those parameter counts and the 288-experts-per-layer layout are Z.ai's published figures, read from the model card and its config.json. They are not ours to verify.

Pick a build

Two cards with 128 GB of RAM is the build to aim for: good answer quality at a usable speed. The other two are a cheaper start and a single-card option.

Build Cards System RAM Model quality (file) Estimated speed
Recommended 2× 170HX 128 GB Good — UD-Q4_K_XL (200 GB) 8–20 tokens/s, set by the RAM platform
Cheapest that works 2× 170HX 64 GB Usable, noticeably weaker — UD-Q2_K_XL (109 GB) 40–55 tokens/s (the whole model fits on the two cards)
Single card 1× 170HX 192 GB, 8-channel platform Good — UD-Q4_K_XL (200 GB) about 15 tokens/s

A token is roughly three quarters of a word. 10 tokens/s reads about as fast as a person reads; 20+ feels quick.

How much system RAM you need

A small, fully lit block beside a much larger block with only a few cells lit — the card holds what every word needs, system RAM holds the rest

RAM needed = the part of the model that does not fit on the cards + about 16 GB for the operating system. One card leaves about 46 GB of its memory for model weights; two cards leave about 105 GB.

Model quality File size 1× 170HX: min / recommended RAM 2× 170HX: min / recommended RAM
UD-Q2_K_XL — usable, noticeably weaker 109 GB 64 GB / 96 GB 32 GB / 64 GB
UD-IQ3_XXS — better 120 GB 96 GB / 128 GB 48 GB / 64 GB
UD-Q4_K_XL — good (recommended) 200 GB 160 GB / 192 GB 96 GB / 128 GB
Q8_0 — close to the original 341 GB 320 GB / 384 GB 256 GB / 256 GB

The smallest version (93 GB) does not fit on one card alone, so a single-card system always needs a large RAM pool. File sizes are unsloth's published GGUF sizes, not ours.

Expected speed

Estimated answer speed by memory platform: an 8-channel platform is about three times a desktop

The number of memory channels matters more than the CPU model. A second card helps most on desktop platforms, because it moves a bigger share of the model out of slow RAM.

More cards: 3 to 6

From three cards on, the whole model fits in card memory, so system RAM no longer limits speed.

Cards Card memory Largest model that fits on the cards Quality System RAM Speed, one user (estimate)
1 64 GB none — good quality needs RAM offload good, with offload 192 GB, 8-channel 14–17 tokens/s
2 128 GB UD-Q2_K_XL (109 GB) usable, weaker 32–64 GB 40–55 tokens/s
3 192 GB UD-IQ4_XS (157 GB) good, slightly below Q4 64 GB 30–40 tokens/s
4 256 GB UD-Q4_K_XL (200 GB), ~30 GB left for long context good 64 GB 27–34 tokens/s
5 320 GB UD-Q5_K_XL (240 GB) very good 64 GB 22–28 tokens/s
6 384 GB UD-Q6_K_XL (292 GB), or the official FP8 release (328 GB) close to original / original 128 GB 18–23 tokens/s
  • More cards buy quality and room, not speed per word. The cards work one after another, and a higher-quality file reads more data for each word, so speed falls slightly from four cards to six.
  • Four cards is the sweet spot for one user: good quality, about 30 tokens/s, no fast-RAM platform needed.
  • Six cards suit a team. The official FP8 weights fit and a vLLM server could answer several users at once. Space for context is tight at about 37 GB, and vLLM support for this model still needs checking.
  • What changes from three cards: a workstation or server board with one ×16 slot per card (EPYC or Threadripper; ×8 slots also work, only loading gets slower), a power supply of about 250 W per card plus 300 W (about 1.3 kW for four, 1.8–2 kW for six), and a chassis with airflow for every card.

These estimates assume each card keeps about 58 GB for model weights and reaches about 35% of its 950 GB/s on this model — typical for similar models on cards of this generation, but an assumption, not a result.

System checklist

The rows that most often go wrong are memory channels, Above 4G decoding and cooling.

Two flat passive cards in a tower with a sheet-metal duct carrying air from three case fans straight through their heatsinks
Item 1× 170HX 2× 170HX
CPU 8+ cores with AVX2 (AVX-512 helps); 16+ cores if RAM holds most of the model same
System RAM see the RAM table; 192 GB for good quality see the RAM table; 128 GB for good quality
Memory channels as many as possible — an 8-channel platform is about 3× a desktop same
Mainboard 1× PCIe ×16 slot, Above 4G decoding enabled (the card exposes 64 GB) 2× ×16 slots with one free slot between them, Above 4G decoding enabled
PCIe link Gen2 ×16 with the ×16 modification is enough same; the model is split by layers, so little data moves between cards
Power supply 750 W (the card draws 250 W) 1,000–1,200 W
Power connector 1× 8-pin CPU-style (EPS) per card; adapter supplied 2×
Cooling passive card: a fan duct or blower per card is required one per card; two side by side need a dual duct
Storage NVMe, at least 1.5× the model file (good quality: 512 GB min, 1 TB recommended) same
Operating system Linux same
Driver NVIDIA open driver 610.43 + cmpunlocker 64 GB unlock; re-install after every kernel update same
Inference software llama.cpp or ik_llama.cpp (faster on the CPU side) same

On the PCIe link: we have measured the link state (Gen2 ×16). We have not benchmarked transfer throughput on it. A Gen2 ×16 link tops out around 8 GB/s in theory and real traffic lands nearer 6–7 GB/s after overhead — treat that as derived from the specification, not as a measurement.

What to expect

  • Long prompts start slowly. Reading a long document or a large code base takes seconds to minutes before the first word appears, longer with one card. The Gen2 link is the limit there; answer speed is not affected.
  • Lower-bit versions lose quality. UD-Q2_K_XL works but makes more mistakes than UD-Q4_K_XL. For coding or reasoning, choose the 4-bit version.
  • Long context is cheap. Most layers keep a fixed-size state, so a long conversation costs little extra card memory.
  • New model, young software support. GLM-5.3-Flash uses a new attention design. llama.cpp loads it, but parts may run less efficiently than they could; this improves with updates.
  • The 64 GB comes from a patched driver. After a kernel update the cards fall back to 8 GB until the unlock is re-installed.
  • No ECC. The card does not correct or report memory errors, so use cards that have passed a full memory test. Ours ship with a per-card memory verification report carrying that card's serial number.

Where the numbers come from

Claim Status
320B total / 18B active, 288 experts per layer, 8 per token Z.ai's figures — model card and config.json
Quantized file sizes, 93–341 GB unsloth's published GGUF sizes
All tokens/s figures, RAM minimums, 46 / 105 / 58 GB per-card capacities Kentino estimate — modelled, not measured
Usable RAM bandwidth per platform (35 / 60 / 110 / 300 GB/s) Kentino estimate — typical real-world values
64 GB unlock, 931–970 GB/s, Gen2 ×16 link state, 244–245 W, no ECC Measured by Kentino on our own cards

Sources: zai-org/GLM-5.3-Flash · unsloth/GLM-5.3-Flash-GGUF


commissioned here. Every CMP 170HX we sell is memory-tested at full 64 GB capacity and ships with a report carrying its own serial number. If you want this build sized against your actual workload, write to us.*

Tillbaka till blogg