Running GLM-5.3-Flash on CMP 170HX — a sizing guide
Condividi

Read this before the numbers. Every tokens-per-second figure, every RAM minimum and every "fits / does not fit" on this page is calculated from memory bandwidth, not measured. We publish it because the arithmetic is useful for choosing parts, and because nobody else seems to have written it down. We will replace it with real figures once we have run the model — the benchmark is planned on an 8-channel, 512 GB machine and has not happened yet. Plan for the lower end of every range.
GLM-5.3-Flash runs on one or two unlocked CMP 170HX 64 GB cards when the rest of the model sits in system RAM. With two cards and 128 GB of RAM it should run at good quality; the amount and speed of system memory decide how fast it answers. From four cards on, the good-quality model fits entirely on the cards and system RAM stops mattering.
What we have actually measured
These are ours, from our own bench and from a four-card system we commissioned in September 2026. They are the foundation the estimates are built on.
| Measured on our hardware | Result |
|---|---|
| Unlocked capacity per card | 65,536 MiB (64 GB), confirmed across many cards |
| Peak VRAM bandwidth | 931–970 GB/s (memtest_vulkan; four-card system read 931.3 / 944.2 / 947.3 / 970.0) |
| PCIe link state with our x16 modification | Gen2 ×16 |
| Sustained power draw | 244–245 W held over a 60-minute soak (250 W board limit) |
| ECC | None. Fused off — the card neither corrects nor reports memory errors |
Everything below this line is modelling.
How it works

GLM-5.3-Flash is a mixture-of-experts model: 320 billion parameters in total, but only 18 billion are used for each word. The card holds the parts every word needs; the rest stays in system RAM, and only the few experts each word picks are read from there.
Those parameter counts and the 288-experts-per-layer layout are Z.ai's published figures, read from the model card and its config.json. They are not ours to verify.
Pick a build
Two cards with 128 GB of RAM is the build to aim for: good answer quality at a usable speed. The other two are a cheaper start and a single-card option.
| Build | Cards | System RAM | Model quality (file) | Estimated speed |
|---|---|---|---|---|
| Recommended | 2× 170HX | 128 GB | Good — UD-Q4_K_XL (200 GB) | 8–20 tokens/s, set by the RAM platform |
| Cheapest that works | 2× 170HX | 64 GB | Usable, noticeably weaker — UD-Q2_K_XL (109 GB) | 40–55 tokens/s (the whole model fits on the two cards) |
| Single card | 1× 170HX | 192 GB, 8-channel platform | Good — UD-Q4_K_XL (200 GB) | about 15 tokens/s |
A token is roughly three quarters of a word. 10 tokens/s reads about as fast as a person reads; 20+ feels quick.
How much system RAM you need

RAM needed = the part of the model that does not fit on the cards + about 16 GB for the operating system. One card leaves about 46 GB of its memory for model weights; two cards leave about 105 GB.
| Model quality | File size | 1× 170HX: min / recommended RAM | 2× 170HX: min / recommended RAM |
|---|---|---|---|
| UD-Q2_K_XL — usable, noticeably weaker | 109 GB | 64 GB / 96 GB | 32 GB / 64 GB |
| UD-IQ3_XXS — better | 120 GB | 96 GB / 128 GB | 48 GB / 64 GB |
| UD-Q4_K_XL — good (recommended) | 200 GB | 160 GB / 192 GB | 96 GB / 128 GB |
| Q8_0 — close to the original | 341 GB | 320 GB / 384 GB | 256 GB / 256 GB |
The smallest version (93 GB) does not fit on one card alone, so a single-card system always needs a large RAM pool. File sizes are unsloth's published GGUF sizes, not ours.
Expected speed

The number of memory channels matters more than the CPU model. A second card helps most on desktop platforms, because it moves a bigger share of the model out of slow RAM.
More cards: 3 to 6
From three cards on, the whole model fits in card memory, so system RAM no longer limits speed.
| Cards | Card memory | Largest model that fits on the cards | Quality | System RAM | Speed, one user (estimate) |
|---|---|---|---|---|---|
| 1 | 64 GB | none — good quality needs RAM offload | good, with offload | 192 GB, 8-channel | 14–17 tokens/s |
| 2 | 128 GB | UD-Q2_K_XL (109 GB) | usable, weaker | 32–64 GB | 40–55 tokens/s |
| 3 | 192 GB | UD-IQ4_XS (157 GB) | good, slightly below Q4 | 64 GB | 30–40 tokens/s |
| 4 | 256 GB | UD-Q4_K_XL (200 GB), ~30 GB left for long context | good | 64 GB | 27–34 tokens/s |
| 5 | 320 GB | UD-Q5_K_XL (240 GB) | very good | 64 GB | 22–28 tokens/s |
| 6 | 384 GB | UD-Q6_K_XL (292 GB), or the official FP8 release (328 GB) | close to original / original | 128 GB | 18–23 tokens/s |
- More cards buy quality and room, not speed per word. The cards work one after another, and a higher-quality file reads more data for each word, so speed falls slightly from four cards to six.
- Four cards is the sweet spot for one user: good quality, about 30 tokens/s, no fast-RAM platform needed.
- Six cards suit a team. The official FP8 weights fit and a vLLM server could answer several users at once. Space for context is tight at about 37 GB, and vLLM support for this model still needs checking.
- What changes from three cards: a workstation or server board with one ×16 slot per card (EPYC or Threadripper; ×8 slots also work, only loading gets slower), a power supply of about 250 W per card plus 300 W (about 1.3 kW for four, 1.8–2 kW for six), and a chassis with airflow for every card.
These estimates assume each card keeps about 58 GB for model weights and reaches about 35% of its 950 GB/s on this model — typical for similar models on cards of this generation, but an assumption, not a result.
System checklist
The rows that most often go wrong are memory channels, Above 4G decoding and cooling.

| Item | 1× 170HX | 2× 170HX |
|---|---|---|
| CPU | 8+ cores with AVX2 (AVX-512 helps); 16+ cores if RAM holds most of the model | same |
| System RAM | see the RAM table; 192 GB for good quality | see the RAM table; 128 GB for good quality |
| Memory channels | as many as possible — an 8-channel platform is about 3× a desktop | same |
| Mainboard | 1× PCIe ×16 slot, Above 4G decoding enabled (the card exposes 64 GB) | 2× ×16 slots with one free slot between them, Above 4G decoding enabled |
| PCIe link | Gen2 ×16 with the ×16 modification is enough | same; the model is split by layers, so little data moves between cards |
| Power supply | 750 W (the card draws 250 W) | 1,000–1,200 W |
| Power connector | 1× 8-pin CPU-style (EPS) per card; adapter supplied | 2× |
| Cooling | passive card: a fan duct or blower per card is required | one per card; two side by side need a dual duct |
| Storage | NVMe, at least 1.5× the model file (good quality: 512 GB min, 1 TB recommended) | same |
| Operating system | Linux | same |
| Driver | NVIDIA open driver 610.43 + cmpunlocker 64 GB unlock; re-install after every kernel update | same |
| Inference software | llama.cpp or ik_llama.cpp (faster on the CPU side) | same |
On the PCIe link: we have measured the link state (Gen2 ×16). We have not benchmarked transfer throughput on it. A Gen2 ×16 link tops out around 8 GB/s in theory and real traffic lands nearer 6–7 GB/s after overhead — treat that as derived from the specification, not as a measurement.
What to expect
- Long prompts start slowly. Reading a long document or a large code base takes seconds to minutes before the first word appears, longer with one card. The Gen2 link is the limit there; answer speed is not affected.
- Lower-bit versions lose quality. UD-Q2_K_XL works but makes more mistakes than UD-Q4_K_XL. For coding or reasoning, choose the 4-bit version.
- Long context is cheap. Most layers keep a fixed-size state, so a long conversation costs little extra card memory.
- New model, young software support. GLM-5.3-Flash uses a new attention design. llama.cpp loads it, but parts may run less efficiently than they could; this improves with updates.
- The 64 GB comes from a patched driver. After a kernel update the cards fall back to 8 GB until the unlock is re-installed.
- No ECC. The card does not correct or report memory errors, so use cards that have passed a full memory test. Ours ship with a per-card memory verification report carrying that card's serial number.
Where the numbers come from
| Claim | Status |
|---|---|
| 320B total / 18B active, 288 experts per layer, 8 per token |
Z.ai's figures — model card and config.json
|
| Quantized file sizes, 93–341 GB | unsloth's published GGUF sizes |
| All tokens/s figures, RAM minimums, 46 / 105 / 58 GB per-card capacities | Kentino estimate — modelled, not measured |
| Usable RAM bandwidth per platform (35 / 60 / 110 / 300 GB/s) | Kentino estimate — typical real-world values |
| 64 GB unlock, 931–970 GB/s, Gen2 ×16 link state, 244–245 W, no ECC | Measured by Kentino on our own cards |
Sources: zai-org/GLM-5.3-Flash · unsloth/GLM-5.3-Flash-GGUF
commissioned here. Every CMP 170HX we sell is memory-tested at full 64 GB capacity and ships with a report carrying its own serial number. If you want this build sized against your actual workload, write to us.*