Building a 576 GB VRAM Inference Server: A Commissioning Log

Most server write-ups describe the machine that was ordered. This one describes the machine that was built — including the two things that were wrong with it when it first powered on, and the one we never fully explained.

Build: 7U rack · 2× AMD EPYC 7643 (96C/192T) · 512 GB DDR4 ECC · 6× NVIDIA RTX PRO 6000 Blackwell Max-Q, 576 GB GDDR7 · 2 TB + 4 TB NVMe
Built and commissioned: August 2026, Kentino workshop, Prague · Status: delivered

Kentino AI-7 7U chassis, sealed and powered on the bench
The finished machine on the commissioning bench, front panel removed for airflow testing.

We think that is the more useful article. Anybody can copy a parts list.

What it is

Six RTX PRO 6000 Blackwell Max-Q cards, 96 GB each, on a dual-socket EPYC Rome platform with 512 GB of ECC memory. 576 GB of VRAM in one box, which is enough to hold a 753-billion-parameter model at 4-bit with room left for a million tokens of context.

The Max-Q variant matters. At 300 W per card instead of 600, six of them fit inside a thermal and electrical envelope a workshop — and later a server room — can actually deliver. Full-load wall power measured ~2.9 kW, about 80% of a single 16 A/230 V circuit.

The chassis is our own 7U design: two independently cooled blocks of four GPU bays, so no card sits in the exhaust of more than three others.

Open 7U chassis from above showing the GPU bay dividers
Two independently cooled blocks of four bays. No card sits in the exhaust of more than three others.

The first fault: two GPUs quietly running at half speed

The system booted, all six cards enumerated, nvidia-smi looked perfect.

Two of them were running at PCIe Gen4 x8 instead of x16.

Nothing reported an error. The driver was happy, the cards worked, inference ran. The only way to see it was to measure host-to-device bandwidth per card and notice that two of them returned 13.9 GB/s where the others returned 27.7.

The cause was a physical-layer fault on two chassis connectors — not the cards, not the board, not the riser design. Both links were moved to spare connectors and all six then confirmed Gen4 x16 under load. The two suspect connectors are recorded by name in the build file and left unused for the life of the machine.

A GPU that trains at half its lane width does not announce itself. If your commissioning process does not measure per-card PCIe bandwidth under load, you will ship this fault and never know.

Chassis interior with GPU bays and cooling
Interior during assembly. Each card gets its own riser run to the board — which is exactly where the fault was.

The second fault: an ECC pattern that was not the memory

Partway through commissioning, two DIMMs began logging correctable ECC errors. Correctable means the hardware detected and fixed them: no data loss, no wrong answers, no impact on any result. But two modules reporting errors is normally two modules to replace.

We did not replace them, and the reason is the interesting part.

The events were spaced 4 h 30 m and 4 h 01 m apart, and each one asserted on both modules within the same second. Two independently marginal memory modules do not fail on a near-fixed interval, and they certainly do not fail in the same second three times running. That is not what degrading silicon looks like. That is the signature of something external and periodic.

The interval matched a high-power switching load in the commissioning environment. We then ran ~20 hours of instrumented soak under sustained real load, with per-minute telemetry on all six GPUs, CPU, memory, disk and the BMC event log. Zero ECC events.

So the modules stayed.

What we are careful to say: we never correlated the events against a logged cycle of the suspected source. The attribution is inference, not proof. We reported the observation, the interval and the test window to the customer, and said plainly that it is not a proven root cause. If the cause was environmental it does not travel with the machine — and twenty hours of silence under load is good evidence, not a guarantee.

One practical note that cost us time: this platform uses firmware-first error handling, so Linux-side ECC counters read zero no matter what is happening. Memory health has to be read from the BMC event log. If you monitor edac and nothing else, you are monitoring nothing.

What the machine actually does

All figures measured on this specific system, not from a datasheet.

Compute and consistency

Measurement Result
FP16 dense matmul, per card 291–302 TFLOPS (3.5% spread across six)
VRAM bandwidth, per card ~1,471 GB/s (uniform to 0.2%)
Six-way data-parallel inference 3,859 tok/s · 90.8% scaling efficiency
Per-replica spread across six GPUs 0.45%

That half-percent spread across six cards is the number we care about most. It means no card is quietly underperforming — and it is only meaningful because the PCIe fault above was found first.

Large-model inference

GLM-5.2 — 753 billion parameters (40 B active), 4-bit, 341 GB resident across six cards:

Concurrent requests 1 4 16 32 48
Generation tok/s 38.0 78.5 117–120 148.1 172.2

It also loads the model's native 1,048,576-token context on this hardware — 494 GiB of VRAM, leaving 6.6 GiB free on the tightest card. That works, but it is a thin margin and we said so.

Tuning the prefill micro-batch (-b/-ub 4096) lifted prompt processing from 403 to 606 tok/s, a 50% gain for a one-line change. Speculative decoding added +53% single-stream but cost 6% on long-prompt batch work, so it is enabled per workload rather than globally.

An RTX PRO 6000 Blackwell Max-Q card installed
One of the six RTX PRO 6000 Blackwell Max-Q cards. 96 GB each, 300 W each.

Burn-in

810,033 matmuls across six GPUs at 100% load. Zero computation errors. Peak 88 °C, sustained 300 W per card, no thermal throttling — clocks_event_reasons read Not Active throughout.

Where this machine is tight, and we said so

Thermal margin is ~2 °C. Peak 88 °C against a ~90 °C throttle point, at an ambient of 28–33 °C in an uncontrolled workshop in August. In a room below 25 °C there is comfortable room. In a warm one there is not. Sustained real workloads reached the same temperatures as synthetic burn-in, so this is not a synthetic-only concern.

The obvious lever is a GPU power cap — 260 W instead of 300 W. We did not measure what that costs in throughput on this machine, so we did not recommend it as though we had. Benchmark it against your own workload.

The software stack is specific and fragile. GLM-5.2 will not run under vLLM or SGLang on this GPU generation — three separate upstream limitations. llama.cpp is the working path. Separately, the vLLM installation carries a local patch fixing a crash under concurrency with quantised models, and that patch does not survive pip install -U vllm. Pin your versions.

Dual-socket means NUMA. Host-to-device bandwidth splits by socket: GPUs 0–2 on one, 3–5 on the other, with roughly a 20% penalty for host buffers allocated on the wrong node. Not a fault — physics — but you have to pin with numactl where host↔device bandwidth matters.

One test is genuinely incomplete. The final combined CPU+GPU maximum-load stage was interrupted by a failure of workshop power distribution equipment. The ~20 hours of mixed-load endurance testing covers the real operating case, so we reported the gap rather than quietly re-labelling the endurance runs as the missing test.

The finished server, sealed and ready for dispatch
Sealed and ready for dispatch.

Why we publish the faults

A commissioning report that says everything passed tells you nothing about the people who wrote it. Every machine of this complexity has something wrong with it on first power-on. The question is whether it gets found on our bench or in your rack.

On this build it was two connectors degrading a quarter of the GPU fabric to half speed, invisible to every status indicator on the system. It took measurement to find and ten minutes to fix.

That is what commissioning is for.

Retour au blog