We Built a 576 GB AI Server That Runs GLM 5.2 in Your Own Building
Share
Kentino's on-premise AI server packs six NVIDIA RTX PRO 6000 Blackwell Max-Q cards, dual EPYC CPUs and enough redundancy to survive a pulled power supply — so frontier-class models run on your floor, not someone else's cloud.
Everyone knows One Miners. Now meet Kentino.
If you've been around our world for a while, you know One Miners. Kentino is its sister company, and Kentino builds machines: AI servers, workstations, gaming PCs — more or less whatever the spec sheet in your head looks like.
This article is about one of those builds. It's a system we assembled for a customer, so we're not naming who it's going to. What we can do is open the lid and show you exactly what goes into an on-premise AI server when you stop compromising.
The pitch: Claude-class AI that never leaves the building
The whole point of this machine is simple. You get an assistant in the same class as Claude Opus or GPT — and it runs on your premises. Your office, your rack room, your home if that's where you work. Every prompt, every document, every line of code you paste into it stays inside your own network.
That's the deal on offer: no per-token bill that scales with how useful the thing becomes, no queue behind other tenants, and no third party holding your context window. You own the hardware, so you own the workload.
Inside: six professional GPUs, 96 GB each
Open the chassis and the first thing you see is the GPU wall — three cards on one side, three more on the other. Each one is an NVIDIA RTX PRO 6000 Blackwell Max-Q with 96 GB of VRAM. Six cards, roughly 576 GB of video memory in a single box.
This configuration is six, but the chassis isn't full. It takes up to eight cards, which leaves room to add GPUs later or to spend the slots on something else entirely, like fiber networking. Behind the GPU plate sit two AMD EPYC 7643 processors feeding them.
The unglamorous parts are the ones that matter
A machine like this lives or dies on the boring specs. Memory: 512 GB of DDR4 installed as 8x 64 GB LRDIMMs, expandable to 1.5 TB when the workload asks for it. Storage: a 2 TB system drive plus 4 TB of NVMe, because model weights have to come off disk fast.
Power is five 2000 W redundant supplies. Pull one out mid-run and the server keeps going — it will make its displeasure known, but it keeps going. Cooling is twelve 15,000 RPM high-pressure fans, hot-swappable, arranged in rows across the chassis. You can yank a dead fan while the machine is under load, drop a new one in, close it up and walk away. It screams at you for a few seconds. That's the whole maintenance procedure.
The motherboard carries dedicated remote management, dual 10 Gigabit LAN, onboard VGA for a terminal, USB 3.2, and both a primary and a secondary power button so you don't have to unrack anything to cycle it.
Running GLM 5.2 locally
The headline workload is GLM 5.2 at Q4 — about 341 GB of model weights, resident across the six GPUs. That is a genuinely large model, and it lives in your building. Other open models run here too; there's more than enough VRAM to keep several loaded or to swap between them.
A cold start is not instant, and we'd rather tell you that than have you discover it. Between BIOS enumerating this much hardware, Linux bringing up a very large device tree, and then streaming 341 GB from SSD into VRAM, you're waiting on the order of minutes before the first token. Once it's warm, there is no single throughput number, and anyone quoting you one is skipping the question. It depends on three things: how many requests are in flight, whether speculative decoding is on, and whether the model fits on one card.
For GLM 5.2 sharded across all six GPUs: pure single-concurrency decode without speculative decoding is roughly 25–40 tokens/sec — that is the arithmetic, and it is the figure on our public product page. Turn speculative decoding on, still one request at a time, and you are at 40–60. That is the regime in the video, which is why the on-camera run lands at about 42 tokens/sec. Add concurrency on top of speculative decoding and the same box does 60–120.
Those are all big-model numbers, with 341 GB spread over six cards. Drop to a model that fits inside a single 96 GB GPU and you are into the thousands of tokens per second aggregate, because you are running six independent instances instead of one sharded one. Same hardware, an order of magnitude apart — which is the whole reason a single headline number is meaningless.
Beyond chat and code, the same silicon is a serious image generation box, and it will do video work as well. The practical difference from a cloud endpoint isn't only speed — it's that nobody is in front of you in the queue.
Why Max-Q, and what that trade actually costs
One detail worth being honest about: these are the Max-Q variants of the RTX PRO 6000. Max-Q runs at a 300 W board power. The non-Max-Q card is a 600 W part that you can cap to around 400 W in the driver.
Max-Q trades a little on-chip cache for that much lower power draw. In a six-card chassis, that trade is the reason the thermal and power budget works at all. It's the kind of decision you make once, at build time, and then live with happily — as long as you knew you were making it.
Coming next — and how to get one
There were a couple of other things on the bench during this shoot that don't belong to this server. An RTX 4090, and an NVIDIA CMP 170HX — an Ethereum-era mining card, listed at €1,600 on our site. That card gets a proper review and proper testing in the next video, so we're not going to half-explain it here.
We're also putting together a real benchmark table across several open models. Numbers, methodology, no hand-waving. That's the next one too.
If you want a machine like this one — or a smaller version of it, or a workstation, or a gaming build — that's what Kentino does. Browse the AI Servers collection, or tell us what you need and we will spec it.