NVIDIA RTX PRO 6000 Blackwell Max-Q 96 GB GDDR7 ECC
Enterprise-grade platform, assembled and tested in the EU.
NVIDIA RTX PRO 6000 Blackwell Max-Q — 96 GB for on-prem AI
The full 96 GB Blackwell GPU in a 300 W envelope. Same unified ECC VRAM pool and same compute silicon as the 600 W card, drawing half the power — the card to choose when thermals, noise or a shared workstation chassis set the limit.
96 GB unified VRAM at 300 W
The Max-Q variant carries the same Blackwell die and the same 96 GB of ECC GDDR7 as the full-power RTX PRO 6000, but is power-limited to 300 W. That matters more than the headline suggests: a single card holds a 70B-class model at 4-bit entirely in one unified memory pool, so there is no tensor-parallel split, no inter-GPU PCIe hop, and no partitioning work. Compared with stacking consumer cards to reach the same VRAM, you get ECC memory, a fraction of the power draw, and one slot instead of several.
Technical data
| GPU | NVIDIA GB202 — Blackwell |
| Part number | 900-5G153 |
| Memory | 96 GB GDDR7 ECC, 512-bit |
| Memory bandwidth | ~1.8 TB/s |
| CUDA cores | 24,064 |
| Tensor cores | 752 (5th gen) · ~2,000 TOPS INT8 |
| RT cores | 188 (4th gen) |
| Interface | PCIe 5.0 ×16 |
| Board power | 300 W (16-pin 12V-2×6) |
| Cooling | Active blower, dual-slot |
| Display outputs | 4× DisplayPort 2.1b |
| Warranty | 36 months |
Where it fits
- Single-card 70B inference — 96 GB holds a 70B model at 4-bit in one unified pool, no model splitting.
- Quiet workstations and offices where a 600 W card is too much heat and noise.
- Multi-GPU builds on a constrained power budget — four Max-Q draw about the same as two full-power cards.
- Fine-tuning mid-size models where ECC memory and long-run stability matter.
- Image and video generation with large models held resident in VRAM.
FAQ
How does Max-Q differ from the 600 W card?
Same GPU, same 96 GB of ECC GDDR7, same memory bandwidth. The Max-Q is capped at 300 W instead of 600 W, so sustained throughput is lower on compute-bound work — but any model that fits in 96 GB still fits, and VRAM capacity is what usually decides whether a model runs at all.
Can one card really run a 70B model?
Yes. A 70B model at 4-bit needs roughly 40 GB, so it fits in 96 GB with substantial room left for KV cache and long context. That is the main reason to choose this card over several smaller ones.
Do I need ECC memory?
For interactive work, no. For unattended training runs, long batch jobs or anything where a silent bit-flip would corrupt a result, ECC is the difference between a reliable machine and an unexplained failure.
Will it fit my chassis?
It is a dual-slot active-cooled card, so it fits standard tower and 4U builds. Confirm you have a free PCIe 5.0 ×16 slot and a 16-pin 12V-2×6 power connector available.
What is the lead time?
These are sourced to order — typically 10–21 days. Availability on this part moves quickly, so ask us to confirm before you plan around a date.
The questions buyers ask us most often before ordering a server.
How long does it take?
Machines built from components we hold ship quickly; anything requiring a specific GPU generation depends on supply. We give you a date before you pay, and if it moves we tell you rather than letting you find out.
Can the configuration be changed before you build it?
Almost always. GPUs, memory, storage and cooling are chosen per order, and the listed configuration is a starting point rather than a fixed package. If you need more VRAM, faster storage or a different cooling approach, say so before you order and we will quote the change.
Can I collect the server in person?
You can. Our warehouse is in Prague, and collection in person is welcome — most people who come use the visit to go through the machine with our engineer and ask the questions that are awkward over email. For orders within the Czech Republic we also try to deliver personally and walk you through the setup on site.
Can I talk to someone who actually understands the workload?
Yes. We have an engineer who works on AI systems specifically, not a general sales desk. If your question is about batch sizes, quantisation, interconnect or where your bottleneck will be, ask it — that conversation usually changes the configuration for the better.
Which model can I run on this configuration?
Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.
How do you test a server before shipping?
We run it against actual AI workloads rather than synthetic scores: inference throughput, sustained load behaviour and thermals under continuous operation. You get the benchmark results with the machine, so the performance you were promised is the performance you can verify on day one.
Is the server ready to run when it arrives?
Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.
Ships from our EU warehouse. Heavy items may require freight arrangement — contact us for a shipping quote and lead time. 2-year limited warranty with advanced RMA support; extended warranty available.
Not exactly what you need?
Tell us your workload and we'll spec this platform around it — GPUs, memory, storage and cooling matched to what you actually run.