NVIDIA

NVIDIA RTX PRO 6000 Blackwell Max-Q 96 GB GDDR7 ECC

NVIDIA RTX PRO 6000 Blackwell Max-Q 96 GB GDDR7 ECC

常规价格 €13.565,40 EUR
常规价格 促销价 €13.565,40 EUR
促销 售罄
结账时计算的运费
Blackwell · 96 GB ECC · 300 W

NVIDIA RTX PRO 6000 Blackwell Max-Q — 96 GB for on-prem AI

The full 96 GB Blackwell GPU in a 300 W envelope. Same unified ECC VRAM pool and same compute silicon as the 600 W card, drawing half the power — the card to choose when thermals, noise or a shared workstation chassis set the limit.

96 GB
GDDR7 ECC VRAM
24,064
CUDA cores
~1.8 TB/s
Memory bandwidth
300 W
Board power
Overview

96 GB unified VRAM at 300 W

The Max-Q variant carries the same Blackwell die and the same 96 GB of ECC GDDR7 as the full-power RTX PRO 6000, but is power-limited to 300 W. That matters more than the headline suggests: a single card holds a 70B-class model at 4-bit entirely in one unified memory pool, so there is no tensor-parallel split, no inter-GPU PCIe hop, and no partitioning work. Compared with stacking consumer cards to reach the same VRAM, you get ECC memory, a fraction of the power draw, and one slot instead of several.

Specifications

Technical data

GPU NVIDIA GB202 — Blackwell
Part number 900-5G153
Memory 96 GB GDDR7 ECC, 512-bit
Memory bandwidth ~1.8 TB/s
CUDA cores 24,064
Tensor cores 752 (5th gen) · ~2,000 TOPS INT8
RT cores 188 (4th gen)
Interface PCIe 5.0 ×16
Board power 300 W (16-pin 12V-2×6)
Cooling Active blower, dual-slot
Display outputs 4× DisplayPort 2.1b
Warranty 36 months
Best for

Where it fits

  • Single-card 70B inference — 96 GB holds a 70B model at 4-bit in one unified pool, no model splitting.
  • Quiet workstations and offices where a 600 W card is too much heat and noise.
  • Multi-GPU builds on a constrained power budget — four Max-Q draw about the same as two full-power cards.
  • Fine-tuning mid-size models where ECC memory and long-run stability matter.
  • Image and video generation with large models held resident in VRAM.
Indicative single-stream throughput: roughly 25–35 tok/s on a 70B model at 4-bit, and well over 100 tok/s on an 8B model. Figures are indicative estimates from published external references, not measured on Kentino hardware — tell us your model and context length and we will size it properly.
Questions

FAQ

How does Max-Q differ from the 600 W card?

Same GPU, same 96 GB of ECC GDDR7, same memory bandwidth. The Max-Q is capped at 300 W instead of 600 W, so sustained throughput is lower on compute-bound work — but any model that fits in 96 GB still fits, and VRAM capacity is what usually decides whether a model runs at all.

Can one card really run a 70B model?

Yes. A 70B model at 4-bit needs roughly 40 GB, so it fits in 96 GB with substantial room left for KV cache and long context. That is the main reason to choose this card over several smaller ones.

Do I need ECC memory?

For interactive work, no. For unattended training runs, long batch jobs or anything where a silent bit-flip would corrupt a result, ECC is the difference between a reliable machine and an unexplained failure.

Will it fit my chassis?

It is a dual-slot active-cooled card, so it fits standard tower and 4U builds. Confirm you have a free PCIe 5.0 ×16 slot and a 16-pin 12V-2×6 power connector available.

What is the lead time?

These are sourced to order — typically 10–21 days. Availability on this part moves quickly, so ask us to confirm before you plan around a date.

Want this card in a finished machine? Kentino builds, benchmarks and commissions complete RTX PRO 6000 AI servers and workstations — ask us for a configured build.
查看完整详细信息