NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB GDDR7 ECC
Enterprise-grade platform, assembled and tested in the EU.
NVIDIA RTX PRO 6000 Blackwell Server Edition — 96 GB
The datacenter variant: 96 GB of unified ECC GDDR7, passively cooled for front-to-back rack airflow, with a configurable 400–600 W power envelope. Built to run continuously in a server chassis rather than on a desk.
Built for 24/7 rack operation
The Server Edition carries the same Blackwell die and the same 96 GB ECC memory pool as the workstation cards, but is engineered for a different environment. It has no fan and no display outputs: cooling comes from the chassis, in a front-to-back path, which is what makes dense multi-GPU rack builds possible in the first place. The power envelope is configurable between 400 W and 600 W, so density and thermal budget can be traded against per-card throughput. This is the card behind Kentino's multi-GPU Kentino AI rack servers.
Technical data
| GPU | NVIDIA GB202 — Blackwell |
| Memory | 96 GB GDDR7 ECC, 512-bit |
| Memory bandwidth | ~1.8 TB/s |
| CUDA cores | 24,064 |
| Tensor cores | 752 (5th gen) · ~2,000 TOPS INT8 |
| RT cores | 188 (4th gen) |
| Interface | PCIe 5.0 ×16 |
| Board power | 400–600 W configurable |
| Cooling | Passive — requires chassis front-to-back airflow |
| Display outputs | None |
| Form factor | Dual-slot, full-height |
| Warranty | 3-year manufacturer warranty |
Where it fits
- Dense multi-GPU rack servers — 2, 4, 6 or 8 cards in a single chassis.
- Continuous 24/7 inference serving in a datacenter or server room.
- Shared or multi-tenant AI infrastructure where ECC and stability are requirements.
- Deployments where per-card power must be capped to fit a rack thermal budget.
- Large-model serving across several cards with 96 GB of unified VRAM each.
FAQ
Can I use this in a desktop PC?
No. It has no fan and relies entirely on chassis airflow designed for passive GPUs. In a normal tower case it will throttle and then overheat. Use the Workstation or Max-Q edition for tower and desk builds.
How does it compare with the workstation cards?
Identical GPU and identical 96 GB of ECC memory. The differences are physical: passive cooling, no display outputs, and a configurable 400–600 W envelope for rack density. Choose it only if you have a server chassis to put it in.
Why cap it at 400 W?
In a dense build, total rack power and cooling — not the individual card — is usually the limit. Capping each card lets you fit more GPUs into the same thermal budget, which typically yields more total throughput than fewer cards running flat out.
Do you supply complete servers with these?
Yes. The Kentino AI rack line is built around this card in 1, 2, 4, 6 and 8-GPU configurations, assembled, benchmarked and commissioned before delivery.
The questions buyers ask us most often before ordering a server.
How long does it take?
Machines built from components we hold ship quickly; anything requiring a specific GPU generation depends on supply. We give you a date before you pay, and if it moves we tell you rather than letting you find out.
Can the configuration be changed before you build it?
Almost always. GPUs, memory, storage and cooling are chosen per order, and the listed configuration is a starting point rather than a fixed package. If you need more VRAM, faster storage or a different cooling approach, say so before you order and we will quote the change.
Can I collect the server in person?
You can. Our warehouse is in Prague, and collection in person is welcome — most people who come use the visit to go through the machine with our engineer and ask the questions that are awkward over email. For orders within the Czech Republic we also try to deliver personally and walk you through the setup on site.
Can I talk to someone who actually understands the workload?
Yes. We have an engineer who works on AI systems specifically, not a general sales desk. If your question is about batch sizes, quantisation, interconnect or where your bottleneck will be, ask it — that conversation usually changes the configuration for the better.
Which model can I run on this configuration?
Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.
How do you test a server before shipping?
We run it against actual AI workloads rather than synthetic scores: inference throughput, sustained load behaviour and thermals under continuous operation. You get the benchmark results with the machine, so the performance you were promised is the performance you can verify on day one.
Is the server ready to run when it arrives?
Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.
Ships from our EU warehouse. Heavy items may require freight arrangement — contact us for a shipping quote and lead time. 2-year limited warranty with advanced RMA support; extended warranty available.
Not exactly what you need?
Tell us your workload and we'll spec this platform around it — GPUs, memory, storage and cooling matched to what you actually run.