Kentino · Custom AI Server

HX Cooler Kit - 3 GPU - 120 mm PWM Fan Manifold for CMP 170HX / A100

Enterprise-grade platform, assembled and tested in the EU.

€70,00 ex-VAT · shipping calculated at checkout
In stock — ships from EU warehouse
ISO 9001 · verified hardwareEU warehouse & warranty7+ years in enterprise AIBurn-in tested before shipping
Overview
FAQ
Shipping & warranty
3 GPU · 120 mm PWM · 3 EPS converters

HX Cooler Kit — 3 GPU

A three-wide plenum for a triple-accelerator desktop build. One 120 mm PWM fan, one duct spanning all three heatsinks, and three PCIe-to-EPS converters — the whole cooling and power interface for 192 GB of aggregated HBM in a tower.

3 GPUs
Cards cooled
120 mm
PWM fan
6,000 RPM
Peak fan speed
3 × EPS
Power converters
The problem

A datacenter GPU has no fan of its own

Every passive accelerator — the CMP 170HX, the A100, the whole Tesla line before them — ships as a bare finned heatsink with no fan attached. In a rack that is deliberate: a 2U chassis puts a wall of 15,000 RPM fans in front of the cards and forces air down their length. Drop the same card into a desktop tower and there is no such wall. Case airflow drifts past the heatsink instead of through it, the card heat-soaks within minutes, and you watch it clock down to nothing — or trip thermal shutdown.

This kit supplies the missing half of the design. A 120 mm fan is sealed into a manifold that mates to the intake face of the heatsink, so every litre the fan moves is forced through the fins and out the back of the card rather than around it. It is the same principle as the server shroud, sized for three cards side by side on a desktop board.

Without ducting, a 250 W passive card in a tower is not a slow card — it is a card that works for four minutes and then stops being useful. Ducted, it holds full clocks indefinitely.
In the box

What you get

  • The manifold — a sealed plenum sized for three cards side by side, mating to the heatsink intake face.
  • 1 × 120 mm PWM fan — 1,000–6,000 RPM, 2.7 A, 4-pin PWM control.
  • 3 × PCIe-to-EPS converters — one per card, so nothing else is needed to power the GPUs.
  • Mounting hardware for the fan-to-manifold and manifold-to-card interfaces.
Three cards, one fan

Running the fan higher on a wide plenum

Three heatsinks in one plenum is the point at which fan speed starts to matter. The wider the duct, the more the fan is working against combined fin resistance rather than open air — which is precisely why this kit uses a fan that reaches 6,000 RPM instead of a quiet 1,800 RPM case fan. Expect to sit higher on the curve than a single-card build, roughly the 2,500–3,500 RPM band under sustained three-card load.

If the noise at that point is more than you want in the room, the honest answer is that three 250 W passive cards belong in a rack. Ask us about the rackmount route — we build both.
Power

Why the card needs a converter, not just a cable

Datacenter cards take power through an 8-pin EPS inlet at the end of the card — the CPU-style connector, not the graphics-card one. It is the single most common way people destroy one of these on the first boot: an EPS 8-pin and a PCIe 8-pin look alike, and in the wrong hands one will go into the other.

They are not the same. A PCIe 8-pin carries three 12 V lines and five grounds; an EPS 8-pin carries four and four, in a different arrangement. Cross them and you feed 12 V into a ground pin. The converters in this kit remap the pins correctly and take their feed from your PSU's PCIe cables, which every desktop supply has spare.

Never adapt an EPS inlet with a plain re-pinned cable or a straight-through "8-pin to 8-pin" lead. Use a converter that is explicitly built for the direction you need. That is what ships here.
Noise and control

6,000 RPM is the ceiling, not the setting

The fan in this kit is an industrial-grade unit with a very wide range: it will idle near 1,000 RPM and it will pull 2.7 A flat out at 6,000. At the top of that range it is genuinely loud — this is server-class air movement and we are not going to pretend otherwise.

The point of the 4-pin PWM control is that you decide where on that range you live. Put the fan on a curve tied to GPU temperature: it will idle quietly when the cards are idle and only climb when you are actually loading all three. The headroom is there for the days you saturate every card at once, or for a warm summer with no air conditioning.

Check your fan header first. At full speed this fan draws up to 2.7 A — well beyond the ~1 A a typical motherboard chassis header is rated for. Feed the fan from the PSU (a hub with its own SATA or Molex input is the usual answer) and take only the PWM signal wire from the header. Ask us if you want the wiring spelled out for your board.
Fitment

Which cards this fits

The manifold is cut for the standard NVIDIA full-height, full-length, dual-slot passive body — the 267 mm card with a finned tunnel running its whole length and the 8-pin EPS inlet at the far end. That covers CMP 170HX, A100 PCIe, A30, A40, L40, V100, P100 and P40, and the AMD Instinct cards of the same generation share the geometry closely enough to work.

Three dual-slot cards side by side needs a board with six usable slot positions and the x16 pitch to match. Not every desktop board has it. Tell us the board and we will confirm the layout works before you order.
Specifications

Technical data

Cards supported 3 × full-height, full-length dual-slot passive accelerators
Fan 1 × 120 mm, 1,000–6,000 RPM
Fan current Up to 2.7 A at 12 V, per fan
Fan control 4-pin PWM
Power converters included 3 × PCIe to 8-pin EPS
Card power inlet 8-pin EPS at the end of the card
Verified cards CMP 170HX, A100 PCIe, A30, A40, L40, V100, P100 and P40
Price €70 ex VAT
Questions

FAQ

Do I need this, or will good case airflow do?

You need it. Case airflow moves air through the case; a passive accelerator needs air forced through a narrow finned tunnel against real static pressure. Those are different jobs. Builders who skip the duct report the card idling fine and then collapsing under the first sustained load — which is exactly the load you bought it for.

Is one fan really enough for three cards?

It is, because the plenum is sealed and the fan is rated far beyond a normal case fan — 2.7 A and 6,000 RPM exists precisely to push against three heatsinks of fin resistance. You will run it higher on the curve than a one- or two-card build, around 2,500–3,500 RPM under sustained load.

Is it loud?

Under full three-card load, yes — noticeably. This is the configuration where we would genuinely ask whether the machine wants to live in the room with you or in a rack somewhere else.

Can I run the fan from a motherboard header?

Take the PWM signal from the header, but not the power. At 2.7 A the fan exceeds what a standard chassis header is rated to deliver. Feed it from the PSU through a powered hub.

What PSU do I need for three cards?

Budget around 250 W per accelerator on top of the rest of the system — so roughly 750 W of GPU load before the CPU, drives and fans. A 1,200–1,600 W supply with enough spare PCIe outputs to feed three converters is the sensible bracket.

Do you have a four-card version?

Yes — a dual-fan, full-board manifold built specifically for the Huananzhi H12D-8D, which also ducts air over the chipset and the board's three M.2 drives. Ask us about it.

Buying the cards too? We sell the CMP 170HX 64 GB and we build complete machines around these accelerators — see the AI Servers collection. Order the cooler alongside a card and it ships together.

The questions buyers ask us most often before ordering a server.

How long does it take?

Machines built from components we hold ship quickly; anything requiring a specific GPU generation depends on supply. We give you a date before you pay, and if it moves we tell you rather than letting you find out.

Can the configuration be changed before you build it?

Almost always. GPUs, memory, storage and cooling are chosen per order, and the listed configuration is a starting point rather than a fixed package. If you need more VRAM, faster storage or a different cooling approach, say so before you order and we will quote the change.

Can I collect the server in person?

You can. Our warehouse is in Prague, and collection in person is welcome — most people who come use the visit to go through the machine with our engineer and ask the questions that are awkward over email. For orders within the Czech Republic we also try to deliver personally and walk you through the setup on site.

Can I talk to someone who actually understands the workload?

Yes. We have an engineer who works on AI systems specifically, not a general sales desk. If your question is about batch sizes, quantisation, interconnect or where your bottleneck will be, ask it — that conversation usually changes the configuration for the better.

Which model can I run on this configuration?

Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.

How do you test a server before shipping?

We run it against actual AI workloads rather than synthetic scores: inference throughput, sustained load behaviour and thermals under continuous operation. You get the benchmark results with the machine, so the performance you were promised is the performance you can verify on day one.

Is the server ready to run when it arrives?

Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.

Ships from our EU warehouse. Heavy items may require freight arrangement — contact us for a shipping quote and lead time. 2-year limited warranty with advanced RMA support; extended warranty available.

Not exactly what you need?

Tell us your workload and we'll spec this platform around it — GPUs, memory, storage and cooling matched to what you actually run.