Home AI Server — Punkstation Pro - 88 GB VRAM - EPYC
Home AI Server — Punkstation Pro - 88 GB VRAM - EPYC
Home AI Server — Punkstation Pro - 88 GB VRAM - EPYC
Home AI Server — Punkstation Pro - 88 GB VRAM - EPYC
Home AI Server — Punkstation Pro - 88 GB VRAM - EPYC
Kentino · Custom AI Server

Home AI Server — Punkstation Pro - 88 GB VRAM - EPYC

Enterprise-grade platform, assembled and tested in the EU.

€4.599,00ex-VAT · shipping calculated at checkout
In stock — ships from EU warehouse
ISO 9001 · verified hardwareEU warehouse & warranty7+ years in enterprise AIBurn-in tested before shipping
Overview
FAQ
Shipping & warranty

The Punkstation on a server-grade platform: a consumer RTX 4090 paired with a driver-unlocked NVIDIA CMP 170HX (64 GB HBM2e) — 88 GB of aggregated VRAM — on a single-socket AMD EPYC board with ECC memory, IPMI remote management and four PCIe 4.0 ×16 slots to grow into. Not the budget desktop — a robust, expandable machine, built and configured for maximum performance. Available as a quiet tower or a rack-mount server.

88 GBaggregated VRAM (24 + 64)
EPYCserver platform · ECC · IPMI
2 → 4GPU slots (room to expand)

What it's for

Same aggregated-VRAM strengths as the Punkstation, on hardware that lasts and scales. VRAM is aggregated, not pooled — the two cards are independent memories, so it's at its best on video and image generation, large mixture-of-experts models that fit one card, and mixture-of-agents / multi-model pipelines. The RTX 4090 drives the desktop and fast image generation; the 170HX's 64 GB carries the big model. The EPYC platform adds ECC memory, out-of-band IPMI management, more PCIe lanes and NVMe, and a clear upgrade path to the 4× 170HX build.

Image generationMixture-of-agentsVideo generationMoE reasoning

  • Validated: Qwen3 35B-A3B (mixture-of-experts, W8A8, compiled) runs entirely in the 170HX's 64 GB with no offload at ~140 tokens/s single-stream (and ~3,270 tok/s across a 48-request batch) — faster than a 7B dense model while reasoning coherently.
  • Video / image: generation models can't be split across cards — each fits one GPU. The 170HX's 64 GB hosts large diffusion and video models at high resolution and batch; the 4090 gives fast SDXL/Flux-class image generation. Run both in parallel for concurrent jobs.

Specification

Component Detail
Accelerator 1 NVIDIA GeForce RTX 4090 · 24 GB GDDR6X · drives display + fast single-card work
Accelerator 2 NVIDIA CMP 170HX · 64 GB HBM2e (modified, driver-unlocked, compute-only)
Total VRAM 88 GB aggregated (24 + 64), both fully accessible
CPU AMD EPYC (Rome / Milan, SP3) — model configured to order
Mainboard Single-socket EPYC server board · 8-channel DDR4 ECC · 4× PCIe 4.0 ×16 · onboard graphics + IPMI/BMC remote management
Memory DDR4 ECC, octa-channel — configurable (128 GB typical)
Storage M.2 NVMe (up to 3× PCIe 4.0 ×4)
Networking Dual 2.5 GbE + dedicated IPMI remote-management port
Power Uprated PSU sized to the build
Chassis Quiet desktop tower — or rack-mount server (choose above)
Operating system Ubuntu, NVIDIA drivers + CUDA pre-configured, compiled/cudagraph runtime tuned
Assembly Built, burned-in and validated at maximum performance before dispatch

Please read before ordering

⚠ Driver-unlocked card. The CMP 170HX is a compute card unlocked to full speed outside NVIDIA's specification. The configuration is confirmed working and ships validated, but the 170HX carries no manufacturer warranty — Kentino covers the build and the standard components; the 170HX is supported on a best-effort basis.
  • The 170HX is compute-only (no display) — the RTX 4090 (or onboard BMC graphics) drives the console.
  • The 170HX links at PCIe 2.0 (×4 standard, ×16 mod available) — limits host↔device transfer, not resident-model inference.

Ordering

Price in preparation — request a quote. Choose Tower or Server (rack) above. Lead time 10–28 days. Prices are EUR ex-VAT. Bespoke CPU / memory / storage configured to order.

The questions buyers ask us most often before ordering a server.

How long does it take?

Machines built from components we hold ship quickly; anything requiring a specific GPU generation depends on supply. We give you a date before you pay, and if it moves we tell you rather than letting you find out.

Can the configuration be changed before you build it?

Almost always. GPUs, memory, storage and cooling are chosen per order, and the listed configuration is a starting point rather than a fixed package. If you need more VRAM, faster storage or a different cooling approach, say so before you order and we will quote the change.

Can I collect the server in person?

You can. Our warehouse is in Prague, and collection in person is welcome — most people who come use the visit to go through the machine with our engineer and ask the questions that are awkward over email. For orders within the Czech Republic we also try to deliver personally and walk you through the setup on site.

Can I talk to someone who actually understands the workload?

Yes. We have an engineer who works on AI systems specifically, not a general sales desk. If your question is about batch sizes, quantisation, interconnect or where your bottleneck will be, ask it — that conversation usually changes the configuration for the better.

Which model can I run on this configuration?

Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.

How do you test a server before shipping?

We run it against actual AI workloads rather than synthetic scores: inference throughput, sustained load behaviour and thermals under continuous operation. You get the benchmark results with the machine, so the performance you were promised is the performance you can verify on day one.

Is the server ready to run when it arrives?

Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.

Ships from our EU warehouse. Heavy items may require freight arrangement — contact us for a shipping quote and lead time. 2-year limited warranty with advanced RMA support; extended warranty available.

Not exactly what you need?

Tell us your workload and we'll spec this platform around it — GPUs, memory, storage and cooling matched to what you actually run.