NVIDIA GeForce RTX 4090 24 GB GDDR6X (Refurbished)
Enterprise-grade platform, assembled and tested in the EU.
NVIDIA GeForce RTX 4090 — 24 GB for on-prem AI
A professionally refurbished RTX 4090: 24 GB of GDDR6X and Ada Lovelace compute, tested and verified stable under sustained load. The workhorse card for multi-GPU LLM inference and workstation AI at a fraction of datacenter-GPU cost.
The 24 GB workhorse for AI inference
The RTX 4090 pairs 24 GB of GDDR6X with 16,384 Ada Lovelace CUDA cores and 4th-gen Tensor cores — enough VRAM to serve quantized 7B–13B models comfortably, or to run a 70B-class model across two to four cards. It is the price-performance backbone of Kentino's multi-GPU AI builds. Each refurbished unit is cleaned, inspected and stress-tested before it ships.
Technical data
| GPU | NVIDIA AD102 — Ada Lovelace |
| Memory | 24 GB GDDR6X, 384-bit |
| Memory bandwidth | ~1,008 GB/s |
| CUDA cores | 16,384 |
| Tensor cores | 512 (4th gen) · up to ~330 TFLOPS FP16 |
| RT cores | 128 (3rd gen) |
| Interface | PCIe 4.0 ×16 |
| Board power | 450 W (16-pin 12VHPWR) |
| Condition | Refurbished — bench-tested, verified stable |
Where it fits
- On-prem LLM inference — 24 GB hosts quantized 7B–13B models; scale to 70B across multiple cards.
- AI workstations and fine-tuning of small-to-mid models.
- Rendering, simulation and content-creation workloads.
- Cost-effective multi-GPU servers where datacenter GPUs are overkill.
FAQ
What does "refurbished" mean here?
The card has been professionally cleaned, visually inspected and stress-tested for stability and full VRAM function before dispatch. It is not a new retail unit.
Can it run a 70B model?
Not on a single card — a 70B model at 4-bit needs ~40 GB of VRAM. Two to four RTX 4090 (48–96 GB total) handle it via tensor-parallel inference.
Does it fit a standard server?
The RTX 4090 is a large triple-slot card. It fits tower and 4U rack builds; confirm slot clearance and 16-pin power availability in your chassis.
The questions buyers ask us most often before ordering a server.
How long does it take?
Machines built from components we hold ship quickly; anything requiring a specific GPU generation depends on supply. We give you a date before you pay, and if it moves we tell you rather than letting you find out.
Can the configuration be changed before you build it?
Almost always. GPUs, memory, storage and cooling are chosen per order, and the listed configuration is a starting point rather than a fixed package. If you need more VRAM, faster storage or a different cooling approach, say so before you order and we will quote the change.
Can I collect the server in person?
You can. Our warehouse is in Prague, and collection in person is welcome — most people who come use the visit to go through the machine with our engineer and ask the questions that are awkward over email. For orders within the Czech Republic we also try to deliver personally and walk you through the setup on site.
Can I talk to someone who actually understands the workload?
Yes. We have an engineer who works on AI systems specifically, not a general sales desk. If your question is about batch sizes, quantisation, interconnect or where your bottleneck will be, ask it — that conversation usually changes the configuration for the better.
Which model can I run on this configuration?
Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.
How do you test a server before shipping?
We run it against actual AI workloads rather than synthetic scores: inference throughput, sustained load behaviour and thermals under continuous operation. You get the benchmark results with the machine, so the performance you were promised is the performance you can verify on day one.
Is the server ready to run when it arrives?
Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.
Ships from our EU warehouse. Heavy items may require freight arrangement — contact us for a shipping quote and lead time. 2-year limited warranty with advanced RMA support; extended warranty available.
Not exactly what you need?
Tell us your workload and we'll spec this platform around it — GPUs, memory, storage and cooling matched to what you actually run.