Design your AI server

Pick what you want to run. We draft a machine, show the models it can host on-prem, and price it live under Kentino's build conditions — form factor, GPUs, memory and power, all adjustable.

1What will it do?sets a starting build
2Adjust the buildconstrained to what we ship
Estimates are ex-VAT indications under Kentino's standard build conditions; final quote confirmed at order (lead time 10–28 days). LLM-class models pool VRAM across GPUs via tensor parallelism (scaling isn't linear; dense parallelises worse than MoE). Video, image and audio models cannot be split across GPUs — each must fit one card's VRAM; extra cards add parallel jobs, not bigger models. Redundant power is not silent (~60 dB). Bespoke builds (AMD Instinct, integrated UPS/NAS) are quoted separately.