INSTRUCT L40 12 GPU Server (Legacy)
Enterprise-grade platform, assembled and tested in the EU.
This product listing is kept for reference only.
This server has been replaced by the new Kentino AI product line. For the current equivalent or an upgraded configuration, please see our AI Servers collection.
Recommended replacement: Kentino AI 288 Rome L40 (6x L40, passive) or Kentino AI 576 Genoa RTXPro6000 (6x RTX Pro 6000, 576 GB ECC VRAM)
⚠️ LEGACY PRODUCT - DEPRECATED CONFIGURATION
This 12 GPU multi-case server configuration is no longer in active production due to reliability issues with multi-case interconnects. Listed for reference and historical purposes only. For current production AI builds, please see our 8 GPU systems with PCIe-based GPUs (RTX 5090, RTX Pro 6000 Blackwell, L40, L4, Intel Arc Pro B70, AMD configurations) or contact us for a custom configuration.
Specifications
- GPU: 12x NVIDIA L40 48GB VRAM (576GB total)
- Motherboard: ASRock Rack ROMED8-2T
- CPU: AMD EPYC 7713
- RAM: 1024GB CT128G4ZFJ426S
- Power Supply: 4x AX1600i 6400W
- Case: Non-standard U8 Rack Mount
-
Storage:
- 4TB NVMe SSD
- 500GB SATA Drive
Key Features
- High-Performance Computing: Equipped with 12 powerful NVIDIA L40 GPUs, providing an impressive 576GB of VRAM for intensive AI and machine learning tasks.
- Server-Grade Components: Features the reliable ASRock Rack ROMED8-2T motherboard and a top-tier AMD EPYC 7713 CPU for exceptional processing power.
- Massive Memory: 1024GB of high-speed RAM ensures smooth multitasking and efficient data processing for even the most demanding workloads.
- Robust Power Supply: Quad power supply setup with 4x AX1600i 1600W units, ensuring stable and ample power delivery under heavy loads.
- Expandable Storage: Comes with a fast 4TB NVMe SSD for primary storage and an additional 500GB SATA drive for extra capacity.
- Compact Design: Housed in a non-standard U8 rack mount case, optimizing space efficiency while maintaining performance.
- Scalable Configuration: Supports up to 4 nodes per rack, allowing for impressive computational density and scalability.
Ideal Use Cases
- Deep Learning and AI Model Training
- High-Performance Computing (HPC) Applications
- Large-Scale Data Analysis and Visualization
- Scientific Simulations and Research
- Rendering and 3D Modeling
Special Notes
- Non-Standard Case: This server utilizes a custom U8 case, providing a more compact form factor compared to traditional server cases. This design allows for improved space efficiency in data centers and server rooms.
- Multi-Node Configuration: The rack design supports up to 4 nodes per rack, enabling a highly dense and powerful computing cluster. This configuration is ideal for organizations requiring massive parallel processing capabilities or looking to maximize computational power per square foot.
This INSTRUCT12 GPU Server is the ultimate solution for organizations and researchers requiring extreme computational power in a space-efficient design. With its massive GPU capacity, server-grade components, and the ability to house up to 4 nodes per rack, it's built to tackle the most complex calculations and data processing tasks while optimizing data center space utilization.
Delivery 2 - 6 weeks
The questions buyers ask us most often before ordering a server.
How long does it take?
Machines built from components we hold ship quickly; anything requiring a specific GPU generation depends on supply. We give you a date before you pay, and if it moves we tell you rather than letting you find out.
Can the configuration be changed before you build it?
Almost always. GPUs, memory, storage and cooling are chosen per order, and the listed configuration is a starting point rather than a fixed package. If you need more VRAM, faster storage or a different cooling approach, say so before you order and we will quote the change.
Can I collect the server in person?
You can. Our warehouse is in Prague, and collection in person is welcome — most people who come use the visit to go through the machine with our engineer and ask the questions that are awkward over email. For orders within the Czech Republic we also try to deliver personally and walk you through the setup on site.
Can I talk to someone who actually understands the workload?
Yes. We have an engineer who works on AI systems specifically, not a general sales desk. If your question is about batch sizes, quantisation, interconnect or where your bottleneck will be, ask it — that conversation usually changes the configuration for the better.
Which model can I run on this configuration?
Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.
How do you test a server before shipping?
We run it against actual AI workloads rather than synthetic scores: inference throughput, sustained load behaviour and thermals under continuous operation. You get the benchmark results with the machine, so the performance you were promised is the performance you can verify on day one.
Is the server ready to run when it arrives?
Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.
Ships from our EU warehouse. Heavy items may require freight arrangement — contact us for a shipping quote and lead time. 2-year limited warranty with advanced RMA support; extended warranty available.
Not exactly what you need?
Tell us your workload and we'll spec this platform around it — GPUs, memory, storage and cooling matched to what you actually run.