Inference 70B L40 Serwer AI
Inference 70B L40 Serwer AI
Inference 70B L40 Serwer AI
Inference 70B L40 Serwer AI
Inference 70B L40 Serwer AI
Kentino · Custom AI Server

Inference 70B L40 Serwer AI

Enterprise-grade platform, assembled and tested in the EU.

€44.214,00ex-VAT · shipping calculated at checkout
In stock — ships from EU warehouse
ISO 9001 · verified hardwareEU warehouse & warranty7+ years in enterprise AIBurn-in tested before shipping
Overview
FAQ
Shipping & warranty

This product listing is kept for reference only.

This server has been replaced by the new Kentino AI product line. For the current equivalent or an upgraded configuration, please see our AI Servers collection.

Recommended replacement: Kentino AI 192 Rome L40 1448TOPS (4x NVIDIA L40, same platform, updated build)

70B L40 Computer

Specifications

  • GPU: 6x NVIDIA L40 (288 GB VRAM total)
  • Motherboard: ASRock Rack ROMED8-2T
  • CPU: AMD EPYC 7542
  • RAM: 512GB SK Hynix 2666MHz REG ECC DDR4 LRDIMM (8 x 64GB)
  • GPU-Motherboard Connection: RYSER PCIe 4.0 x16 Cable
  • Power Supply: 2x AX1600i 1000W
  • Case: 4U Rack Mount
  • Storage:
    • 2TB NVMe SSD
    • 500GB SATA Drive

Key Features

  1. High-Performance GPU Compute: Equipped with 6 NVIDIA L40 GPUs, providing a total of 288 GB VRAM for demanding AI, machine learning, and visualization workloads.
  2. Server-Grade Components: Features the reliable ASRock Rack ROMED8-2T motherboard and a powerful AMD EPYC 7542 CPU for exceptional processing capabilities.
  3. Ample Memory: 512GB of high-speed SK Hynix DDR4 RAM ensures smooth multitasking and efficient data processing for complex computations.
  4. High-Speed GPU Integration: Utilizes the RYSER PCIe 4.0 x16 cable for fast, full-bandwidth connection between the GPUs and the motherboard, ensuring optimal performance and data transfer speeds.
  5. Robust Power Supply: Dual AX1600i 1000W units provide stable and ample power delivery to support the high-performance components under heavy loads.
  6. Expandable Storage: Comes with a fast 2TB NVMe SSD for primary storage and an additional 500GB SATA drive for extra capacity.
  7. Professional-Grade Cooling: Housed in a spacious 24U rack mount case, providing optimal airflow and thermal management for sustained high-performance operation.
  8. Versatile Configuration: Designed for a wide range of high-performance computing tasks, from AI and machine learning to professional visualization and rendering.

Ideal Use Cases

  • Large Language Model Inference (e.g., 70B parameter models)
  • AI and Machine Learning Research
  • Data Analytics and Visualization
  • Professional 3D Rendering and Animation
  • Scientific Simulations
  • High-Performance Computing (HPC) Applications
  • Computer Vision and Image Processing
  • Financial Modeling and Risk Analysis

Special Notes

  • Optimized for 70B Models: With 288 GB of total GPU VRAM, this system is specifically designed to handle large language models with up to 70 billion parameters, making it ideal for cutting-edge AI research and applications.
  • NVIDIA L40 Advantage: The L40 GPUs offer a balance of compute performance and memory, suitable for a wide range of AI, HPC, and professional visualization workloads.
  • PCIe 4.0 Performance: The RYSER PCIe 4.0 x16 cable ensures that each GPU can operate at full bandwidth, maximizing data throughput and minimizing latency.
  • Scalable Design: While optimized for 70B parameter models, this system can be easily scaled or clustered for even larger workloads or multi-user environments.

The 70B L40 Computer represents a powerful and versatile solution for organizations and researchers working with large AI models, particularly in the realm of natural language processing and generation. Its balanced configuration of NVIDIA L40 GPUs, AMD EPYC CPU, and high-speed memory makes it suitable for a wide range of high-performance computing tasks beyond AI, including scientific simulations, data analytics, and professional visualization.

Delivery 2 - 6 weeks 

The questions buyers ask us most often before ordering a server.

How long does it take?

Machines built from components we hold ship quickly; anything requiring a specific GPU generation depends on supply. We give you a date before you pay, and if it moves we tell you rather than letting you find out.

Can the configuration be changed before you build it?

Almost always. GPUs, memory, storage and cooling are chosen per order, and the listed configuration is a starting point rather than a fixed package. If you need more VRAM, faster storage or a different cooling approach, say so before you order and we will quote the change.

Can I collect the server in person?

You can. Our warehouse is in Prague, and collection in person is welcome — most people who come use the visit to go through the machine with our engineer and ask the questions that are awkward over email. For orders within the Czech Republic we also try to deliver personally and walk you through the setup on site.

Can I talk to someone who actually understands the workload?

Yes. We have an engineer who works on AI systems specifically, not a general sales desk. If your question is about batch sizes, quantisation, interconnect or where your bottleneck will be, ask it — that conversation usually changes the configuration for the better.

Which model can I run on this configuration?

Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.

How do you test a server before shipping?

We run it against actual AI workloads rather than synthetic scores: inference throughput, sustained load behaviour and thermals under continuous operation. You get the benchmark results with the machine, so the performance you were promised is the performance you can verify on day one.

Is the server ready to run when it arrives?

Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.

Ships from our EU warehouse. Heavy items may require freight arrangement — contact us for a shipping quote and lead time. 2-year limited warranty with advanced RMA support; extended warranty available.

Not exactly what you need?

Tell us your workload and we'll spec this platform around it — GPUs, memory, storage and cooling matched to what you actually run.