NVIDIA L4 24 GB GDDR6 — Low Profile Passive
Enterprise-grade platform, assembled and tested in the EU.
NVIDIA L4 24 GB — the 72 W inference card
A single-slot, low-profile, passively cooled datacenter GPU that draws just 72 W straight from the PCIe slot — no power cable, no extra airflow design. The densest way to add 24 GB of inference capacity to a server.
Density and efficiency, not peak throughput
The L4 is built around a different priority to the big accelerator cards: performance per watt and per slot. At 72 W it needs no supplementary power connector at all, and being single-slot low-profile it fits chassis that physically cannot accept a full-height card. That combination is what makes it the practical choice for packing many independent inference workloads into one server, or for adding capable AI acceleration to compact and edge systems where power and space are the binding constraints. It is also a strong video card — hardware encode and decode including AV1 — which makes it a common pick for transcoding alongside inference.
Technical data
| GPU | NVIDIA AD104 — Ada Lovelace |
| Part number | TCSL4PCIE-PB |
| Memory | 24 GB GDDR6 with ECC |
| Memory bandwidth | ~300 GB/s |
| CUDA cores | 7,680 |
| Tensor cores | 240 (4th gen) · ~242 TOPS INT8 |
| RT cores | 60 (3rd gen) |
| Interface | PCIe 4.0 ×16 |
| Board power | 72 W — drawn from the slot, no power connector |
| Form factor | Single-slot, low profile, full-length |
| Cooling | Passive — requires chassis airflow |
| Display outputs | None — compute only |
| Video engines | Hardware encode / decode incl. AV1 |
| Warranty | 36 months |
Where it fits
- Serving small and mid-size models — 24 GB comfortably holds a quantised 7B–13B model.
- Many-GPU inference servers: at 72 W a card, you can fit several without redesigning power or cooling.
- Low-profile and compact chassis that cannot take a full-height dual-slot card.
- Video transcoding pipelines, including AV1, alongside inference on the same card.
- Edge and on-premise deployments where total power draw is capped.
FAQ
Does it need a power cable?
No. The whole card runs inside the PCIe slot's 75 W budget, so there is no 8-pin or 12-pin connector to route. That is a large part of its appeal in dense builds — it removes PSU cabling as a constraint entirely.
Will it cool itself?
No. Like other datacenter cards it is passive and relies on chassis airflow. In a server with a proper front-to-back path this is ideal; in a quiet desktop case it will overheat. Ask us if you are unsure about your chassis.
What model sizes fit in 24 GB?
A quantised 7B–13B model fits with room for context. A 70B model does not — that needs roughly 40 GB at 4-bit, so look at a 96 GB card or a multi-GPU configuration instead.
How does it compare with the L40?
The L40 has 48 GB and far more compute, at around 300 W and a full-height dual-slot form factor. The L4 trades that throughput for 72 W, one slot and low profile. Choose the L4 when power and density decide the build, the L40 when you need the performance.
Can I put several in one server?
Yes, and this is where the L4 is at its best — several cards at 72 W each stay within budgets that would be impossible with high-power GPUs. Kentino builds multi-L4 inference servers if you would rather buy the finished machine.
What is the lead time?
Sourced to order, typically 10–21 days. Ask us to confirm current availability before planning around a date.
The questions buyers ask us most often before ordering a server.
How long does it take?
Machines built from components we hold ship quickly; anything requiring a specific GPU generation depends on supply. We give you a date before you pay, and if it moves we tell you rather than letting you find out.
Can the configuration be changed before you build it?
Almost always. GPUs, memory, storage and cooling are chosen per order, and the listed configuration is a starting point rather than a fixed package. If you need more VRAM, faster storage or a different cooling approach, say so before you order and we will quote the change.
Can I collect the server in person?
You can. Our warehouse is in Prague, and collection in person is welcome — most people who come use the visit to go through the machine with our engineer and ask the questions that are awkward over email. For orders within the Czech Republic we also try to deliver personally and walk you through the setup on site.
Can I talk to someone who actually understands the workload?
Yes. We have an engineer who works on AI systems specifically, not a general sales desk. If your question is about batch sizes, quantisation, interconnect or where your bottleneck will be, ask it — that conversation usually changes the configuration for the better.
Which model can I run on this configuration?
Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.
How do you test a server before shipping?
We run it against actual AI workloads rather than synthetic scores: inference throughput, sustained load behaviour and thermals under continuous operation. You get the benchmark results with the machine, so the performance you were promised is the performance you can verify on day one.
Is the server ready to run when it arrives?
Yes. Every machine is assembled, burn-in tested and benchmarked on real AI workloads before it ships, and it leaves us with an LLM already installed and running. You plug it in, connect it to your network and start work — the only decision left is which project it runs first.
Ships from our EU warehouse. Heavy items may require freight arrangement — contact us for a shipping quote and lead time. 2-year limited warranty with advanced RMA support; extended warranty available.
Not exactly what you need?
Tell us your workload and we'll spec this platform around it — GPUs, memory, storage and cooling matched to what you actually run.