NVIDIA

NVIDIA L4 24 GB GDDR6 — Low Profile Passive

NVIDIA L4 24 GB GDDR6 — Low Profile Passive

Ordinarie pris €2.657,42 EUR
Ordinarie pris Försäljningspris €2.657,42 EUR
Rea Slutsåld
Frakt beräknas i kassan.
Ada Lovelace · 24 GB · 72 W · Low Profile

NVIDIA L4 24 GB — the 72 W inference card

A single-slot, low-profile, passively cooled datacenter GPU that draws just 72 W straight from the PCIe slot — no power cable, no extra airflow design. The densest way to add 24 GB of inference capacity to a server.

24 GB
GDDR6 ECC VRAM
7,680
CUDA cores
72 W
Slot-powered
1 slot
Low profile, passive
Overview

Density and efficiency, not peak throughput

The L4 is built around a different priority to the big accelerator cards: performance per watt and per slot. At 72 W it needs no supplementary power connector at all, and being single-slot low-profile it fits chassis that physically cannot accept a full-height card. That combination is what makes it the practical choice for packing many independent inference workloads into one server, or for adding capable AI acceleration to compact and edge systems where power and space are the binding constraints. It is also a strong video card — hardware encode and decode including AV1 — which makes it a common pick for transcoding alongside inference.

Specifications

Technical data

GPU NVIDIA AD104 — Ada Lovelace
Part number TCSL4PCIE-PB
Memory 24 GB GDDR6 with ECC
Memory bandwidth ~300 GB/s
CUDA cores 7,680
Tensor cores 240 (4th gen) · ~242 TOPS INT8
RT cores 60 (3rd gen)
Interface PCIe 4.0 ×16
Board power 72 W — drawn from the slot, no power connector
Form factor Single-slot, low profile, full-length
Cooling Passive — requires chassis airflow
Display outputs None — compute only
Video engines Hardware encode / decode incl. AV1
Warranty 36 months
Best for

Where it fits

  • Serving small and mid-size models — 24 GB comfortably holds a quantised 7B–13B model.
  • Many-GPU inference servers: at 72 W a card, you can fit several without redesigning power or cooling.
  • Low-profile and compact chassis that cannot take a full-height dual-slot card.
  • Video transcoding pipelines, including AV1, alongside inference on the same card.
  • Edge and on-premise deployments where total power draw is capped.
The L4 is a capacity-and-efficiency card, not a throughput champion. For heavy fine-tuning, large-model work or maximum tokens per second, an RTX PRO 6000 or a multi-GPU build is the better spend — tell us the workload and we will point you at the right one.
Questions

FAQ

Does it need a power cable?

No. The whole card runs inside the PCIe slot's 75 W budget, so there is no 8-pin or 12-pin connector to route. That is a large part of its appeal in dense builds — it removes PSU cabling as a constraint entirely.

Will it cool itself?

No. Like other datacenter cards it is passive and relies on chassis airflow. In a server with a proper front-to-back path this is ideal; in a quiet desktop case it will overheat. Ask us if you are unsure about your chassis.

What model sizes fit in 24 GB?

A quantised 7B–13B model fits with room for context. A 70B model does not — that needs roughly 40 GB at 4-bit, so look at a 96 GB card or a multi-GPU configuration instead.

How does it compare with the L40?

The L40 has 48 GB and far more compute, at around 300 W and a full-height dual-slot form factor. The L4 trades that throughput for 72 W, one slot and low profile. Choose the L4 when power and density decide the build, the L40 when you need the performance.

Can I put several in one server?

Yes, and this is where the L4 is at its best — several cards at 72 W each stay within budgets that would be impossible with high-power GPUs. Kentino builds multi-L4 inference servers if you would rather buy the finished machine.

What is the lead time?

Sourced to order, typically 10–21 days. Ask us to confirm current availability before planning around a date.

Want it in a finished machine? Kentino builds, benchmarks and commissions multi-L4 inference servers — ask us for a configured build.
Visa alla uppgifter