NVIDIA
NVIDIA L4 24 GB GDDR6 — Low Profile Passive
NVIDIA L4 24 GB GDDR6 — Low Profile Passive
Δεν ήταν δυνατή η φόρτωση της διαθεσιμότητας παραλαβής
NVIDIA L4 24 GB — the 72 W inference card
A single-slot, low-profile, passively cooled datacenter GPU that draws just 72 W straight from the PCIe slot — no power cable, no extra airflow design. The densest way to add 24 GB of inference capacity to a server.
Density and efficiency, not peak throughput
The L4 is built around a different priority to the big accelerator cards: performance per watt and per slot. At 72 W it needs no supplementary power connector at all, and being single-slot low-profile it fits chassis that physically cannot accept a full-height card. That combination is what makes it the practical choice for packing many independent inference workloads into one server, or for adding capable AI acceleration to compact and edge systems where power and space are the binding constraints. It is also a strong video card — hardware encode and decode including AV1 — which makes it a common pick for transcoding alongside inference.
Technical data
| GPU | NVIDIA AD104 — Ada Lovelace |
| Part number | TCSL4PCIE-PB |
| Memory | 24 GB GDDR6 with ECC |
| Memory bandwidth | ~300 GB/s |
| CUDA cores | 7,680 |
| Tensor cores | 240 (4th gen) · ~242 TOPS INT8 |
| RT cores | 60 (3rd gen) |
| Interface | PCIe 4.0 ×16 |
| Board power | 72 W — drawn from the slot, no power connector |
| Form factor | Single-slot, low profile, full-length |
| Cooling | Passive — requires chassis airflow |
| Display outputs | None — compute only |
| Video engines | Hardware encode / decode incl. AV1 |
| Warranty | 36 months |
Where it fits
- Serving small and mid-size models — 24 GB comfortably holds a quantised 7B–13B model.
- Many-GPU inference servers: at 72 W a card, you can fit several without redesigning power or cooling.
- Low-profile and compact chassis that cannot take a full-height dual-slot card.
- Video transcoding pipelines, including AV1, alongside inference on the same card.
- Edge and on-premise deployments where total power draw is capped.
FAQ
Does it need a power cable?
No. The whole card runs inside the PCIe slot's 75 W budget, so there is no 8-pin or 12-pin connector to route. That is a large part of its appeal in dense builds — it removes PSU cabling as a constraint entirely.
Will it cool itself?
No. Like other datacenter cards it is passive and relies on chassis airflow. In a server with a proper front-to-back path this is ideal; in a quiet desktop case it will overheat. Ask us if you are unsure about your chassis.
What model sizes fit in 24 GB?
A quantised 7B–13B model fits with room for context. A 70B model does not — that needs roughly 40 GB at 4-bit, so look at a 96 GB card or a multi-GPU configuration instead.
How does it compare with the L40?
The L40 has 48 GB and far more compute, at around 300 W and a full-height dual-slot form factor. The L4 trades that throughput for 72 W, one slot and low profile. Choose the L4 when power and density decide the build, the L40 when you need the performance.
Can I put several in one server?
Yes, and this is where the L4 is at its best — several cards at 72 W each stay within budgets that would be impossible with high-power GPUs. Kentino builds multi-L4 inference servers if you would rather buy the finished machine.
What is the lead time?
Sourced to order, typically 10–21 days. Ask us to confirm current availability before planning around a date.