{"product_id":"nvidia-l4-24-gb-gddr6-low-profile-passive","title":"NVIDIA L4 24 GB GDDR6 — Low Profile Passive","description":"\u003cstyle\u003e\n.rtx{--bl:#334fb4;--dk:#26408f;--ink:#17171a;--soft:#f3f7fd;--muted:#6c6660;--line:#dce6f5;font-family:-apple-system,BlinkMacSystemFont,\"Segoe UI\",Roboto,Helvetica,Arial,sans-serif;color:var(--ink);line-height:1.6;font-size:15px;}\n.rtx *{box-sizing:border-box;}\n.rtx-head{background:var(--ink);color:#fff;border-radius:14px;padding:30px 28px;position:relative;overflow:hidden;}\n.rtx-head:before{content:\"\";position:absolute;inset:0;background:radial-gradient(120% 100% at 100% 0,rgba(51,79,180,.32),transparent 55%);}\n.rtx-head\u003e*{position:relative;}\n.rtx-head h2{font-size:clamp(22px,3.2vw,30px);font-weight:800;margin:0 0 10px;color:#fff !important;}\n.rtx-head p{color:#d7d2cb;max-width:64ch;margin:0;}\n.rtx-badge{display:inline-block;background:var(--bl);color:#fff !important;font-size:12px;font-weight:700;letter-spacing:.08em;text-transform:uppercase;padding:5px 11px;border-radius:999px;margin-bottom:14px;}\n.rtx-hi{display:grid;grid-template-columns:repeat(4,1fr);gap:1px;background:var(--line);border:1px solid var(--line);border-radius:12px;overflow:hidden;margin:18px 0 6px;}\n.rtx-hi div{background:#fff;padding:18px 12px;text-align:center;}\n.rtx-hi .v{font-size:20px;font-weight:800;}\n.rtx-hi .v small{color:var(--bl);font-size:12px;}\n.rtx-hi .k{font-size:11.5px;color:var(--muted);margin-top:4px;}\n.rtx-sec{padding:24px 0 4px;}\n.rtx-ey{font-size:12px;font-weight:700;letter-spacing:.16em;text-transform:uppercase;color:var(--dk);}\n.rtx h2.t{font-size:clamp(19px,2.4vw,24px);font-weight:800;margin:6px 0 12px;}\n.rtx p.lead{font-size:16px;color:#3a3a36;}\n.rtx table{width:100%;border-collapse:collapse;font-size:14px;margin:6px 0;}\n.rtx table.spec td{padding:9px 12px;border-bottom:1px solid var(--line);}\n.rtx table.spec tr:nth-child(odd) td{background:#f6f9fc;}\n.rtx table.spec td:first-child{width:42%;color:var(--muted);font-weight:600;}\n.rtx ul.tick{list-style:none;margin:6px 0;padding:0;}\n.rtx ul.tick li{position:relative;padding:7px 0 7px 26px;border-bottom:1px solid var(--line);font-size:14.5px;}\n.rtx ul.tick li:last-child{border-bottom:none;}\n.rtx ul.tick li:before{content:\"\";position:absolute;left:4px;top:14px;width:9px;height:9px;border-radius:2px;background:var(--bl);}\n.rtx-note{background:var(--soft);border-left:3px solid var(--bl);padding:13px 16px;border-radius:0 8px 8px 0;font-size:13.5px;color:#4a4641;margin-top:14px;}\n.rtx-faq details{border-bottom:1px solid var(--line);padding:14px 0;}\n.rtx-faq summary{font-weight:700;cursor:pointer;list-style:none;display:flex;justify-content:space-between;font-size:15px;}\n.rtx-faq summary:after{content:\"+\";color:var(--bl);font-weight:700;}\n.rtx-faq details[open] summary:after{content:\"–\";}\n.rtx-faq p{margin:9px 0 0;color:var(--muted);font-size:14px;}\n@media(max-width:760px){.rtx-hi{grid-template-columns:repeat(2,1fr);}}\n\u003c\/style\u003e\n\u003cdiv class=\"rtx\"\u003e\n\u003cdiv class=\"rtx-head\"\u003e\n\u003cspan class=\"rtx-badge\"\u003eAda Lovelace · 24 GB · 72 W · Low Profile\u003c\/span\u003e\n\u003ch2\u003eNVIDIA L4 24 GB — the 72 W inference card\u003c\/h2\u003e\n\u003cp\u003eA single-slot, low-profile, passively cooled datacenter GPU that draws just 72 W straight from the PCIe slot — no power cable, no extra airflow design. The densest way to add 24 GB of inference capacity to a server.\u003c\/p\u003e\n\u003c\/div\u003e\n\u003cdiv class=\"rtx-hi\"\u003e\n\u003cdiv\u003e\n\u003cdiv class=\"v\"\u003e24 \u003csmall\u003eGB\u003c\/small\u003e\n\u003c\/div\u003e\n\u003cdiv class=\"k\"\u003eGDDR6 ECC VRAM\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv class=\"v\"\u003e7,680\u003c\/div\u003e\n\u003cdiv class=\"k\"\u003eCUDA cores\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv class=\"v\"\u003e72 \u003csmall\u003eW\u003c\/small\u003e\n\u003c\/div\u003e\n\u003cdiv class=\"k\"\u003eSlot-powered\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cdiv\u003e\n\u003cdiv class=\"v\"\u003e1 \u003csmall\u003eslot\u003c\/small\u003e\n\u003c\/div\u003e\n\u003cdiv class=\"k\"\u003eLow profile, passive\u003c\/div\u003e\n\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cdiv class=\"rtx-sec\"\u003e\n\u003cdiv class=\"rtx-ey\"\u003eOverview\u003c\/div\u003e\n\u003ch2 class=\"t\"\u003eDensity and efficiency, not peak throughput\u003c\/h2\u003e\n\u003cp class=\"lead\"\u003eThe L4 is built around a different priority to the big accelerator cards: performance per watt and per slot. At 72 W it needs no supplementary power connector at all, and being single-slot low-profile it fits chassis that physically cannot accept a full-height card. That combination is what makes it the practical choice for packing many independent inference workloads into one server, or for adding capable AI acceleration to compact and edge systems where power and space are the binding constraints. It is also a strong video card — hardware encode and decode including AV1 — which makes it a common pick for transcoding alongside inference.\u003c\/p\u003e\n\u003c\/div\u003e\n\u003cdiv class=\"rtx-sec\"\u003e\n\u003cdiv class=\"rtx-ey\"\u003eSpecifications\u003c\/div\u003e\n\u003ch2 class=\"t\"\u003eTechnical data\u003c\/h2\u003e\n\u003ctable class=\"spec\"\u003e\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd\u003eGPU\u003c\/td\u003e\n\u003ctd\u003eNVIDIA AD104 — Ada Lovelace\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003ePart number\u003c\/td\u003e\n\u003ctd\u003eTCSL4PCIE-PB\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eMemory\u003c\/td\u003e\n\u003ctd\u003e24 GB GDDR6 with ECC\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eMemory bandwidth\u003c\/td\u003e\n\u003ctd\u003e~300 GB\/s\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eCUDA cores\u003c\/td\u003e\n\u003ctd\u003e7,680\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eTensor cores\u003c\/td\u003e\n\u003ctd\u003e240 (4th gen) · ~242 TOPS INT8\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eRT cores\u003c\/td\u003e\n\u003ctd\u003e60 (3rd gen)\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eInterface\u003c\/td\u003e\n\u003ctd\u003ePCIe 4.0 ×16\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eBoard power\u003c\/td\u003e\n\u003ctd\u003e72 W — drawn from the slot, no power connector\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eForm factor\u003c\/td\u003e\n\u003ctd\u003eSingle-slot, low profile, full-length\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eCooling\u003c\/td\u003e\n\u003ctd\u003ePassive — requires chassis airflow\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eDisplay outputs\u003c\/td\u003e\n\u003ctd\u003eNone — compute only\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eVideo engines\u003c\/td\u003e\n\u003ctd\u003eHardware encode \/ decode incl. AV1\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003eWarranty\u003c\/td\u003e\n\u003ctd\u003e36 months\u003c\/td\u003e\n\u003c\/tr\u003e\n\u003c\/tbody\u003e\u003c\/table\u003e\n\u003c\/div\u003e\n\u003cdiv class=\"rtx-sec\"\u003e\n\u003cdiv class=\"rtx-ey\"\u003eBest for\u003c\/div\u003e\n\u003ch2 class=\"t\"\u003eWhere it fits\u003c\/h2\u003e\n\u003cul class=\"tick\"\u003e\n\u003cli\u003eServing small and mid-size models — 24 GB comfortably holds a quantised 7B–13B model.\u003c\/li\u003e\n\u003cli\u003eMany-GPU inference servers: at 72 W a card, you can fit several without redesigning power or cooling.\u003c\/li\u003e\n\u003cli\u003eLow-profile and compact chassis that cannot take a full-height dual-slot card.\u003c\/li\u003e\n\u003cli\u003eVideo transcoding pipelines, including AV1, alongside inference on the same card.\u003c\/li\u003e\n\u003cli\u003eEdge and on-premise deployments where total power draw is capped.\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003cdiv class=\"rtx-note\"\u003eThe L4 is a capacity-and-efficiency card, not a throughput champion. For heavy fine-tuning, large-model work or maximum tokens per second, an RTX PRO 6000 or a multi-GPU build is the better spend — tell us the workload and we will point you at the right one.\u003c\/div\u003e\n\u003c\/div\u003e\n\u003cdiv class=\"rtx-sec\"\u003e\n\u003cdiv class=\"rtx-ey\"\u003eQuestions\u003c\/div\u003e\n\u003ch2 class=\"t\"\u003eFAQ\u003c\/h2\u003e\n\u003cdiv class=\"rtx-faq\"\u003e\n\u003cdetails open\u003e\u003csummary\u003eDoes it need a power cable?\u003c\/summary\u003e\u003cp\u003eNo. The whole card runs inside the PCIe slot's 75 W budget, so there is no 8-pin or 12-pin connector to route. That is a large part of its appeal in dense builds — it removes PSU cabling as a constraint entirely.\u003c\/p\u003e\u003c\/details\u003e\n\u003cdetails\u003e\u003csummary\u003eWill it cool itself?\u003c\/summary\u003e\u003cp\u003eNo. Like other datacenter cards it is passive and relies on chassis airflow. In a server with a proper front-to-back path this is ideal; in a quiet desktop case it will overheat. Ask us if you are unsure about your chassis.\u003c\/p\u003e\u003c\/details\u003e\n\u003cdetails\u003e\u003csummary\u003eWhat model sizes fit in 24 GB?\u003c\/summary\u003e\u003cp\u003eA quantised 7B–13B model fits with room for context. A 70B model does not — that needs roughly 40 GB at 4-bit, so look at a 96 GB card or a multi-GPU configuration instead.\u003c\/p\u003e\u003c\/details\u003e\n\u003cdetails\u003e\u003csummary\u003eHow does it compare with the L40?\u003c\/summary\u003e\u003cp\u003eThe L40 has 48 GB and far more compute, at around 300 W and a full-height dual-slot form factor. The L4 trades that throughput for 72 W, one slot and low profile. Choose the L4 when power and density decide the build, the L40 when you need the performance.\u003c\/p\u003e\u003c\/details\u003e\n\u003cdetails\u003e\u003csummary\u003eCan I put several in one server?\u003c\/summary\u003e\u003cp\u003eYes, and this is where the L4 is at its best — several cards at 72 W each stay within budgets that would be impossible with high-power GPUs. Kentino builds multi-L4 inference servers if you would rather buy the finished machine.\u003c\/p\u003e\u003c\/details\u003e\n\u003cdetails\u003e\u003csummary\u003eWhat is the lead time?\u003c\/summary\u003e\u003cp\u003eSourced to order, typically 10–21 days. Ask us to confirm current availability before planning around a date.\u003c\/p\u003e\u003c\/details\u003e\n\u003c\/div\u003e\n\u003cdiv class=\"rtx-note\"\u003eWant it in a finished machine? Kentino builds, benchmarks and commissions multi-L4 inference servers — ask us for a configured build.\u003c\/div\u003e\n\u003c\/div\u003e\n\u003c\/div\u003e","brand":"NVIDIA","offers":[{"title":"Default Title","offer_id":53618834899272,"sku":"TCSL4PCIE-PB","price":2657.42,"currency_code":"EUR","in_stock":true}],"url":"https:\/\/kentino.com\/he\/products\/nvidia-l4-24-gb-gddr6-low-profile-passive","provider":"Kentino","version":"1.0","type":"link"}