AI corner
Building a 576 GB VRAM Inference Server: A Comm...
Most server write-ups describe the machine that was ordered. This one describes the machine that was built — including the two things that were wrong with it when it first...
Building a 576 GB VRAM Inference Server: A Comm...
Most server write-ups describe the machine that was ordered. This one describes the machine that was built — including the two things that were wrong with it when it first...
We Built a 576 GB AI Server That Runs GLM 5.2 i...
Kentino's on-premise AI server packs six NVIDIA RTX PRO 6000 Blackwell Max-Q cards, dual EPYC CPUs and enough redundancy to survive a pulled power supply — so frontier-class models run...
We Built a 576 GB AI Server That Runs GLM 5.2 i...
Kentino's on-premise AI server packs six NVIDIA RTX PRO 6000 Blackwell Max-Q cards, dual EPYC CPUs and enough redundancy to survive a pulled power supply — so frontier-class models run...
Complete Technical Guide: Flashing CMP 170HX & ...
# Complete Technical Guide: Flashing CMP 170HX & Unlocking 64 GB HBM2e on Ubuntu 22.04 LTS This guide walks through the complete procedure to bypass factory OTP restrictions on the...
Complete Technical Guide: Flashing CMP 170HX & ...
# Complete Technical Guide: Flashing CMP 170HX & Unlocking 64 GB HBM2e on Ubuntu 22.04 LTS This guide walks through the complete procedure to bypass factory OTP restrictions on the...
Unlocking the Future of AI Compute: How Budget ...
The global explosion of artificial intelligence has created an unprecedented demand for high-density GPU compute. For startups, enterprise developers, and cloud infrastructure providers, acquiring top-tier enterprise accelerators like the NVIDIA...
Unlocking the Future of AI Compute: How Budget ...
The global explosion of artificial intelligence has created an unprecedented demand for high-density GPU compute. For startups, enterprise developers, and cloud infrastructure providers, acquiring top-tier enterprise accelerators like the NVIDIA...
Case Study: 4x RTX 4090 AI Workstation
This article documents a complete build commissioned for a research customer who needed a rack-mountable, 24/7-capable LLM inference workstation with enough VRAM to host 70B-class models without cloud dependency. Everything...
Case Study: 4x RTX 4090 AI Workstation
This article documents a complete build commissioned for a research customer who needed a rack-mountable, 24/7-capable LLM inference workstation with enough VRAM to host 70B-class models without cloud dependency. Everything...
TurboQuant: Reading the KV Cache Compression Br...
Reading time: 10 min | How Google's 3-bit compression makes long-context LLMs cheaper, and what it tells us about the next 18 months of AI inference There is a quiet...
TurboQuant: Reading the KV Cache Compression Br...
Reading time: 10 min | How Google's 3-bit compression makes long-context LLMs cheaper, and what it tells us about the next 18 months of AI inference There is a quiet...