KI-Ecke

Building a 576 GB VRAM Inference Server: A Comm...

Most server write-ups describe the machine that was ordered. This one describes the machine that was built — including the two things that were wrong with it when it first...

Building a 576 GB VRAM Inference Server: A Comm...

Most server write-ups describe the machine that was ordered. This one describes the machine that was built — including the two things that were wrong with it when it first...

We Built a 576 GB AI Server That Runs GLM 5.2 in Your Own Building

We Built a 576 GB AI Server That Runs GLM 5.2 i...

Kentino's on-premise AI server packs six NVIDIA RTX PRO 6000 Blackwell Max-Q cards, dual EPYC CPUs and enough redundancy to survive a pulled power supply — so frontier-class models run...

We Built a 576 GB AI Server That Runs GLM 5.2 i...

Kentino's on-premise AI server packs six NVIDIA RTX PRO 6000 Blackwell Max-Q cards, dual EPYC CPUs and enough redundancy to survive a pulled power supply — so frontier-class models run...

Complete Technical Guide: Flashing CMP 170HX & Unlocking 64 GB HBM2e on Ubuntu 22.04 LTS

Complete Technical Guide: Flashing CMP 170HX & ...

# Complete Technical Guide: Flashing CMP 170HX & Unlocking 64 GB HBM2e on Ubuntu 22.04 LTS This guide walks through the complete procedure to bypass factory OTP restrictions on the...

Complete Technical Guide: Flashing CMP 170HX & ...

# Complete Technical Guide: Flashing CMP 170HX & Unlocking 64 GB HBM2e on Ubuntu 22.04 LTS This guide walks through the complete procedure to bypass factory OTP restrictions on the...

Unlocking the Future of AI Compute: How Budget GPUs and Unlocked CMP 170HX Are Revolutionizing LLM Hosting

Unlocking the Future of AI Compute: How Budget ...

The global explosion of artificial intelligence has created an unprecedented demand for high-density GPU compute. For startups, enterprise developers, and cloud infrastructure providers, acquiring top-tier enterprise accelerators like the NVIDIA...

Unlocking the Future of AI Compute: How Budget ...

The global explosion of artificial intelligence has created an unprecedented demand for high-density GPU compute. For startups, enterprise developers, and cloud infrastructure providers, acquiring top-tier enterprise accelerators like the NVIDIA...

Dark navy tech cover showing a 4x RTX 4090 GPU workstation build for a local AI case study.

Case Study: 4x RTX 4090 AI Workstation

This article documents a complete build commissioned for a research customer who needed a rack-mountable, 24/7-capable LLM inference workstation with enough VRAM to host 70B-class models without cloud dependency. Everything...

Case Study: 4x RTX 4090 AI Workstation

This article documents a complete build commissioned for a research customer who needed a rack-mountable, 24/7-capable LLM inference workstation with enough VRAM to host 70B-class models without cloud dependency. Everything...

Dark navy tech cover illustrating TurboQuant's KV cache compression breakthrough for large language models.

TurboQuant: Reading the KV Cache Compression Br...

Reading time: 10 min | How Google's 3-bit compression makes long-context LLMs cheaper, and what it tells us about the next 18 months of AI inference There is a quiet...

TurboQuant: Reading the KV Cache Compression Br...

Reading time: 10 min | How Google's 3-bit compression makes long-context LLMs cheaper, and what it tells us about the next 18 months of AI inference There is a quiet...