SGLang HiCache using NIXL (NVIDIA)

Prev Next

SGLang HiCache

As long-context LLM inference scales, GPU High Bandwidth Memory (HBM) rapidly becomes the primary bottleneck due to linear Key-Value (KV) cache expansion. Integrating SGLang HiCache, host DRAM staging, and a POSIX L3 backend (NIXL) creates a disaggregated path that spills KV from GPU → host → storage, bypassing GPU memory bounds and cutting warm Time-To-First-Token (TTFT) when prefixes hit L3.

This KB guide details how to configure vLLM OffloadingConnector against high-performance VAST storage volumes over RDMA-enabled fabrics for production inference.

Prerequisites

Please make sure you have previously gone through the File Prerequisites before starting this guide, as it could drastically impact performance.

Motivation

GPUs: The testing environment utilizes 8× NVIDIA B200

  • Network:

    • The host is equipped with 2x Mellanox CX-7 400GbE Single-Port NICs

    • The network configuration for VAST storage is NFSv3 over RDMA

  • Storage: The team tested two primary configurations for KV cache offloading:

    • Local host: DRAM used as a transient HiCache staging buffer (buffer_only; pages free after storage ack)

    • Remote VAST Data storage: FS secondary tier on a VAST partition via NFSv3 over RDMA, with NFS multipath

  • Software:

    • SGLang: stock lmsysorg/sglang:v0.5.20-cu130 (CUDA 13)

    • HiCache L3: NIXL POSIX plugin (O_DIRECT + io_uring)

    • NVIDIA CUDA: image-bundled CUDA 13 runtime

  • Model: Llama-4-Maverick-17B-128E-Instruct-FP8, TP-8

We evaluated a tp=8 configuration across concurrency levels, tuning the system according to the File Prerequisites guidelines to maximize performance:

  • nconnect: Set to 32

  • Read-ahead increased to 8192

  • vLLM OffloadingConnector optimized configurations, detailed below

  • Rank-local host pin baked into the image (required for large CPU tiers)

As for the evaluated parameters, we used:

  • ISL≈120k, OSL=1,

  • total prompts=concurrency

  • HiCache page size 8192

  • Layout / IO: page_first_direct + direct

  • Host mode: buffer_only (staging only; no host cache tier)

  • Write policy: write_through

  • Prefetch policy: timeout (production default)

  • Host pool: 200 GiB/rank (--hicache-size 200), mem-fraction-static 0.9

We don’t recommend using the builtinfile HiCache backend for this workload. It is single-threaded and page-cached; on the same shape it delivered only ~2.5× warm TTFT with ~72% storage load. Prefer NIXL (see Results).

Results

TTFT was measured cold (empty L3) vs warm (after full L3 dump + engine restart). Warm means L3 prefetch into GPU — not host-DRAM reuse.

Backend

Cold mean TTFT

Warm mean TTFT

Speedup

Storage tokens loaded

HiCacheFile (reference)

~146 s

~58 s

~2.5×

~72%

NIXL POSIX + uring (recommended)

~147 s

~30 s

~4.9×

~98.6%

Software Stack

We'll be using the following NVIDIA stack:

  1. SGLang — lmsysorg/sglang:v0.5.20-cu130

  2. Model — nvidia/Llama-4-Maverick-17B-128E-Instruct-FP8

Installation - NVIDIA

We’ll be working in the working directory /root/vastdata, and have cfg and models underneath it. We’ll also be using the /mnt/kvcache-2ports from earlier, so be sure you have this tree ready

BASE_DIR="/root/vastdata"
MODEL_DIR="${BASE_DIR}/models/llama-4-maverick"
CACHE_DIR="/mnt/kvcache-2ports"
CFG="${BASE_DIR}/cfg"            # contains start.sh
mkdir -p "${BASE_DIR}" "${MODEL_DIR}" "${CFG}" "${CACHE_DIR}"

Downloading the model locally:

python3 -m venv "${BASE_DIR}/.venv"
source "${BASE_DIR}/.venv/bin/activate"
pip install "huggingface_hub[cli]"
hf download nvidia/Llama-4-Maverick-17B-128E-Instruct-FP8 \
  --local-dir "${MODEL_DIR}"

Why NIXL (not HiCacheFile)

Stock HiCacheFile uses one synchronous buffered writer/reader per rank. The default timeout prefetch policy then admits requests before L3 IO finishes.

NIXL POSIX with use_direct_io + uring uses async multi-QD O_DIRECT IO. That is the lever that got us from ~2.5× to ~4.9× warm TTFT.

./cfg/start.sh

#!/usr/bin/env bash
set -euo pipefail

# HiCache host pool size per TP rank (GiB). Fleet host DRAM ≈ this × TP.
HICACHE_GB="${SGLANG_HICACHE_SIZE_GB:-200}"
PAGE_SIZE="${SGLANG_HICACHE_PAGE_SIZE:-8192}"

# NIXL POSIX extra config: O_DIRECT + io_uring (posix_aio is the fallback).
EXTRA_CONFIG='{"use_direct_io": true, "use_uring": true, "l3_cleaner_enabled": false}'

exec python3 -m sglang.launch_server \
  --model-path /models/model \
  --served-model-name llama4-maverick \
  --host 0.0.0.0 --port 8000 \
  --tp-size 8 \
  --mem-fraction-static 0.9 \
  --page-size "${PAGE_SIZE}" \
  --chunked-prefill-size "${PAGE_SIZE}" \
  --enable-hierarchical-cache \
  --hicache-size "${HICACHE_GB}" \
  --hicache-mem-layout page_first_direct \
  --hicache-io-backend direct \
  --hicache-write-policy write_through \
  --hicache-host-memory-mode buffer_only \
  --hicache-storage-backend nixl \
  --hicache-storage-prefetch-policy timeout \
  --hicache-storage-backend-extra-config "${EXTRA_CONFIG}" \
  --skip-server-warmup

./run.sh

#!/usr/bin/env bash
set -euo pipefail

BASE_DIR="/root/vastdata"
MODEL_DIR="${BASE_DIR}/models/llama-4-maverick"
CACHE_DIR="/mnt/kvcache-2ports"
CFG="${BASE_DIR}/cfg"
IMAGE="lmsysorg/sglang:v0.5.20-cu130"

mkdir -p "${CACHE_DIR}"
chmod +x "${CFG}/start.sh"

docker run --name sglang-hicache-nixl -d \
  --network host --ipc host --gpus all \
  --ulimit memlock=-1 --ulimit stack=67108864 \
  -v "${MODEL_DIR}:/models/model:ro" \
  -v "${CACHE_DIR}:/hicache-nfs:rw" \
  -v "${CFG}:/cfg:ro" \
  -e SGLANG_HICACHE_NIXL_BACKEND_STORAGE_DIR=/hicache-nfs \
  -e SGLANG_HICACHE_NIXL_BACKEND_PLUGIN=POSIX \
  -e SGLANG_HICACHE_SIZE_GB=200 \
  -e SGLANG_HICACHE_PAGE_SIZE=8192 \
  -e SAFETENSORS_FAST_GPU=1 \
  --entrypoint bash \
  "${IMAGE}" /cfg/start.sh

If NIXL fails to create the uring backend (EBADF / plugin errors), change the extra-config selector to posix_aio (or set the harness --nixl-posix-async-io posix_aio) and keep use_direct_io: true.

Image

Pull the stock image — no bake step required for HiCache NIXL:

docker pull lmsysorg/sglang:v0.5.20-cu130
chmod +x cfg/start.sh run.sh
./run.sh

Operational notes (from field validation)

  1. buffer_only for L3 persistence cells. Default host-as-cache mode can couple SWA eviction to incomplete dumps. Staging-only host mode + write_through is the clean L3 path.

  2. Hybrid SWA byte volume. For Llama-4, SGLang stores a full SWA window with every page — roughly 3× the disk bytes of vLLM OffloadingConnector for the same token set. NIXL removes the IO bottleneck; it does not remove that multiplier.


Additional Resources

Final Note

As workloads and software stacks frequently change, performance opportunities and compatibility shifts may occur. VAST actively expands its capabilities in the KV cache space—contact our team directly to leverage our latest updates and maximize your performance.