SGLang HiCache
As long-context LLM inference scales, GPU High Bandwidth Memory (HBM) rapidly becomes the primary bottleneck due to linear Key-Value (KV) cache expansion. Integrating SGLang HiCache, host DRAM staging, and a POSIX L3 backend (NIXL) creates a disaggregated path that spills KV from GPU → host → storage, bypassing GPU memory bounds and cutting warm Time-To-First-Token (TTFT) when prefixes hit L3.
This KB guide details how to configure vLLM OffloadingConnector against high-performance VAST storage volumes over RDMA-enabled fabrics for production inference.
Prerequisites
Please make sure you have previously gone through the File Prerequisites before starting this guide, as it could drastically impact performance.
Motivation
GPUs: The testing environment utilizes 8× NVIDIA B200
Network:
The host is equipped with 2x Mellanox CX-7 400GbE Single-Port NICs
The network configuration for VAST storage is NFSv3 over RDMA
Storage: The team tested two primary configurations for KV cache offloading:
Local host: DRAM used as a transient HiCache staging buffer (
buffer_only; pages free after storage ack)Remote VAST Data storage: FS secondary tier on a VAST partition via NFSv3 over RDMA, with NFS multipath
Software:
SGLang: stock
lmsysorg/sglang:v0.5.20-cu130(CUDA 13)HiCache L3: NIXL POSIX plugin (
O_DIRECT+io_uring)NVIDIA CUDA: image-bundled CUDA 13 runtime
Model: Llama-4-Maverick-17B-128E-Instruct-FP8, TP-8
We evaluated a tp=8 configuration across concurrency levels, tuning the system according to the File Prerequisites guidelines to maximize performance:
nconnect: Set to 32
Read-ahead increased to 8192
vLLM OffloadingConnector optimized configurations, detailed below
Rank-local host pin baked into the image (required for large CPU tiers)
As for the evaluated parameters, we used:
ISL≈120k, OSL=1,
total prompts=concurrency
HiCache page size 8192
Layout / IO:
page_first_direct+directHost mode:
buffer_only(staging only; no host cache tier)Write policy:
write_throughPrefetch policy:
timeout(production default)Host pool: 200 GiB/rank (
--hicache-size 200),mem-fraction-static0.9
We don’t recommend using the builtinfile HiCache backend for this workload. It is single-threaded and page-cached; on the same shape it delivered only ~2.5× warm TTFT with ~72% storage load. Prefer NIXL (see Results).
Results
TTFT was measured cold (empty L3) vs warm (after full L3 dump + engine restart). Warm means L3 prefetch into GPU — not host-DRAM reuse.
Backend | Cold mean TTFT | Warm mean TTFT | Speedup | Storage tokens loaded |
|---|---|---|---|---|
HiCacheFile (reference) | ~146 s | ~58 s | ~2.5× | ~72% |
NIXL POSIX + uring (recommended) | ~147 s | ~30 s | ~4.9× | ~98.6% |
Software Stack
We'll be using the following NVIDIA stack:
SGLang — lmsysorg/sglang:v0.5.20-cu130
Model —
nvidia/Llama-4-Maverick-17B-128E-Instruct-FP8
Installation - NVIDIA
We’ll be working in the working directory /root/vastdata, and have cfg and models underneath it. We’ll also be using the /mnt/kvcache-2ports from earlier, so be sure you have this tree ready
BASE_DIR="/root/vastdata"
MODEL_DIR="${BASE_DIR}/models/llama-4-maverick"
CACHE_DIR="/mnt/kvcache-2ports"
CFG="${BASE_DIR}/cfg" # contains start.sh
mkdir -p "${BASE_DIR}" "${MODEL_DIR}" "${CFG}" "${CACHE_DIR}"
Downloading the model locally:
python3 -m venv "${BASE_DIR}/.venv"
source "${BASE_DIR}/.venv/bin/activate"
pip install "huggingface_hub[cli]"
hf download nvidia/Llama-4-Maverick-17B-128E-Instruct-FP8 \
--local-dir "${MODEL_DIR}"
Why NIXL (not HiCacheFile)
Stock HiCacheFile uses one synchronous buffered writer/reader per rank. The default timeout prefetch policy then admits requests before L3 IO finishes.
NIXL POSIX with use_direct_io + uring uses async multi-QD O_DIRECT IO. That is the lever that got us from ~2.5× to ~4.9× warm TTFT.
./cfg/start.sh
#!/usr/bin/env bash
set -euo pipefail
# HiCache host pool size per TP rank (GiB). Fleet host DRAM ≈ this × TP.
HICACHE_GB="${SGLANG_HICACHE_SIZE_GB:-200}"
PAGE_SIZE="${SGLANG_HICACHE_PAGE_SIZE:-8192}"
# NIXL POSIX extra config: O_DIRECT + io_uring (posix_aio is the fallback).
EXTRA_CONFIG='{"use_direct_io": true, "use_uring": true, "l3_cleaner_enabled": false}'
exec python3 -m sglang.launch_server \
--model-path /models/model \
--served-model-name llama4-maverick \
--host 0.0.0.0 --port 8000 \
--tp-size 8 \
--mem-fraction-static 0.9 \
--page-size "${PAGE_SIZE}" \
--chunked-prefill-size "${PAGE_SIZE}" \
--enable-hierarchical-cache \
--hicache-size "${HICACHE_GB}" \
--hicache-mem-layout page_first_direct \
--hicache-io-backend direct \
--hicache-write-policy write_through \
--hicache-host-memory-mode buffer_only \
--hicache-storage-backend nixl \
--hicache-storage-prefetch-policy timeout \
--hicache-storage-backend-extra-config "${EXTRA_CONFIG}" \
--skip-server-warmup
./run.sh
#!/usr/bin/env bash
set -euo pipefail
BASE_DIR="/root/vastdata"
MODEL_DIR="${BASE_DIR}/models/llama-4-maverick"
CACHE_DIR="/mnt/kvcache-2ports"
CFG="${BASE_DIR}/cfg"
IMAGE="lmsysorg/sglang:v0.5.20-cu130"
mkdir -p "${CACHE_DIR}"
chmod +x "${CFG}/start.sh"
docker run --name sglang-hicache-nixl -d \
--network host --ipc host --gpus all \
--ulimit memlock=-1 --ulimit stack=67108864 \
-v "${MODEL_DIR}:/models/model:ro" \
-v "${CACHE_DIR}:/hicache-nfs:rw" \
-v "${CFG}:/cfg:ro" \
-e SGLANG_HICACHE_NIXL_BACKEND_STORAGE_DIR=/hicache-nfs \
-e SGLANG_HICACHE_NIXL_BACKEND_PLUGIN=POSIX \
-e SGLANG_HICACHE_SIZE_GB=200 \
-e SGLANG_HICACHE_PAGE_SIZE=8192 \
-e SAFETENSORS_FAST_GPU=1 \
--entrypoint bash \
"${IMAGE}" /cfg/start.sh
If NIXL fails to create the uring backend (EBADF / plugin errors), change the extra-config selector to posix_aio (or set the harness --nixl-posix-async-io posix_aio) and keep use_direct_io: true.
Image
Pull the stock image — no bake step required for HiCache NIXL:
docker pull lmsysorg/sglang:v0.5.20-cu130
chmod +x cfg/start.sh run.sh
./run.shOperational notes (from field validation)
buffer_onlyfor L3 persistence cells. Default host-as-cache mode can couple SWA eviction to incomplete dumps. Staging-only host mode +write_throughis the clean L3 path.Hybrid SWA byte volume. For Llama-4, SGLang stores a full SWA window with every page — roughly 3× the disk bytes of vLLM OffloadingConnector for the same token set. NIXL removes the IO bottleneck; it does not remove that multiplier.
Additional Resources
Sibling pages:
Final Note
As workloads and software stacks frequently change, performance opportunities and compatibility shifts may occur. VAST actively expands its capabilities in the KV cache space—contact our team directly to leverage our latest updates and maximize your performance.