vLLM & LMCache Multi-Process (AMD)

Prev Next

vLLM & LMCache Multi-Process

As long-context LLM inference scales, GPU High Bandwidth Memory (HBM) rapidly becomes the primary bottleneck due to the linear expansion of the Key-Value (KV) cache. Integrating vLLM, LMCache, and VAST Data creates an ideal disaggregated infrastructure that offloads KV cache to a scale-out storage fabric, bypassing GPU memory bounds and slashing Time-To-First-Token (TTFT) by up to 10x.

This KB guide details how to configure LMCache and vLLM to mount high-performance VAST storage volumes over RDMA-enabled fabrics for production inference.

Prerequisites

Please make sure you have completed the File KVCache Prerequisites before starting this guide, as it can drastically impact performance.

Motivation

GPUs: The testing environment utilizes an AMD MI350X node.

  • Network:

    • The host is equipped with 2x Mellanox CX-7 400GbE Single-Port NICs.

    • The network configuration for VAST storage is NFSv3 over RDMA.

  • Storage: The team tested two primary configurations for KV cache offloading:

    • Local host: offloading to local host RAM.

    • Remote VAST Data storage: Offloading to a VAST storage partition via NFSv3 over RDMA, with NFS multipath.

  • Software: 

    • vLLM: Version 0.24

    • LMCache: Version 0.53

    • AMD ROCm: open foundation with programming models, compilers, libraries, runtimes, and deployment tools.

  • Model: GPT-oss 120B, TP-8

We evaluated a tp=8 configuration across concurrency levels ranging from 100 to 600, tuning the system according to the File Prerequisites guidelines to maximize performance:

  • nconnect: Set to 32

  • Read-ahead increased to 8192

  • vLLM & LMCache optimized configurations, detailed below

Results

TTFT was measured with and without the VAST storage backend (Layer 2). Both setups utilize L0/L1 caching (GPU & CPU).

~10x Warm tok/s gain

~10x Warm tok/s gain

~10x Warm mean TTFT gain

~10x Warm mean TTFT gain

Software Stack

We’ll be using the following AMD Stack:

  1. vLLM(*) - https://hub.docker.com/r/vllm/vllm-openai-rocm  (0.24.0+rocm723, vllm/vllm-openai-rocm@sha256:3832d79d9e514ce2e072580689da078726454596d833c8ab803f29f3cea5ea28)

  2. LMCache(*) - https://github.com/LMCache/LMCache/releases#release-v0.5.3-rocm (Release v0.5.3 · ROCm (gfx942, gfx950))

  3. Model - amd/gpt-oss-120b-w-mxfp4-a-fp8

Installation - AMD

We’ll be working in the working directory /root/vastdata, and have cfg and modelsunderneath it. We’ll also be using the /mnt/kvcache-2ports from earlier, so be sure you have this tree ready

BASE_DIR="/root/vastdata"
MODEL_DIR="${BASE_DIR}/models/gpt-oss-120b-w-mxfp4-a-fp8"
CACHE_DIR="/mnt/kvcache-2ports"
CFG="${BASE_DIR}/cfg"   # contains lmcache.yaml + start.sh

mkdir -p ${BASE_DIR} ${MODEL_DIR} ${CFG}

Downloading the model locally

python3 -m venv .venv
source .venv/bin/activate

pip install "huggingface_hub[cli]"

hf download amd/gpt-oss-120b-w-mxfp4-a-fp8 --local-dir "${MODEL_DIR}"

Scripts

./cfg/start.sh

#!/usr/bin/env bash
set -euo pipefail

lmcache server \
  --host 127.0.0.1 --port 5555 \
  --l1-size-gb 1600 \
  --eviction-policy noop \
  --eviction-trigger-watermark 0.8 --eviction-ratio 0.2 \
  --chunk-size 8192 \
  --l2-prefetch-max-in-flight 4 \
  --max-gpu-workers 8 --max-cpu-workers 8 \
  --worker-reap-timeout-seconds 180 \
  --l2-adapter '{"type":"fs_native","base_path":"/kvcache","num_workers":4,"use_odirect":true,"max_capacity_gb":0,"read_ahead_size":8192}' \
  --l1-align-bytes 1048576 \
  --l2-store-policy skip_l1 &

for _ in $(seq 1 60); do
  (echo > /dev/tcp/127.0.0.1/5555) >/dev/null 2>&1 && break
  sleep 1
done
(echo > /dev/tcp/127.0.0.1/5555) >/dev/null 2>&1 \
  || { echo "LMCache server failed readiness on :5555" >&2; exit 1; }

exec vllm serve /models/model \
  --host 0.0.0.0 --port 8000 \
  --served-model-name gpt-oss-120b \
  --tensor-parallel-size 8 \
  --max-model-len 131072 \
  --disable-hybrid-kv-cache-manager \
  --enable-prefix-caching \
  --block-size=64 \
  --max-num-seqs 256 \
  --kv-transfer-config \
  '{"kv_connector":"LMCacheMPConnector","kv_connector_module_path":"lmcache.integration.vllm.lmcache_mp_connector","kv_role":"kv_both","kv_connector_extra_config":{"lmcache.mp.host":"tcp://127.0.0.1","lmcache.mp.port":5555,"lmcache.mp.heartbeat_interval":60.0}}'

./cfg/lmcache.yaml

chunk_size: 8192
save_decode_cache: false
enable_async_loading: true
blocking_timeout_secs: 120
lookup_timeout_ms: 1800000
mp_host: 127.0.0.1
mp_port: 5555

./run.sh

BASE_DIR="/root/vastdata"
MODEL_DIR="${BASE_DIR}/models/gpt-oss-120b-w-mxfp4-a-fp8"
CACHE_DIR="/mnt/kvcache-2ports"
CFG="${BASE_DIR}/cfg"   # contains lmcache.yaml + start.sh
mkdir -p "$CFG"
chmod +x "$CFG/start.sh"

docker run --name vllm-lmcache-mp-fs -d \
  --network host --ipc host \
  --ulimit memlock=-1 --ulimit stack=67108864 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add render \
  -v "${MODEL_DIR}:/models/model:ro" \
  -v "${CACHE_DIR}:/kvcache:rw" \
  -v "${CFG}:/cfg:ro" \
  -e LMCACHE_CONFIG_FILE=/cfg/lmcache.yaml \
  -e PYTHONHASHSEED=0 \
  -e HIP_FORCE_DEV_KERNARG=1 -e HSA_NO_SCRATCH_RECLAIM=1 \
  -e TORCH_BLAS_PREFER_HIPBLASLT=1 -e SAFETENSORS_FAST_GPU=1 \
  -e VLLM_ROCM_USE_AITER=1 -e VLLM_ROCM_USE_AITER_MHA=0 \
  -e VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1 \
  -e VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4 \
  -e NCCL_MIN_NCHANNELS=112 \
  --entrypoint bash vllm-lmcache:0.24.0-rocm-lmcache0.5.3 /cfg/start.sh

./Dockerfile

# vllm-lmcache:0.24.0-rocm-lmcache0.5.3
FROM vllm/vllm-openai-rocm@sha256:3832d79d9e514ce2e072580689da078726454596d833c8ab803f29f3cea5ea28
ARG LMCACHE_VERSION=0.5.3

# PyPI deps, then overwrite with the official ROCm wheel (--no-deps keeps
# the base image's ROCm torch). Swap CUDA CuPy for ROCm CuPy.
RUN python3 -m pip install --no-cache-dir "lmcache==${LMCACHE_VERSION}" \
 && python3 -m pip uninstall -y cupy-cuda13x cupy-cuda12x cufile-python || true \
 && python3 -m pip install --no-cache-dir "cupy-rocm-7-0==14.1.1" \
 && python3 -m pip install --no-cache-dir --force-reinstall --no-deps --no-index \
      --find-links "https://github.com/LMCache/LMCache/releases/expanded_assets/v${LMCACHE_VERSION}-rocm" \
      "lmcache==${LMCACHE_VERSION}" \
 && python3 -c "import lmcache; print('lmcache', lmcache.__version__)"

Build the image & Run it

docker build -t vllm-lmcache:0.24.0-rocm-lmcache0.5.3 -f Dockerfile .

chmod +x run.sh
./run.sh

Additional Resources

https://hub.docker.com/r/rocm/vllm/tags (AMD published images, not used here but shared for knowledge)

https://github.com/ROCm/rocm-aic  (AMD AIC, not production-ready at time of posting)

(*) While optimized, they are frequently not the latest vLLM builds

Final Note

As workloads and software stacks frequently change, performance opportunities and compatibility shifts may occur. VAST actively expands its capabilities in the KV cache space—contact our team directly to leverage our latest updates and maximize your performance.