Dynamo using vLLM

Prev Next

Dynamo + vLLM OffloadingConnector

As long-context LLM inference scales, GPU High Bandwidth Memory (HBM) rapidly becomes the primary bottleneck due to linear Key-Value (KV) cache expansion. Integrating NVIDIA Dynamo (OpenAI-compatible frontend + worker orchestration on Kubernetes) with vLLM’s native OffloadingConnector, host DRAM, and VAST Data creates a disaggregated path that spills KV from GPU → CPU → NFS, bypassing GPU memory bounds and cutting warm Time-To-First-Token (TTFT) when prefixes hit the CPU/FS tiers.

This guide details how to run Dynamo with a vLLM OffloadingConnector worker against high-performance VAST storage — on Kubernetes, with optional VAST CSI for the KV volume, and with either a bundled (combined) or unbundled (split) frontend layout.

Prerequisites

  1. Complete the VAST-specific File Prerequisites first — NFS/RDMA tuning here can drastically affect OffloadingConnector FS-tier performance.

  2. You need a Kubernetes v1.30+ cluster with NVIDIA GPU nodes, kubectl v1.30+, and Helm v3+.

  3. Install the Dynamo platform (GPU Operator + dynamo-platform) using the Dynamo platform setup below (based on NVIDIA’s Kubernetes Quickstart).

Dynamo platform setup

Note:

NVIDIA’s Dynamo documentation changes over time. The commands in this section are based on the Dynamo 1.5.0 vLLM Kubernetes Quickstart. Prefer the live page if it differs: Kubernetes Quickstart.

Install accelerator support

Install the NVIDIA GPU Operator:

helm repo add nvidia https://helm.ngc.nvidia.com/nvidia --force-update
helm repo update nvidia
helm upgrade --install gpu-operator nvidia/gpu-operator \
  --namespace gpu-operator \
  --create-namespace \
  --wait \
  --timeout=10m

Info

Tip: If your cluster provider installs the NVIDIA driver, add --set driver.enabled=false. Add --set toolkit.enabled=false only when the provider also configures the GPU container runtime.

Install Dynamo

export NAMESPACE=dynamo-system
export DYNAMO_VERSION=1.5.0

helm upgrade --install dynamo-platform \
  "https://helm.ngc.nvidia.com/nvidia/ai-dynamo/charts/dynamo-platform-${DYNAMO_VERSION}.tgz" \
  --namespace "$NAMESPACE" \
  --create-namespace \
  --wait \
  --timeout=10m

kubectl get pods --namespace "$NAMESPACE"

Hugging Face token secret (optional — Quickstart only)

NVIDIA’s Dynamo Quickstart creates a Secret named hf-token-secret and wires it into sample DGDs via envFromSecret. That is fine to create if you follow their examples, but this guide does not use it: workers load Scout from a local hostPath (--model /models/llama-4-scout), so there is no runtime Hugging Face pull, and our YAML below omits envFromSecret.

If you still want the Secret for Quickstart parity (or other Dynamo samples), you can create it — it is unused by the instructions here:

kubectl create secret generic hf-token-secret \
  --namespace "$NAMESPACE" \
  --from-literal=HF_TOKEN=<YOUR_HF_TOKEN> \
  --dry-run=client -o yaml | kubectl apply -f -

Replace <YOUR_HF_TOKEN> with your Hugging Face access token (example shape: hf_xxxxxxxx), or leave an empty string for public models.

At this point, the platform is installed. The Quickstart’s sample Qwen DynamoGraphDeploymentRequest deploy is optional validation only — this guide continues with a Scout + OffloadingConnector DynamoGraphDeployment and VAST-backed KV instead.

Motivation

GPUs: The testing environment utilizes 2x NVIDIA RTX PRO 6000 Blackwell

  • Network:

    • The host is equipped with 2x Mellanox CX-7 200GbE Single-Port NICs

    • The network configuration for VAST storage is NFSv3 over RDMA

  • Storage: The team tested two primary configurations for KV cache offloading:

    • Local host: OffloadingConnector CPU DRAM tier (pod sharedMemory / mmap)

    • Remote VAST Data storage: FS secondary tier on a VAST partition via NFSv3 over RDMA, with NFS multipath — either as a hostPath mount of the multipath export, or via VAST CSI (see optional CSI guide)

  • Host memory: The GPU node used for these recipes has ~500 GiB of system RAM. That leaves room for the OS, the model runtime, and a large OffloadingConnector CPU tier without swapping.

  • Software:

    • Dynamo: Version 1.5.0 (dynamo-platform on Kubernetes)

    • vLLM: Version 0.28.0 (OffloadingConnector worker image)

    • NVIDIA CUDA: proprietary accelerated computing platform

  • Model: Llama-4-Scout-17B-16E-Instruct-FP8, TP-2

We evaluated a tp=2 configuration, tuning the system according to the File Prerequisites guidelines:

  • nconnect: Set to 32

  • Read-ahead increased to 8192

  • vLLM OffloadingConnector optimized configurations, detailed below

  • Rank-local host pin baked into the image (required for large CPU tiers)

  • CPU offload tier sized to 160 GiB (cpu_bytes_to_use: 171798691840 = (160 X 1024^3)): large enough to hold warm prefixes for Scout at this concurrency, while staying well under the ~500 GiB host RAM after OS, Dynamo/vLLM, and Scout weights. Pod sharedMemorySize is set to 200Gi so the mmap backing that 160 GiB tier has margin.

Results

Performance numbers TBD — same cold/warm TTFT methodology as the vLLM (NVIDIA) OffloadingConnector guide. Placeholder until Dynamo + OffloadingConnector measurements are published here.

Software Stack

We’ll be using the following NVIDIA stack:

  1. Dynamo platform — Helm chart dynamo-platform-1.5.0 (see Quickstart link above)

  2. Dynamo + vLLM OffloadingConnector image — built below as dynamo-vllm-offloading:1.5.0-vllm0.28.0    - Base: nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.5.0 (already ships vLLM 0.28.0 — no separate vllm/vllm-openai layer)    - Overlay: rank-local host-pin bake only (required for a large OffloadingConnector CPU tier)

  3. Model — nvidia/Llama-4-Scout-17B-16E-Instruct-FP8

Installation - NVIDIA

We’ll be working in the working directory /root/vastdata, with cfg, yamls, and models underneath it. We’ll also be using the /mnt/kvcache-2ports multipath mount from the File Prerequisites, so be sure you have this tree ready:

BASE_DIR="/root/vastdata"
MODEL_DIR="${BASE_DIR}/models/llama-4-scout"
CACHE_DIR="/mnt/kvcache-2ports"
CFG="${BASE_DIR}/cfg"
YAMLS="${BASE_DIR}/yamls"
mkdir -p "${BASE_DIR}" "${MODEL_DIR}" "${CFG}" "${YAMLS}"

Build package (attached)

Download dynamo-w-vllm-package.zip from this page’s attachments, then unpack it into ${BASE_DIR}:

dynamo-w-vllm-package
22.60 KB

# from the machine where you will docker-build (usually the GPU node)
unzip dynamo-w-vllm-package.zip -d "${BASE_DIR}"
cd "${BASE_DIR}/dynamo-w-vllm-package"

Path

Role

Dockerfile

NGC vllm-runtime:1.5.0 + rank-local host pin

build_images.sh

docker build wrapper

scripts/bake_vllm_rank_local_pin.sh

Installs the patched modules into the image

patches/gpu_worker.py

Rank-local pin_mmap_region

patches/shared_offload_region.py

Matching multi-pointer cleanup

yamls/dgd-combined-hostpath.yaml

Path A — bundled + hostPath

yamls/dgd-split-hostpath.yaml

Path B — unbundled + hostPath

yamls/dgd-combined-csi.yaml

Path C — bundled + CSI PVC

yamls/dgd-split-csi.yaml

Path D — unbundled + CSI PVC

Copy the YAML you need into ${YAMLS}/ (or apply from yamls/ in place). Replace <GPU-NODE-NAME> (and for C/D, <NAME-FROM-CSI-GUIDE-pvc.yaml>) before kubectl apply. The Dockerfile / bake script / Path A–B snippets below match the zip — kept inline for review.

Downloading the model locally

Scout FP8 is gated on Hugging Face — export a token in this shell for the download only (not a Kubernetes Secret):

export HF_TOKEN=<YOUR_HF_TOKEN>   # e.g. hf_xxxxxxxx; required for nvidia/Llama-4-Scout-…

python3 -m venv "${BASE_DIR}/.venv"
source "${BASE_DIR}/.venv/bin/activate"
pip install "huggingface_hub[cli]"
hf download nvidia/Llama-4-Scout-17B-16E-Instruct-FP8 \
  --local-dir "${MODEL_DIR}"

Why the Dockerfile patches two files

Stock (unpatched) vLLM pin_mmap_region registers the entire shared offload mmap in every tensor-parallel rank. At higher TP with a large CPU tier, it asks the driver for an excessive amount of pinned mappings and can fail or destabilize the process.

The files under patches/ change that to rank-local registration (page-aligned strided slots) and matching unregister in cleanup. Aggregate pinned bytes stay ~tier size. This is a source replacement in the image, not a runtime mount.

Look for this line once per rank after engine start:

Rank-local host registration complete: rank=N rows=... ... GiB (shared region ... GiB)

./Dockerfile

vllm-runtime:1.5.0 already includes vLLM 0.28. Build from that single base and bake the rank-local pin. (Older Dynamo 1.4.2 runtimes shipped vLLM 0.26 and needed a second-stage swap from vllm/vllm-openai — that is not required here.)

# Dynamo frontend + vLLM OffloadingConnector worker:
# NGC Dynamo 1.5.0 runtime (ships vLLM 0.28) + rank-local host pin.
# No second vLLM image — unlike the Dynamo 1.4.2 + swap recipe.

ARG DYNAMO_IMAGE=nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.5.0

FROM ${DYNAMO_IMAGE}

# NGC runtime defaults to uid 1000 (dynamo); installs need root.
USER root

ENV CUDA_HOME=/usr/local/cuda

COPY scripts/bake_vllm_rank_local_pin.sh /tmp/scripts/
COPY patches/gpu_worker.py patches/shared_offload_region.py /tmp/vllm-rank-local-pin/

RUN chmod +x /tmp/scripts/*.sh \
 && bash /tmp/scripts/bake_vllm_rank_local_pin.sh /tmp/vllm-rank-local-pin \
 && rm -rf /tmp/scripts /tmp/vllm-rank-local-pin

RUN python3 - <<'PY'
import importlib
from pathlib import Path
import vllm
import vllm.v1.kv_offload.cpu.gpu_worker as gw
import vllm.v1.kv_offload.cpu.shared_offload_region as sor
assert vllm.__version__.startswith("0.28"), vllm.__version__
assert "Rank-local host registration complete" in Path(gw.__file__).read_text()
assert "_rank_local_registered_ptrs" in Path(sor.__file__).read_text()
for name in ("dynamo", "dynamo.frontend", "dynamo.vllm"):
    importlib.import_module(name)
print("ok dynamo-1.5.0 vllm", vllm.__version__)
PY

USER dynamo

Build the image

Log in to NGC so Docker can pull the private base, then build from the unpacked package directory (the tree that contains Dockerfile, scripts/, and patches/):

# NGC API key that can pull nvcr.io/nvidia/ai-dynamo/vllm-runtime
docker login nvcr.io

cd "${BASE_DIR}/dynamo-w-vllm-package"
chmod +x build_images.sh scripts/bake_vllm_rank_local_pin.sh
./build_images.sh
# equivalent:
# docker build \
#   --build-arg DYNAMO_IMAGE=nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.5.0 \
#   -t dynamo-vllm-offloading:1.5.0-vllm0.28.0 .

Load the image into the node container runtime used by Kubernetes if the DGD uses imagePullPolicy: Never (for example docker save … | ctr -n k8s.io images import -, or your cluster’s usual private registry push/pull).

Choose a deployment layout

Layout

What you apply

When to use

Bundled frontend (combined)

One DynamoGraphDeployment with Frontend + VllmDecodeWorker

Simplest path; frontend and worker share one DGD

Unbundled frontend (split)

Two DGDs: dynamo-vllm-frontend then dynamo-vllm-worker, both with globalDynamoNamespace: true

Frontend Ready before the worker attaches; clearer lifecycle split

KV volume

What you use

When to use

hostPath

Node mount /mnt/kvcache-2ports → in-pod /kvcache

Multipath NFS already mounted on the GPU node

VAST CSI

PVC from the CSI guide → in-pod /kvcache

You want Kubernetes to manage the VAST view/volume

If VAST CSI is desired, complete Dynamo - VAST CSI (Optional) through a Bound PVC first, then return here and apply the CSI DGD variant whose claimName matches that PVC. Do not duplicate CSI Helm/secret/PV steps in this guide.

Otherwise, continue with hostPath below.

Values you must set before applying

The YAML examples use placeholders. Replace them before kubectl apply:

Placeholder

Meaning

How to find / example

<GPU-NODE-NAME>

Kubernetes node that has the GPUs (and, for hostPath, the /mnt/kvcache-2ports mount)

kubectl get nodes -o wide → e.g. gpu-node-01

<NAME-FROM-CSI-GUIDE-pvc.yaml> (CSI paths only)

PVC metadata name created in the CSI guide

kubectl get pvc -A after that guide → e.g. whatever pvc.yaml sets as metadata.name

Also confirm the DynamoGraphDeployment API version matches your CRD (kubectl api-resources | grep -i dynamograph). These snippets use nvidia.com/v1beta1 with spec.components + podTemplate (catalog default for Dynamo platform 1.5.0). Platform 1.4.2 still defaults to nvidia.com/v1alpha1 (spec.services); override with --k8s-api-version / node kubernetes.api_version. Do not change only the apiVersion string — the schemas differ.

Worker OffloadingConnector config (shared)

All variants pass the same --kv-transfer-config into python3 -m dynamo.vllm.

Why cpu_bytes_to_use: 171798691840? That integer is 160 GiB in bytes ((160 \times 1024^3 = 171798691840)). On a ~500 GiB RAM host, we reserve:

  • ~160 GiB for the OffloadingConnector CPU tier (warm KV in DRAM),

  • ~200 GiB pod sharedMemory so the mmap backing that tier fits with margin,

  • remaining RAM for the OS, Dynamo frontend, vLLM engine, and Scout weights in GPU/host page cache.

Shrink or grow cpu_bytes_to_use (and sharedMemorySize) if your node has less or more RAM; keep the CPU tier clearly below total RAM minus a comfortable OS/runtime reserve.

FS secondary root is the in-pod path /kvcache (hostPath or CSI mounts there):

{
  "kv_connector": "OffloadingConnector",
  "kv_role": "kv_both",
  "kv_connector_extra_config": {
    "spec_name": "TieringOffloadingSpec",
    "cpu_bytes_to_use": 171798691840,
    "block_size": 8192,
    "eviction_policy": "lru",
    "offload_prompt_only": true,
    "secondary_tiers": [{
      "type": "fs",
      "root_dir": "/kvcache",
      "n_read_threads": 32,
      "n_write_threads": 16
    }]
  }
}

Path A — Bundled frontend + hostPath KV

Save as ${YAMLS}/dgd-combined-hostpath.yaml. Set nodeName to your GPU node (see Values you must set before apply; example: gpu-node-01).

apiVersion: nvidia.com/v1beta1
kind: DynamoGraphDeployment
metadata:
  name: dynamo-vllm
  namespace: dynamo-system
spec:
  backendFramework: vllm
  components:
  - name: Frontend
    type: frontend
    replicas: 1
    podTemplate:
      spec:
        nodeName: <GPU-NODE-NAME>   # e.g. gpu-node-01
        terminationGracePeriodSeconds: 120
        containers:
        - name: main
          image: dynamo-vllm-offloading:1.5.0-vllm0.28.0
          imagePullPolicy: Never
          command: ["python3", "-m", "dynamo.frontend"]
          args: ["--http-port", "8000", "--router-mode", "round-robin"]
  - name: VllmDecodeWorker
    type: worker
    replicas: 1
    sharedMemorySize: 200Gi
    podTemplate:
      spec:
        nodeName: <GPU-NODE-NAME>   # same node as Frontend
        terminationGracePeriodSeconds: 600
        volumes:
        - name: model
          hostPath:
            path: /root/vastdata/models/llama-4-scout
            type: Directory
        - name: kv-fs
          hostPath:
            path: /mnt/kvcache-2ports
            type: DirectoryOrCreate
        containers:
        - name: main
          image: dynamo-vllm-offloading:1.5.0-vllm0.28.0
          imagePullPolicy: Never
          env:
          - name: PYTHONHASHSEED
            value: "0"
          - name: PYTHONUNBUFFERED
            value: "1"
          - name: DYN_SYSTEM_PORT
            value: "8081"
          - name: VLLM_SERVER_DEV_MODE
            value: "1"
          - name: HF_HOME
            value: /tmp/hf-home
          volumeMounts:
          - name: model
            mountPath: /models/llama-4-scout
            readOnly: true
          - name: kv-fs
            mountPath: /kvcache
          command: ["python3", "-m", "dynamo.vllm"]
          args:
          - --model
          - /models/llama-4-scout
          - --served-model-name
          - llama-4-scout
          - --tensor-parallel-size
          - "2"
          - --gpu-memory-utilization
          - "0.92"
          - --enable-prefix-caching
          - --block-size
          - "64"
          - --max-model-len
          - "131072"
          - --max-num-batched-tokens
          - "8192"
          - --max-num-seqs
          - "256"
          - --kv-transfer-config
          - '{"kv_connector":"OffloadingConnector","kv_role":"kv_both","kv_connector_extra_config":{"spec_name":"TieringOffloadingSpec","cpu_bytes_to_use":171798691840,"block_size":8192,"eviction_policy":"lru","offload_prompt_only":true,"secondary_tiers":[{"type":"fs","root_dir":"/kvcache","n_read_threads":32,"n_write_threads":16}]}}'
          resources:
            requests:
              nvidia.com/gpu: "2"
              cpu: "16"
              memory: 220Gi
            limits:
              nvidia.com/gpu: "2"
              cpu: "32"
              memory: 240Gi
              ephemeral-storage: 50Gi
          startupProbe:
            httpGet: { path: /health, port: 8081 }
            failureThreshold: 360
            periodSeconds: 10
          readinessProbe:
            httpGet: { path: /health, port: 8081 }
            failureThreshold: 30
            periodSeconds: 10
          livenessProbe:
            httpGet: { path: /live, port: 8081 }
            failureThreshold: 10
            periodSeconds: 30
kubectl apply -f "${YAMLS}/dgd-combined-hostpath.yaml"
kubectl get pods -n dynamo-system -w

Port-forward and smoke request (bundled)

When the worker pod is Ready, forward the frontend Service to localhost and hit the OpenAI-compatible API:

# Discover the frontend Service name (bundled DGD → typically dynamo-vllm-frontend)
kubectl get svc -n dynamo-system | grep -i frontend

export FRONTEND_SERVICE=dynamo-vllm-frontend   # change if your Service name differs
kubectl port-forward -n dynamo-system "svc/${FRONTEND_SERVICE}" 8000:8000

In another shell:

# Wait until the model is registered
until curl -sf http://127.0.0.1:8000/v1/models | grep -q llama-4-scout; do sleep 2; done

curl -s http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "llama-4-scout",
    "messages": [{"role": "user", "content": "Say hello in one sentence."}],
    "max_tokens": 32
  }' | python3 -m json.tool

Stop the forwarder with Ctrl-C in the port-forward terminal when finished.


Path B — Unbundled frontend + hostPath KV

Save as ${YAMLS}/dgd-split-hostpath.yaml (two documents). Apply the frontend DGD first; wait until it is Ready; then apply/ensure the worker DGD. Both set globalDynamoNamespace: true so the worker can attach without recreating the frontend.

Set nodeName to your GPU node (see Values you must set before apply; example: gpu-node-01).

apiVersion: nvidia.com/v1beta1
kind: DynamoGraphDeployment
metadata:
  name: dynamo-vllm-frontend
  namespace: dynamo-system
spec:
  backendFramework: vllm
  components:
  - name: Frontend
    type: frontend
    replicas: 1
    globalDynamoNamespace: true
    podTemplate:
      spec:
        nodeName: <GPU-NODE-NAME>   # e.g. gpu-node-01
        terminationGracePeriodSeconds: 120
        containers:
        - name: main
          image: dynamo-vllm-offloading:1.5.0-vllm0.28.0
          imagePullPolicy: Never
          command: ["python3", "-m", "dynamo.frontend"]
          args: ["--http-port", "8000", "--router-mode", "round-robin"]
---
apiVersion: nvidia.com/v1beta1
kind: DynamoGraphDeployment
metadata:
  name: dynamo-vllm-worker
  namespace: dynamo-system
spec:
  backendFramework: vllm
  components:
  - name: VllmDecodeWorker
    type: worker
    replicas: 1
    globalDynamoNamespace: true
    sharedMemorySize: 200Gi
    podTemplate:
      spec:
        nodeName: <GPU-NODE-NAME>   # same node as Frontend
        terminationGracePeriodSeconds: 600
        volumes:
        - name: model
          hostPath:
            path: /root/vastdata/models/llama-4-scout
            type: Directory
        - name: kv-fs
          hostPath:
            path: /mnt/kvcache-2ports
            type: DirectoryOrCreate
        containers:
        - name: main
          image: dynamo-vllm-offloading:1.5.0-vllm0.28.0
          imagePullPolicy: Never
          env:
          - name: PYTHONHASHSEED
            value: "0"
          - name: PYTHONUNBUFFERED
            value: "1"
          - name: DYN_SYSTEM_PORT
            value: "8081"
          - name: VLLM_SERVER_DEV_MODE
            value: "1"
          - name: HF_HOME
            value: /tmp/hf-home
          volumeMounts:
          - name: model
            mountPath: /models/llama-4-scout
            readOnly: true
          - name: kv-fs
            mountPath: /kvcache
          command: ["python3", "-m", "dynamo.vllm"]
          args:
          - --model
          - /models/llama-4-scout
          - --served-model-name
          - llama-4-scout
          - --tensor-parallel-size
          - "2"
          - --gpu-memory-utilization
          - "0.92"
          - --enable-prefix-caching
          - --block-size
          - "64"
          - --max-model-len
          - "131072"
          - --max-num-batched-tokens
          - "8192"
          - --max-num-seqs
          - "256"
          - --kv-transfer-config
          - '{"kv_connector":"OffloadingConnector","kv_role":"kv_both","kv_connector_extra_config":{"spec_name":"TieringOffloadingSpec","cpu_bytes_to_use":171798691840,"block_size":8192,"eviction_policy":"lru","offload_prompt_only":true,"secondary_tiers":[{"type":"fs","root_dir":"/kvcache","n_read_threads":32,"n_write_threads":16}]}}'
          resources:
            requests:
              nvidia.com/gpu: "2"
              cpu: "16"
              memory: 220Gi
            limits:
              nvidia.com/gpu: "2"
              cpu: "32"
              memory: 240Gi
              ephemeral-storage: 50Gi
          startupProbe:
            httpGet: { path: /health, port: 8081 }
            failureThreshold: 360
            periodSeconds: 10
          readinessProbe:
            httpGet: { path: /health, port: 8081 }
            failureThreshold: 30
            periodSeconds: 10
          livenessProbe:
            httpGet: { path: /live, port: 8081 }
            failureThreshold: 10
            periodSeconds: 30
# Frontend first
kubectl apply -f "${YAMLS}/dgd-split-hostpath.yaml" --dry-run=client -o yaml >/dev/null  # optional validate
kubectl apply -f <(sed -n '1,/^---$/p' "${YAMLS}/dgd-split-hostpath.yaml" | sed '/^---$/d')
# or apply the full multi-doc file once both are ready to land together:
kubectl apply -f "${YAMLS}/dgd-split-hostpath.yaml"
kubectl get pods -n dynamo-system -w

Then use the same Port-forward and smoke request steps as Path A (Service name is still typically dynamo-vllm-frontend for the unbundled frontend DGD).


Path C / D — CSI-backed KV (bundled or unbundled)

  1. Finish Dynamo - VAST CSI (Optional) (static or dynamic) until kubectl get pvc shows Bound.

  2. Copy the hostPath DGD of your chosen layout (dgd-combined-hostpath.yaml or dgd-split-hostpath.yaml) to dgd-combined-csi.yaml / dgd-split-csi.yaml.

  3. Replace only the worker kv-fs volume with a claim of that PVC. claimName must match metadata.name in the CSI guide’s pvc.yaml (see Values you must set before apply; example after kubectl get pvc -A: vast-kvcache-pvc):

- name: kv-fs
          persistentVolumeClaim:
            claimName: <NAME-FROM-CSI-GUIDE-pvc.yaml>   # e.g. vast-kvcache-pvc

Keep the same volumeMounts (mountPath: /kvcache). Model weights stay on hostPath. Still set nodeName: <GPU-NODE-NAME> as in Path A/B.

kubectl apply -f "${YAMLS}/dgd-combined-csi.yaml"
# or
kubectl apply -f "${YAMLS}/dgd-split-csi.yaml"
kubectl get pods -n dynamo-system -w

Then use the same Port-forward and smoke request steps as Path A.


Operational notes

  • Pod sharedMemorySize: 200Gi (v1beta1) must fit the OffloadingConnector CPU tier (160 GiB) plus headroom; size it for your cpu_bytes_to_use. On v1alpha1 the same field is sharedMemory.size.

  • Scout FP8 checkpoints sometimes omit chat_template in tokenizer_config.json. If the Dynamo frontend refuses to register the model, add a Llama-4 chat template to the local model directory before starting the worker.

  • Confirm rank-local pin lines in the worker logs after Ready.

  • Clean up: kubectl delete -f "${YAMLS}/<your-dgd>.yaml" (CSI: deleting the PVC is covered in the CSI guide’s reclaim behavior).

Additional Resources

Final Note

As workloads and software stacks frequently change, performance opportunities and compatibility shifts may occur. VAST actively expands its capabilities in the KV cache space — contact our team directly to leverage our latest updates and maximize your performance.