NVMe/TCP and LACP host guidance

Prev Next

Intro

This guide explains how to implement native NVMe multipath and a host LACP bond on Linux (RHEL / Rocky) hosts using VAST Block over NVMe/TCP. Native is the standard build: two unbonded storage NICs, in-kernel multipath, and nvme connect-all so each VIP is its own TCP path. A host LACP bond is the exception, for LAN or management, or when the site will not unbond the NVMe NICs.

Architecture

Linux (RHEL / Rocky) hosts using VAST Block over NVMe/TCP.

Linux (RHEL / Rocky) hosts using VAST Block over NVMe/TCP

Host setup

NICs

Two NICs, not bonded. Prefer two VLANs, two subnets, two switches. One subnet per NIC. MTU 9000 end-to-end. No default route on the storage NICs.

<NVME_IF_A> and <NVME_IF_B> are the two NVMe NICs on the host, for example ens2f0np0. <VIP> is any address in the VAST data VIP pool.

ip -br addr
ip -br link
ip route
ip -d link show <NVME_IF_A> | grep mtu
ip -d link show <NVME_IF_B> | grep mtu
cat /sys/class/net/<NVME_IF_A>/master/uevent 2>/dev/null || echo no bond
cat /sys/class/net/<NVME_IF_B>/master/uevent 2>/dev/null || echo no bond
ping -c 2 -I <NVME_IF_A> <VIP>
ping -c 2 -I <NVME_IF_B> <VIP>
nc -vz <VIP> 4420

Kernel multipath

In-kernel NVMe multipath is usually enabled by default in RHEL 9. Persist retries:

cat /sys/module/nvme_core/parameters/multipath
cat /sys/module/nvme_core/parameters/max_retries

sudo tee /etc/modprobe.d/nvme_core.conf >/dev/null <<'EOF'
options nvme_core multipath=Y max_retries=5
EOF
sudo dracut -f

If parameters/multipath is not Y, the modprobe file above sets it. Rebuild initramfs and reboot if the running values are not already Y and 5.

sudo yum install -y nvme-cli
sudo modprobe nvme nvme-fabrics nvme-tcp
sudo tee /etc/modules-load.d/nvme.conf >/dev/null <<'EOF'
nvme
nvme-fabrics
nvme-tcp
EOF

Discover and connect

Register the host in VAST (Element Store > Block > Hosts) with:

cat /etc/nvme/hostnqn

Map volumes to that NQN, then:

sudo nvme discover -t tcp -a <VIP> -s 4420

sudo tee /etc/nvme/discovery.conf >/dev/null <<EOF
--transport=tcp --traddr=<VIP> --trsvcid=4420 --host-iface=<NVME_IF_A>
--transport=tcp --traddr=<VIP> --trsvcid=4420 --host-iface=<NVME_IF_B>
EOF

sudo nvme connect-all -t tcp -a <VIP> -s 4420 --host-iface=<NVME_IF_A>
sudo nvme connect-all -t tcp -a <VIP> -s 4420 --host-iface=<NVME_IF_B>

After mapping changes:

sudo nvme connect-all
sudo systemctl enable --now nvmf-autoconnect.service

I/O policy

I/O policy should be round-robin. Model string is VastData on 5.3.0-5.3.2 and VASTData on 5.3.2 and later:

sudo tee /lib/udev/rules.d/71-nvmf-vastdata.rules >/dev/null <<'EOF'
ACTION=="add|change", SUBSYSTEM=="nvme-subsystem", ATTR{subsystype}=="nvm", ATTR{model}=="VASTData", RUN+="/bin/sh -c 'echo round-robin > /sys/class/nvme-subsystem/%k/iopolicy'"
EOF
sudo udevadm control --reload-rules
sudo udevadm trigger

Host LACP bond

Use this for LAN and management, or only if the site will not give you two unbonded NICs for NVMe. Build the matching MLAG bond on both switches first (Switch configuration below). Then build the host bond. Only one side will drop the VLAN.

Settings that have to match the switch:

  • mode 802.3ad (LACP)

  • lacp_rate=fast if the switch is lacp-rate fast (slow if the switch is slow)

  • miimon=100

  • xmit_hash_policy=layer3+4 so different TCP flows can use both members

  • MTU 9000 on the bond and on both slaves (switch 9216 / Arista 9214)

<LAN_IF_A> and <LAN_IF_B> are the two physical NICs in the bond, for example ens2f0np0. They must go to different switches in the MLAG pair.

sudo modprobe bonding
echo bonding | sudo tee /etc/modules-load.d/bonding.conf

sudo nmcli connection delete bond0 bond0-p0 bond0-p1 2>/dev/null

sudo nmcli connection add type bond con-name bond0 ifname bond0 \
  bond.options "mode=802.3ad,lacp_rate=fast,miimon=100,xmit_hash_policy=layer3+4" \
  802-3-ethernet.mtu 9000 \
  ipv4.method auto ipv4.never-default yes ipv6.method ignore \
  connection.autoconnect no

sudo nmcli connection add type ethernet con-name bond0-p0 ifname <LAN_IF_A> master bond0 \
  802-3-ethernet.mtu 9000 connection.autoconnect no

sudo nmcli connection add type ethernet con-name bond0-p1 ifname <LAN_IF_B> master bond0 \
  802-3-ethernet.mtu 9000 connection.autoconnect no

If DHCP is tied to the old NIC MAC, clone it onto the bond so the reservation still hits:

sudo nmcli connection modify bond0 802-3-ethernet.cloned-mac-address <OLD_NIC_MAC>

Turn off the old standalone profiles so they do not steal the NICs back, then bring the bond up:

sudo nmcli -g NAME,DEVICE,TYPE connection show
sudo nmcli connection modify <OLD_IF_A_PROFILE> connection.autoconnect no
sudo nmcli connection modify <OLD_IF_B_PROFILE> connection.autoconnect no
sudo nmcli connection down <OLD_IF_A_PROFILE>
sudo nmcli connection down <OLD_IF_B_PROFILE>

sudo nmcli connection up bond0
sudo nmcli connection up bond0-p0
sudo nmcli connection up bond0-p1

If a VLAN rides the bond (example VLAN 69, keep the old device name if scripts depend on it):

sudo nmcli connection add type vlan con-name vlan69 ifname <LAN_IF_A>.69 \
  vlan.parent bond0 vlan.id 69 802-3-ethernet.mtu 9000 \
  ipv4.method auto ipv4.never-default yes
sudo nmcli connection up vlan69

Or move an existing VLAN profile onto the bond:

sudo nmcli connection modify <VLAN_PROFILE> vlan.parent bond0
sudo nmcli connection up <VLAN_PROFILE>

Persist after it is up:

sudo nmcli connection modify bond0 connection.autoconnect yes
sudo nmcli connection modify bond0-p0 connection.autoconnect yes
sudo nmcli connection modify bond0-p1 connection.autoconnect yes

Same thing as ifcfg files if the host still uses /etc/sysconfig/network-scripts:

# /etc/sysconfig/network-scripts/ifcfg-bond0
DEVICE=bond0
NAME=bond0
TYPE=Bond
BONDING_MASTER=yes
BONDING_OPTS="mode=802.3ad lacp_rate=fast miimon=100 xmit_hash_policy=layer3+4"
ONBOOT=yes
BOOTPROTO=dhcp
DEFROUTE=no
MTU=9000

# ifcfg-bond0-p0  /  ifcfg-bond0-p1
DEVICE=<LAN_IF_A>
TYPE=Ethernet
MASTER=bond0
SLAVE=yes
ONBOOT=yes
MTU=9000

Verify the host bond

cat /proc/net/bonding/bond0
ip -br link
ip -br addr show bond0
nmcli connection show --active

Look for:

  • Bonding Mode: IEEE 802.3ad Dynamic link aggregation

  • MII Status: up on the bond and both slaves

  • Number of ports: 2 in the aggregator

  • Partner Mac Address set (that is the switch / MLAG system MAC)

  • both member NICs SLAVE, bond MASTER

If Number of ports stays 1, the far side is not LACP, or only one switch has the MLAG bond.

If NVMe stays on this bond, still run nvme connect-all. Paths will show host_iface=bond0. Each TCP flow still hashes to one member. That is a deviation, not the standard build.

Switch configuration

LACP on an MLAG pair (Cisco: vPC). Same bond / port-channel id, LACP active, same VLAN list on both peers. Stage both, then apply both. Only one side will drop the VLAN. Switch MTU 9216 (Arista 9214).

Cumulus NVUE

Same commands on both peers. Replace <BOND_ID> / <SWP> / <VLAN_LIST>.

nv set interface bond<BOND_ID> bond member <SWP>
nv set interface bond<BOND_ID> bond mode lacp
nv set interface bond<BOND_ID> bond lacp-rate fast
nv set interface bond<BOND_ID> bond mlag id <BOND_ID>
nv set interface bond<BOND_ID> link mtu 9216
nv set interface bond<BOND_ID> link state up
nv set interface bond<BOND_ID> bridge domain br_default vlan <VLAN_LIST>
nv set interface bond<BOND_ID> bridge domain br_default stp admin-edge on
nv unset interface <SWP> bridge
nv config apply
nv config save
nv show mlag
nv show interface bond<BOND_ID> bond
nv show interface bond<BOND_ID>

You want MLAG status dual on both peers, LACP formed, member up.

Cisco NX-OS (vPC)

Same port-channel and vPC id on both peers.

interface Ethernet<PORT>
  mtu 9216
  channel-group <ID> mode active
  no shutdown

interface port-channel<ID>
  switchport
  switchport mode trunk
  switchport trunk allowed vlan <VLAN_LIST>
  mtu 9216
  spanning-tree port type edge trunk
  vpc <ID>
  no shutdown
show vpc
show port-channel summary
show lacp neighbor

Arista EOS

interface Ethernet<PORT>
   channel-group <ID> mode active

interface Port-Channel<ID>
   switchport mode trunk
   switchport trunk allowed vlan <VLAN_LIST>
   mtu 9214
   mlag <ID>
   spanning-tree portfast
show mlag
show lacp neighbor
show port-channel 19

Verify

sudo nvme list
sudo nvme list-subsys
sudo nvme list -v
cat /sys/class/nvme-subsystem/nvme-subsys*/iopolicy
cat /sys/module/nvme_core/parameters/multipath
cat /sys/module/nvme_core/parameters/max_retries

Expect parameters/multipath = Y, max_retries=5, iopolicy=round-robin, one tcp line per VIP all live, and host_iface on the two NICs (not bond0).

Do not let device-mapper own the NVMe devices.

Block process on the cluster

  1. Create a VIP pool with role Protocols, on the storage VLAN the host can reach. The number of addresses in the pool equals the path count VAST returns on discovery. You do not set port 4420 on the pool. The cluster opens NVMe on those VIPs once a Block view uses the pool.

  2. View policy: same tenant; attach that VIP pool. Other policy knobs (NFS flavors and so on) do not drive Block.

  3. Create a view. Path must be new and empty. Protocol Block only. Subsystem name must be unique in the tenant; no spaces. That name is baked into the subsystem NQN. Copy the NQN from the view after creation.

  4. Create a host. The name is cosmetic. NQN must match the initiator, byte for byte. Read it with:

    cat /etc/nvme/hostnqn
  5. Create a volume on that view. Size is the namespace the OS will see as:

    /dev/nvmeXnY
  6. Map the volume to the host. Until this mapping exists, this command returns 0 records even if 8009 and 4420 are open:

    nvme discover
  7. On the host, discover first. Discovery is port 8009. I/O is port 4420. You want one log entry per VIP, trsvcid 4420, and the subsystem NQN. Ping, NFS, or S3 on another VIP pool does not prove Block.

    nvme discover -t tcp -a <VIP> -s 8009

    Then connect once per storage NIC (not on a bond, for the standard build):

    nvme connect-all

What this looks like in VMS

1. VIP pools

Network Access, Virtual IP Pools. Role must be Protocols. A Block view uses one of these pools. Creating a Protocols pool does not open NVMe by itself. The Block view has to use it.

View network access

2. View policy

Element Store, View Policies. For Block, the knobs that matter are tenant, security flavor Block, and which VIP pools are attached. Open the Virtual IP pools picker and select the Block pool. If the policy still points at the NFS pool, discovery can succeed on the wrong addresses or not stick to the Block range. The default for Block views is fine if that policy has the right pool.

Element store

3. Block view

Element Store, Views. Path is a new empty directory. Protocol Block only. The policy is the Block policy with the Block VIP pool. Subsystem name is unique, no spaces. It is baked into the NQN. Copy Subsystem NQN for nvme connect. This screen is the subsystem. It is not the volume and not the host mapping. Without a mapped volume, discovery still returns 0 records.

Element store ->View

Data Flow with NVMe/TCP on a host LACP bond

Data Flow with NVMe/TCP on a host LACP bond

Data Flow with native NVMe multipath

Data Flow with native NVMe multipath

Appendix A

 

Recommended: native NVMe multipath

Exception: NVMe/TCP on a host LACP bond

Host interfaces

Two physical NICs, or a VLAN on each. Not bonded.

bond0 802.3ad, two members

Switch ports facing the host

Access or trunk, no LAG or MLAG, edge / portfast, MTU 9216

MLAG or vPC bond, same id on both peers, LACP active

host_iface in nvme list-subsys

The two NICs

bond0

What balances the load

The NVMe driver, iopolicy=round-robin

The LACP hash, xmit_hash_policy

One TCP session to one VIP

Independent path; the other NIC still carries its own sessions

Pinned to one bond member

Many VIPs via connect-all

Each VIP is its own path, spread over both NICs

Each VIP is still a single flow, though different VIPs can land on different members

Failover

A path drops; the remaining NVMe paths carry the I/O, helped by max_retries=5

The bond moves the member, but there is still one TCP path per VIP

MTU

9000 on the host, 9216 on the switch (9214 Arista)

Same

Cluster uplinks

MLAG with LACP active, unchanged

Same

Does it work

Yes, this is the design

Yes, but not what we size or support as standard

Use it when

Always, when the host can have two storage NICs

LAN and mgmt on a bond, or the site will not unbond the NVMe NICs