Search the docs

NVMe-oF Path Health and Recovery

Each DSX node monitors every path to each NVMe-oF subsystem that it is connected to. A path is one connection from the DSX node to one target address of the subsystem. A subsystem that the array presents on two target addresses has two paths to each DSX node that connects to it.

When a path stops working, the DSX node raises an event and keeps trying to reconnect the path. When the target answers again, the DSX node reconnects the path and clears the event. You do not need to reconnect the path yourself.

Path Health Events

Table 1. NVMe-oF path health events
Event Severity Meaning

NVMEOF_PATH_DEGRADED

Warning

A path has been trying to reconnect for a while, and the subsystem still has at least one working path. I/O continues on the working paths with reduced redundancy.

NVMEOF_PATH_LOST

Warning

The connection for a path no longer exists, and the subsystem still has at least one working path. I/O continues on the working paths with reduced redundancy. This happens, for example, when the enclosure’s controller loss timeout is a finite value and has expired, or when a path did not connect when the DSX node started.

NVMEOF_PATH_LOST

Critical

The subsystem has had no working path from this DSX node for a short time, whether its connections are still trying to reconnect or no longer exist. I/O to volumes on the subsystem’s namespaces waits until a path is restored. If the enclosure’s controller loss timeout is a finite value and it expires first, that I/O fails instead; see Controller Loss Timeout.

Each event names the path it is about, in this order: the subsystem NQN, the target address, the target port, the transport, and the host name of the DSX node. One event is raised for each path of each subsystem, so a subsystem with two paths on a target that goes down raises two events. For example:

NVMe-oF path lost: nqn=<subsystem-nqn>, traddr=<target-address>, trsvcid=<target-port>, transport=tcp, node=<dsx-host-name>

The source of both events is the NVMe-oF subsystem, identified by its NQN. In the Events list, the Source type column shows Nvme-of Subsystem. If Hammerspace cannot match the path to a subsystem that it manages, the source is the DSX node that lost the path.

Expected output of event-list --type NVMEOF_PATH_DEGRADED with one path of a two-path target down:

total 2
ID:                      3c5265f0-bad1-11f1-8660-005056adc30c
Created:                 2026-09-28 00:12:04 UTC
Modified:                2026-09-28 00:12:04 UTC
Type:                    NVMEOF_PATH_DEGRADED
Description:             NVMe-oF path degraded (controller stuck connecting): nqn=nqn.2026-09.com.example:docs-subsys2, traddr=192.0.2.32, trsvcid=8009, transport=tcp, node=dsx1.example.com
Severity:                WARNING
Source type:             Nvme-of Subsystem
Source ID:               7c4b48fa-ce56-5ff2-936f-e68170bb6966
Source name:             nqn.2026-09.com.example:docs-subsys2
Cleared:                 false

ID:                      3c7f2d10-bad1-11f1-8660-005056adc30c
Created:                 2026-09-28 00:12:05 UTC
Modified:                2026-09-28 00:12:05 UTC
Type:                    NVMEOF_PATH_DEGRADED
Description:             NVMe-oF path degraded (controller stuck connecting): nqn=nqn.2026-09.com.example:docs-subsys1, traddr=192.0.2.32, trsvcid=8009, transport=tcp, node=dsx1.example.com
Severity:                WARNING
Source type:             Nvme-of Subsystem
Source ID:               9d588e3f-d951-5d41-b918-d2e557bf8edc
Source name:             nqn.2026-09.com.example:docs-subsys1
Cleared:                 false

How the Events Behave

  • A brief interruption does not raise an event. The DSX node waits to see whether a path recovers on its own before it reports it.

  • When the subsystem’s last working path fails, the events change to Critical. A NVMEOF_PATH_LOST event at Warning rises to Critical in place. A NVMEOF_PATH_DEGRADED event is cleared and replaced by a NVMEOF_PATH_LOST event at Critical for the same path.

  • When any path to the subsystem works again, the events for the paths that are still down step back down to Warning: a NVMEOF_PATH_LOST event for a path whose connection no longer exists returns to Warning in place, and a Critical event for a path that is still trying to reconnect is cleared and replaced by a NVMEOF_PATH_DEGRADED event at Warning.

  • When the path works again, the event clears automatically, usually within ten seconds.

  • Disconnecting a subsystem or removing an enclosure through Hammerspace normally does not raise these events, and any event that is active for those paths is cleared.

  • The DSX node monitors only subsystems that were connected through Hammerspace, by using the Management GUI or the nvmeof-config command.

Responding to a Path Health Event

For a Warning event, check the network path, the DSX node’s network interface, and the array port for the target address named in the event. Data remains available through the subsystem’s other paths.

For a Critical event, treat the problem as urgent and restore connectivity between the DSX node and the array. As soon as any path to the subsystem answers again, the DSX node reconnects it and the event clears.

Do not disconnect and reconnect the subsystem to recover a path. Hammerspace reconnects the path automatically, and you cannot disconnect a subsystem whose namespaces have logical volumes.

SNMP Traps

If SNMP trap forwarding is configured, each event is also sent as an SNMP trap:

  • NvmeofPathDegraded, notification 128

  • NvmeofPathLost, notification 129

Each trap carries the event type, severity, and description, along with the event’s other standard fields.

Connection State Is Not Path Health

The nvmeof-config --list --full output shows a State for each target address of each subsystem:

DISCOVERED

The subsystem was found at this address, but Hammerspace has not connected it.

CONNECTED

Hammerspace connected the subsystem.

The State field records what Hammerspace was asked to do. It does not show whether a path is working right now. A path that has gone down still shows CONNECTED, and so does every address of a subsystem that is connected on any of them. The Connected value of each namespace and the State of each drive in drive-list are the same kind of record.

To find out whether paths are healthy, use the NVMe-oF path health events described at the start of this page.

Expected output of nvmeof-config --list --full with one subsystem connected and one only discovered:

total 2
Type:                    Nvme-of Enclosure
ID:                      9ebbfb00-033c-401b-bb9a-bb4a4ffb3e88
Name:                    dsx1.example.com::nvmeof-9ebbfb00-033c-401b-bb9a-bb4a4ffb3e88
Node:                    dsx1.example.com
Node host NQN:           nqn.2014-08.org.nvmexpress:uuid:0fb02d42-2c53-41ad-3fcf-1ad8256be53e
NVMe-oF config:          Queue depth:             64
                         IO queue pairs:          64
                         I/O policy:              QUEUE_DEPTH
                         Transport:               TCP
                         Target port:             8009
                         Controller loss timeout: -1
Discovery targets:
                         Address:                 192.0.2.32
                         Port:                    8009
                         Transport:               TCP

                         Address:                 192.0.2.31
                         Port:                    8009
                         Transport:               TCP
Subsystems:
                         Type:                    Nvme-of Subsystem
                         ID:                      9d588e3f-d951-5d41-b918-d2e557bf8edc
                         Subsystem NQN:           nqn.2026-09.com.example:docs-subsys1
                         Namespaces:
                                                  [Namespace: nvme0n1, Namespace ID: 1, Size: 4.2GB, Sector size: 4096, Model: Docs Lab NVMe Target, Serial: DOCSLAB00000001, Namespace unique ID: c9eedb06-110b-552b-beaa-1cd7e38ae394, Connected: true]
                                                  [Namespace: nvme0n2, Namespace ID: 2, Size: 4.2GB, Sector size: 4096, Model: Docs Lab NVMe Target, Serial: DOCSLAB00000001, Namespace unique ID: c9eedb06-110b-552b-beaa-1cd7e38ae394-2, Connected: true]
                         Transport addresses:
                                                  [Address: 192.0.2.32, Port: 8009, Transport: TCP, State: CONNECTED]
                                                  [Address: 192.0.2.31, Port: 8009, Transport: TCP, State: CONNECTED]

                         Type:                    Nvme-of Subsystem
                         ID:                      7c4b48fa-ce56-5ff2-936f-e68170bb6966
                         Subsystem NQN:           nqn.2026-09.com.example:docs-subsys2
                         Namespaces:
                                                  [Namespace ID: 1, Size: 4.2GB, Sector size: 4096, Model: Docs Lab NVMe Target, Serial: DOCSLAB00000002, Namespace unique ID: bef19a80-60dd-b2bd-30a3-4965b79d3e20, Connected: false]
                                                  [Namespace ID: 2, Size: 4.2GB, Sector size: 4096, Model: Docs Lab NVMe Target, Serial: DOCSLAB00000002, Namespace unique ID: bef19a80-60dd-b2bd-30a3-4965b79d3e20-2, Connected: false]
                         Transport addresses:
                                                  [Address: 192.0.2.31, Port: 8009, Transport: TCP, State: DISCOVERED]
                                                  [Address: 192.0.2.32, Port: 8009, Transport: TCP, State: DISCOVERED]

Type:                    Nvme-of Enclosure
ID:                      7d53462a-d5d2-4b53-9264-e9883b03a8dd
Name:                    dsx1.example.com::nvmeof-7d53462a-d5d2-4b53-9264-e9883b03a8dd
Node:                    dsx1.example.com
Node host NQN:           nqn.2014-08.org.nvmexpress:uuid:0fb02d42-2c53-41ad-3fcf-1ad8256be53e
NVMe-oF config:          Queue depth:             64
                         IO queue pairs:          64
                         I/O policy:              QUEUE_DEPTH
                         Transport:               TCP
                         Target port:             8009
                         Controller loss timeout: -1
Discovery targets:
                         Address:                 192.0.2.32
                         Port:                    8009
                         Transport:               TCP

                         Address:                 192.0.2.31
                         Port:                    8009
                         Transport:               TCP
Subsystems:
                         Type:                    Nvme-of Subsystem
                         ID:                      9d588e3f-d951-5d41-b918-d2e557bf8edc
                         Subsystem NQN:           nqn.2026-09.com.example:docs-subsys1
                         Namespaces:
                                                  [Namespace: nvme0n1, Namespace ID: 1, Size: 4.2GB, Sector size: 4096, Model: Docs Lab NVMe Target, Serial: DOCSLAB00000001, Namespace unique ID: c9eedb06-110b-552b-beaa-1cd7e38ae394, Connected: true]
                                                  [Namespace: nvme0n2, Namespace ID: 2, Size: 4.2GB, Sector size: 4096, Model: Docs Lab NVMe Target, Serial: DOCSLAB00000001, Namespace unique ID: c9eedb06-110b-552b-beaa-1cd7e38ae394-2, Connected: true]
                         Transport addresses:
                                                  [Address: 192.0.2.31, Port: 8009, Transport: TCP, State: CONNECTED]
                                                  [Address: 192.0.2.32, Port: 8009, Transport: TCP, State: CONNECTED]

                         Type:                    Nvme-of Subsystem
                         ID:                      7c4b48fa-ce56-5ff2-936f-e68170bb6966
                         Subsystem NQN:           nqn.2026-09.com.example:docs-subsys2
                         Namespaces:
                                                  [Namespace ID: 1, Size: 4.2GB, Sector size: 4096, Model: Docs Lab NVMe Target, Serial: DOCSLAB00000002, Namespace unique ID: bef19a80-60dd-b2bd-30a3-4965b79d3e20, Connected: false]
                                                  [Namespace ID: 2, Size: 4.2GB, Sector size: 4096, Model: Docs Lab NVMe Target, Serial: DOCSLAB00000002, Namespace unique ID: bef19a80-60dd-b2bd-30a3-4965b79d3e20-2, Connected: false]
                         Transport addresses:
                                                  [Address: 192.0.2.32, Port: 8009, Transport: TCP, State: DISCOVERED]
                                                  [Address: 192.0.2.31, Port: 8009, Transport: TCP, State: DISCOVERED]

Controller Loss Timeout

The controller loss timeout sets how long a DSX node keeps trying to reconnect a lost connection to an NVMe-oF controller before the node’s kernel gives up and removes the controller. It is one of the connection settings of an enclosure. An enclosure is Hammerspace’s record of one NVMe-oF array as a DSX node sees it: the array’s discovery addresses, the subsystems found there, and the settings used to connect to them. Each enclosure has its own value.

The value is either -1 or a number of seconds from 0 to 3600. The value -1 means that the DSX node never gives up trying to reconnect, and it is the recommended setting. Do not set 0: with 0, the kernel removes the controller as soon as a path fails.

Enclosures added on this release use -1 by default. Enclosures that were added before the cluster was upgraded keep the value they had, so check them after upgrading.

Hammerspace reconnects a lost path automatically when the target answers again, whatever the timeout is set to. The timeout still matters, because it decides what happens to the block device while the path is down:

  • With -1, the NVMe block device stays attached while the DSX node keeps trying to reconnect. If every path to a subsystem is down, I/O to volumes on that subsystem waits until a path returns, and then completes. The storage volume stays online.

  • With a finite value, the kernel removes the controller when the timeout expires. If that was the subsystem’s last working path, the storage volume on the affected namespace goes to the Suspected operational state, I/O that is waiting fails, and the file system on the volume stops serving data. When a path returns, Hammerspace reconnects it and remounts the file system; the volume returns to Up within a few minutes without any action from you.

Checking the Current Value

Command:

nvmeof-config --list --full

In the output, each enclosure’s NVMe-oF configuration shows its Controller loss timeout.

Expected output (one enclosure, trimmed to the configuration block):

total 2
Type:                    Nvme-of Enclosure
ID:                      9ebbfb00-033c-401b-bb9a-bb4a4ffb3e88
Name:                    dsx1.example.com::nvmeof-9ebbfb00-033c-401b-bb9a-bb4a4ffb3e88
Node:                    dsx1.example.com
Node host NQN:           nqn.2014-08.org.nvmexpress:uuid:0fb02d42-2c53-41ad-3fcf-1ad8256be53e
NVMe-oF config:          Queue depth:             64
                         IO queue pairs:          64
                         I/O policy:              QUEUE_DEPTH
                         Transport:               TCP
                         Target port:             8009
                         Controller loss timeout: -1
<Additional output removed for brevity>

Setting the Value

Set the timeout in the Management GUI. The Controller lost timeout field is on the NVMe-oF Config tab that opens when you click the pencil icon of the DSX node on the Storage Systems tab of the Infrastructure page. Enter the value and click Update Storage System. A value saved there applies to every NVMe-oF enclosure on that DSX node, not just one; the tab shows the values of the node’s first enclosure.

The Admin CLI accepts only a value from 0 to 3600 for --ctrl-loss-tmo; it rejects -1 with a syntax error. To set a finite timeout on one enclosure, replace <enclosure-name> with the enclosure’s name as shown by nvmeof-config --list. You can use --enclosure-id <enclosure-uuid> instead.

Command:

nvmeof-config --ctrl-loss-tmo <seconds> --enclosure-name <enclosure-name>

Once you have set a finite value from the Admin CLI, use the Management GUI to return to -1.

Changing the timeout, or any other connection setting except the I/O policy, disconnects and reconnects the controllers of the enclosure’s connected subsystems, because these settings cannot be changed on a live connection. Changing only the I/O policy does not reconnect anything.

Volumes stay mounted through the reconnect. Make the change in a maintenance window anyway.