Upgrading a Cluster That Uses NVMe-oF Storage
If any DSX node in the cluster uses NVMe-oF storage, follow this guidance in addition to the software update procedure in the Hammerspace Installation and Licensing Guide.
Before You Upgrade
-
Make sure that every NVMe-oF path is healthy. If the cluster runs 5.2.14 or later, list the NVMe-oF path events that have not been cleared. In the Management GUI, open Events, select Uncleared events only, and look in the Action column for NVMe-oF path degraded and NVMe-oF path lost. From the Admin CLI, run
event-list --type NVMEOF_PATH_DEGRADED --uncleared, and then the same command with--type NVMEOF_PATH_LOST. Do not start the upgrade while either event is active; restore the path first. Earlier builds do not raise these events, so ask Hammerspace Support to check path health before you upgrade.Do not use the State shown by
nvmeof-config --list --fullfor this check. It showsCONNECTEDfor a path that is down. See NVMe-oF Path Health and Recovery. -
Upgrade to 5.3.1 or later. These builds include fixes for clusters that use NVMe-oF storage.
-
Plan to have out-of-band access to every node during the upgrade, for example through the server’s BMC or through vCenter for a virtual machine. A DSX node that uses NVMe-oF storage can stop responding while it restarts at the end of the upgrade, and then it needs a reset from outside the operating system.
-
Do not change the NVMe-oF configuration while the upgrade is running. That includes adding or removing enclosures, connecting or disconnecting subsystems, and changing connection settings. Some nodes run the new release and some the old one until the upgrade finishes.
After You Upgrade
-
Check that every volume on NVMe-oF storage is online, and that no
NVMEOF_PATH_DEGRADEDorNVMEOF_PATH_LOSTevent is active for a path you expect to work. -
If a volume on NVMe-oF storage is missing or offline, contact Hammerspace Support.
Do not re-create the logical volume, and do not use Force Initialize in the Configure Logical Volumes wizard or logical-volume-create --force. Both format the device and erase the data on it. The data on the array is usually intact, and Hammerspace Support can restore the volume without reformatting it. -
If you upgraded from 5.2, clear any NVMe-oF path event that was raised before the upgrade. Such an event does not clear automatically after the upgrade, because 5.2 attributed it to the DSX node and this release attributes it to the NVMe-oF subsystem.
-
First make sure the paths are healthy (step 1).
-
List the path events that have not been cleared:
Command:
event-list --type NVMEOF_PATH_LOST --unclearedRepeat the command with
--type NVMEOF_PATH_DEGRADED. -
For each event whose source is a DSX node, not an NVMe-oF subsystem, clear it by its ID:
Command:
event-update --clear --id <event-id>In the listing, an event raised before the upgrade shows
Source type: Nodeand the DSX node’s name as its source; an event raised after the upgrade showsSource type: Nvme-of Subsystemand the subsystem’s NQN.
-
-
Check the controller loss timeout of each enclosure. Enclosures keep the value they had before the upgrade. See Controller Loss Timeout.
The upgrade also turns on NFS DIRECT and NFSD DIRECT on every DSX node. See Connecting Storage to a DSX Using NVMe-oF.