Search the docs

Setup

Automation

As you read through these steps and consider the number of nodes involved, it will become obvious that these procedures are repetitive and potentially error-prone. Automation should be utilized to set up, configure, and manage.

Reference the Hammerspace Ansible Tier 0 project to fully automate the setup of RAID arrays, filesystems, NFS exports, and firewall configuration for Hammerspace Tier 0 and Linux Storage Server (LSS) deployments. Refer to the project’s README.md file for full details.

System BIOS, Firmware, NIC Firmware

The following items need to be considered and prepared on the server and peripheral hardware:

  • System BIOS

    • Nodes Per Socket (NPS) - Hammerspace recommends that Nodes Per Socket be set to 1 on AMD platforms. On Intel, this is called Sub-NUMA Clustering (SNC).

    • Hyperthreading - Simultaneous Multithreading (SMT) can depend on server function, but the general recommendation is to set it to off in the system BIOS since it can introduce latency variance and reduce sustained memory copy performance, impacting applications sensitive to predictable I/O.

  • Memory architecture and optimization

    • DIMM types and speeds should match the memory controller of the CPU.

    • Populate all banks appropriately to maintain maximum memory speed. Review CPU specifications for number of memory channels and impact of number of DIMMs per channel (DPC).

  • Drive firmware

  • NIC firmware

  • NIC Driver

  • NVME configuration

    • Optimal slot / bay usage considering PCIe lanes, NUMA node. Some of this is dictated by the server architecture.

    • A single namespace per drive is sufficient. With RAID, the recommendation is a single namespace for each NVME device.

    • Select an NVME format that does not use LBA Metadata space (PI)

    • Sector size (512 vs 4096 byte). 4096 byte sector size is strongly recommended for Tier 0 and LSS nodes. Hammerspace node boot devices currently require 512 bytes per sector.

  • Device drivers and settings especially for NICs

    • There may be different device drivers for advanced networking such as RDMA.

Drives, RAID, Filesystems

Check how NVME drives enumerate (what device name they appear as). Are they consistent between reboots or other events? On some systems, NVME devices may not enumerate as the same /dev/nvmexxx device on reboot. To prevent devices being mounted on the wrong mount point, UUID-based mounts should be used.

NVME drives often support different formats that may include metadata space for each sector (NVME metadata space, not Hammerspace). On some systems, this might be used for additional checksums. With Tier 0 deployments, NVME drives should be configured and formatted with 0 bytes of metadata space.

Determine whether RAID or individual drives will be used. One consideration is how many volumes will need to be managed, which is a function of how many servers times how many drives they have. If the number of volumes is projected to be over 1000, RAID configuration should be favored.

Tips to Improve RAID Performance

  • Use RAID set sizes that are a power of 2 where possible

  • Attempt to build RAID sets on NUMA local NVMe

  • The default mdadm chunk size of 512KiB will suffice for many deployments. For advanced tuning, consult with your Hammerspace tech team to calculate a chunk size matched to I/O profile.

Determine NUMA placement of NVME devices as follows:

# grep "" /sys/block/nvme*n1/device/numa_node

/sys/block/nvme0n1/device/numa_node:0

/sys/block/nvme10n1/device/numa_node:1

/sys/block/nvme11n1/device/numa_node:1

/sys/block/nvme1n1/device/numa_node:0

/sys/block/nvme2n1/device/numa_node:0

/sys/block/nvme3n1/device/numa_node:0

/sys/block/nvme4n1/device/numa_node:0

/sys/block/nvme5n1/device/numa_node:0

/sys/block/nvme6n1/device/numa_node:1

/sys/block/nvme7n1/device/numa_node:1

/sys/block/nvme8n1/device/numa_node:1

/sys/block/nvme9n1/device/numa_node:1

Pre-Setup Checks

Once the base server hardware is set up, operating system installed, and basic networking is configured, you can check the recommendations listed above.

If jumbo frames are in use, check the configuration with ping using a large frame size and the Do Not Fragment flag set. Note that the frame size of the ping is smaller than the MTU because of packet overhead. The most common MTU, 9000, should be checked with a ping size of 8972. You should check node (Tier 0 and LSS) to node, node to Hammerspace nodes, and other combinations including DI, etc. Every node should be checked.

# ping -M do -s 8972 192.168.0.101

PING 192.168.0.101 (192.168.0.101) 8972(9000) bytes of data.

8980 bytes from 192.168.0.101: icmp_seq=1 ttl=64 time=0.013 ms

8980 bytes from 192.168.0.101: icmp_seq=2 ttl=64 time=0.004 ms

8980 bytes from 192.168.0.101: icmp_seq=3 ttl=64 time=0.004 ms

^C

--- 192.168.0.101 ping statistics ---

3 packets transmitted, 3 received, 0% packet loss, time 2038ms

rtt min/avg/max/mdev = 0.004/0.007/0.013/0.004 ms

Listing Storage Devices and Current Usage:

# lsblk -n

NAME MAJ:MIN RM SIZE RO TYPE MOUNTPOINTS

nvme11n1 259:0 0 894.3G 0 disk

nvme10n1 259:1 0 745.1G 0 disk

├─nvme10n1p1 259:2 0 600M 0 part /boot/efi

├─nvme10n1p2 259:3 0 1G 0 part /boot

├─nvme10n1p3 259:4 0 32G 0 part [SWAP]

└─nvme10n1p4 259:5 0 711.5G 0 part /

nvme0n1 259:6 0 7T 0 disk

nvme1n1 259:7 0 7T 0 disk

nvme2n1 259:8 0 7T 0 disk

nvme3n1 259:9 0 7T 0 disk

nvme4n1 259:10 0 7T 0 disk

nvme5n1 259:11 0 7T 0 disk

nvme6n1 259:12 0 7T 0 disk

nvme7n1 259:13 0 7T 0 disk

nvme8n1 259:14 0 7T 0 disk

nvme9n1 259:15 0 7T 0 disk

nvme12n1 259:16 0 7T 0 disk

nvme13n1 259:17 0 7T 0 disk

nvme13n1 259:18 0 7T 0 disk

nvme15n1 259:19 0 7T 0 disk

nvme16n1 259:20 0 7T 0 disk

Configuring the Server: RAID and Filesystems

This shows an example RAID configuration of 3 RAID arrays of 8 drives each. The last two commands ensure that the RAID configuration is saved in the mdadm configuration file.

# mdadm --create --verbose /dev/md0 --level=0 --raid-devices=8 /dev/nvme0n1 /dev/nvme1n1 /dev/nvme2n1 /dev/nvme3n1 /dev/nvme4n1 /dev/nvme5n1 /dev/nvme6n1 /dev/nvme7n1
# mdadm --create --verbose /dev/md1 --level=0 --raid-devices=8 /dev/nvme9n1 /dev/nvme10n1 /dev/nvme11n1 /dev/nvme12n1 /dev/nvme13n1 /dev/nvme14n1 /dev/nvme15n1 /dev/nvme16n1
# mdadm --create --verbose /dev/md2 --level=0 --raid-devices=8 /dev/nvme17n1 /dev/nvme18n1 /dev/nvme19n1 /dev/nvme20n1 /dev/nvme21n1 /dev/nvme22n1 /dev/nvme23n1 /dev/nvme24n1
# mkdir /etc/mdadm
# mdadm --detail --scan | tee -a /etc/mdadm/mdadm.conf

Set up mount points, partition arrays, make a file system, mount, and add to fstab:

# mkdir /hsvol0
# parted /dev/md0 mklabel gpt
# parted -a opt /dev/md0 mkpart primary xfs 0% 100%
# mkfs.xfs -d agcount=512 -L hsvol0 /dev/md0p1
# mount /dev/md0p1 /hsvol0
# echo '/dev/md0p1 /hsvol0 xfs 0 0' >> /etc/fstab

After editing /etc/fstab, you should run 'systemctl daemon-reload' to update systemd units generated from the file.

NFS, Exports, /etc/exports

If the NFS server (nfs-kernel-server) is already installed, check /etc/nfs.conf and /etc/nfs.conf.d/local.conf for non-default TCP/UDP port configuration. Hammerspace recommends the following /etc/nfs.conf settings for use in Tier 0 environments:

Example /etc/nfs.conf
[nfsd]
threads=128
# udp=n
# tcp=y
vers3=y
# vers4=y
vers4.0=n
vers4.1=n
vers4.2=y
rdma=y
rdma-port=20049

NFS v4.2 is required for Parallel NFS, client side mirroring and other advanced features used with Tier 0 implementations. Although clients mount the Hammerspace share(s) via NFS v4.2, the I/O from clients to storage volumes uses NFS v3, so NFS servers including Tier 0 nodes must enable NFS v3.

Run the following command to start NFS and configure it to start on system startup:

# systemctl enable --now nfs-server.service
On some Linux distributions, the NFS service might be nfs-kernel-server, although that is usually linked to nfs-server. Ensure showmount is not disabled (uncommon).

For /etc/exports:

  • Add Hammerspace nodes and Anvil cluster mgmt floating IP RW, no_root_squash,sync,secure,mp,no_subtree_check.

  • Add clients RW,root_squash,sync,secure,mp,no_subtree_check. This can be done by subnet or host

Export Options

The following export options are recommended for nodes that serve NFS, including Tier 0 nodes and LSS:

  • rw - Read/write access

  • no_root_squash - Allows appropriate nodes to have root access on the export. This setting is required for Hammerspace nodes (Anvil and DSX), and nodes running mover services (DI and CM).

  • root_squash - Prohibits nodes from root access on the export. This should be used for Tier 0 and LSS nodes that are not running mover services.

  • sync - Causes the NFS server to ensure writes are committed to storage before replying to the client. This is default behavior in modern NFS server implementations, but specifying the option may suppress warning messages.

  • secure - Requires that clients use a secure, well-known TCP port (less than 1024).

  • mp - Mountpoint indicates that the export should have a filesystem mounted for export. Using this option prevents accidentally exporting the directory under the mounted filesystem, which would result in data being written to the root filesystem.

  • no_subtree_check - This is a default in modern NFS server implementations. The main reason to specify this parameter is to suppress a warning that the default behavior has changed.

The following snippet of /etc/exports shows 3 exported volumes with 3 exports for a Hammerspace Anvil node and 3 exports for a subnet of Tier 0.

Example /etc/exports
/hsvol0 10.1.2.3(rw,no_root_squash,sync,secure,mp,no_subtree_check)
/hsvol1 10.1.2.3(rw,no_root_squash,sync,secure,mp,no_subtree_check)
/hsvol2 10.1.2.3(rw,no_root_squash,sync,secure,mp,no_subtree_check)
/hsvol0 10.2.1.0/24(rw,root_squash,sync,secure,mp,no_subtree_check)
/hsvol1 10.2.1.0/24(rw,root_squash,sync,secure,mp,no_subtree_check)
/hsvol2 10.2.1.0/24(rw,root_squash,sync,secure,mp,no_subtree_check)

The challenge with the above syntax is that although you can put multiple hosts on each line, you have to specify export options in parentheses for each host which gets messy and error prone.

Many Linux distributions allow multiple IPs on a single line with their common export options first:

Example /etc/exports
/mnt/hsvol1 -rw,no_root_squash,sync,secure,mp,no_subtree_check 10.200.104.38 10.200.105.195
/mnt/hsvol1 -rw,root_squash,sync,secure,mp,no_subtree_check 10.200.102.33 10.200.103.223 10.200.101.217 10.200.102.92 10.200.102.34 10.200.103.170 10.200.102.172 10.200.103.173

After updating /etc/exports, run the following command to re-read the exports file:

# exportfs -r

Firewall

You may need to allow ports used by NFS and related protocols and services. Check /etc/nfs.conf and the files in /etc/nfs.conf.d for port specifications, and also rpcinfo -p to confirm. We have seen sites that had other applications on GPU nodes configured to use ports normally used by NFS.

For Rocky, Centos, RHEL using standard ports:

# firewall-cmd --permanent --add-service=nfs --add-service=rpc-bind --add-service=mountd ; firewall-cmd --reload

For Ubuntu and Debian using standard ports:

# ufw allow port 111 # portmapper
# ufw allow port 2049 # NFS
# ufw allow port 20048 # mountd

Confirm it works by checking exports from another client (can be another LSS/Tier 0 server):

# showmount -e 10.200.103.167

Export list for 10.200.103.167:

/mnt/hsvol1 10.200.112.0/21,10.200.10.0/24,10.200.103.166,10.200.103.167,10.200.102.220,10.200.101.3,10.200.106.197

/mnt/hsvol0 10.200.112.0/21,10.200.10.0/24,10.200.103.166,10.200.103.167,10.200.102.220,10.200.101.3,10.200.106.197