Search the docs

Troubleshoot TLS Issues

Hammerspace generates the following events when TLS configuration problems are detected. These events are visible in the Management GUI under Events and through any configured alert channels (SNMP, email, syslog).

Table 1. TLS events

Event

Description

NFS_TLS_CONFIGURATION_UNHEALTHY

Raised per node. The node’s applied TLS policy — or `pd-protod’s policy on the primary Anvil — does not match the cluster policy and the automatic re-apply did not correct it, or the check could not run.

PKI_CA_CREATION_FAILED

The Hammerspace system CA could not be created during initialization.

PKI_TRUST_SYNC_FAILED

CA trust material failed to distribute to one or more cluster nodes.

PKI_NODE_CERT_CREATION_FAILED

A per-node X.509 certificate could not be issued.

PKI_NODE_CERT_SYNC_FAILED

A per-node certificate could not be distributed to its target node.

PKI_NODE_CHECK_FAILED

The periodic health check of the node’s certificate failed. The check covers the CA match, the certificate chain, expiry, and whether the key matches the certificate. It is raised at Critical severity while TLS is enabled and at Warning severity while TLS is off. The event’s details name the check that failed. Hammerspace Support can run hs-nfs-pki-check on the node for more detail, and its output is captured in support bundles.

While TLS is being enabled, each DSX node raises HW_FRU_EVENT notices such as Added hardware device /dev/sda as it re-announces its own disks. These notices are expected and need no action.

Client Cannot Mount After TLS Is Enabled

If a client cannot mount an NFS share after TLS has been enabled on the cluster, verify the following:

  1. Mounts made without TLS before the change are unmounted. A plaintext mount that is still in place loses access when TLS is enabled, and the client kernel logs RPC: server <address> requires stronger authentication about once a second until you unmount it:

    sudo umount /mnt/your_mount_point
  2. The mount uses the xprtsec=mtls option:

    mount -o xprtsec=mtls <server>:/<export> /<mountpoint>
  3. The client mounts by an address that appears in the server’s certificate. Hammerspace-issued certificates carry IP addresses only — node, cluster, and portal floating IPs — plus the internal name data-cluster. Mounting by a customer DNS name fails the handshake.

  4. The client’s ktls-utils is a current release. With ktls-utils 0.9 (the version Ubuntu 24.04 LTS packages) the mount fails with gnutls: Error in the certificate. (-43) in the tlshd journal (preceded by Name or service not known) unless the server address reverse-resolves to a name in the server certificate.

  5. The tlshd.service is running on the client:

    systemctl status tlshd.service
  6. The Hammerspace root CA certificate (or your organization’s CA certificate) is installed in the client’s trust store. See Remount Clients for the trust anchor command.

  7. The client’s /etc/tlshd.conf names a certificate and a private key. The ktls-utils package ships a default configuration file, so a missing file is the less likely fault. The common fault is an [authenticate.client] section with no x509.certificate and no x509.private_key: the client then presents no certificate and the cluster rejects the handshake.

  8. The client certificate is valid and has not expired. Check whatever path x509.certificate under [authenticate.client] in the client’s /etc/tlshd.conf points at:

    openssl x509 -in <path-from-tlshd.conf> -noout -dates
    /etc/pki/hs-tls/nfs-pki/cert.pem is the certificate path on Hammerspace Anvil and DSX nodes. A client uses whatever path its own tlshd.conf specifies.
  9. Restart tlshd after editing /etc/tlshd.conf or replacing certificate files:

    systemctl restart tlshd.service
  10. The client meets the minimum Linux kernel and package version requirements. See Reference: Client Requirements.

  11. Hammerspace trusts the client’s certificate. A client certificate that Hammerspace does not trust fails the handshake even when every item above is correct: the mount reports mount.nfs: access denied by server although the client’s tlshd journal says Handshake with … was successful, because the server closes the connection after the handshake. A certificate issued by an intermediate CA needs the intermediate certificate either in the client’s certificate file or in the Hammerspace trust store. See Reference: Certificate Trust Model for which certificates each side must trust, and Remount Clients for configuring client certificate trust.

File I/O Hangs on a TLS Mount

Symptom: A client mounts a share with xprtsec=mtls and can list directories, but reads and writes never finish. Commands such as dd or cp hang, and dmesg shows no errors. Small writes that appear to succeed leave files of 0 bytes on the share.

Cause: The client’s pNFS data connections to the DSX nodes (port 3049) do not use TLS, so the DSX nodes refuse them and the client retries indefinitely. This happens on Enterprise Linux 10.1 (kernel 6.12.0-124).

To confirm, run these commands on the client while the I/O is hung:

sudo ss -tn state time-wait '( dport = :3049 )' | wc -l
grep -E 'pnfs=|LAYOUTGET' /proc/self/mountstats

Hundreds of TIME-WAIT connections to port 3049, pnfs=LAYOUT_FLEX_FILES, and a LAYOUTGET line with a large error count indicate this problem.

Resolution:

  1. Disable the pNFS flexible-files layout driver on the client, as described in the prerequisites of Remounting an Enterprise Linux 10.1 (or Later) Client.

  2. Reboot the client. The hung mounts keep the driver loaded, so unmounting alone does not clear the problem.

  3. Remount the share with xprtsec=mtls, and check that grep pnfs= /proc/self/mountstats shows pnfs=not configured.

File I/O then goes through the client’s TLS connection to the Anvil.

Backups Fail with "access denied by server"

If a backup mount fails after TLS is required on the cluster with the error:

mount.nfs: access denied by server while mounting <ip>:/backup, error code: 32

the TLS handshake failed. The Linux client reports every failed TLS handshake with this message, so the cause is not necessarily that the backup server has TLS disabled.

The handshake is mutual, so check both directions:

  • The backup server must trust the Hammerspace root certificate authority. The Anvil presents its node certificate to the backup server.

  • The cluster must trust the backup server’s certificate. That certificate must also carry the IP address configured for the backup target as a subject alternative name.

Diagnosing TLS Connectivity to a Storage Node

Hammerspace provides a diagnostic utility, volume_config_test, that tests TLS connectivity to a specific storage node. Run it on an Anvil or DSX node.

volume_config_test --mtls --ip <node_ip> --rootfh <root_filehandle> --node-uuid <node_uuid>
--mtls

Uses the client certificate from the /etc/tlshd.conf of the node you run the command on.

--node-uuid

The UUID of the node you run the command on, not the UUID of the storage node being tested. The tool uses it to name the test directory it creates on the volume.

--port

Required only for a DSX volume, which uses port 3049. Third-party NAS uses the default of 2049, so omit --port for third-party NAS.

--src-ip

Optional.

External Mover Refuses TLS-Required Mobilities

If mobilities to TLS-required storage are not running on a Data Instantiator, check the pd-di log for:

DI is TLS-incapable (PKI material unavailable)

This means pd-di did not find valid PKI material when it started, so it refuses mobilities to TLS-required storage. DME reschedules those mobilities to another DI.

pd-di checks its PKI files only at startup, so putting the files in place is not enough on its own — the mover must be restarted. See External Movers (DI and Cloud Mover Containers).

Local tlshd.conf Edits Lost After a Software Update

Every software update copies the shipped tlshd.conf over /etc/tlshd.conf on Anvil and DSX nodes. Any local edits to that file are lost.

This most often shows up as tlshd debug logging that stops working after an update. Reapply local changes to /etc/tlshd.conf after each software update, and restart tlshd on the node afterwards.

Spurious TLS-unhealthy Alert During a Software Update

You may see an NFS_TLS_CONFIGURATION_UNHEALTHY event fire during a cluster software update, with an accompanying error code 4057 ("prohibited during software update"), even if TLS was never configured on the cluster. This is a known false positive. It clears on its own once the update completes and does not indicate an actual TLS problem.