Search the docs

Monitoring Cluster Health

The local drives on a product node are automatically monitored for space utilization. There is a WARNING event at 70% and an EXCEEDED event at 90%. This is done by a space monitoring service, which checks all the local filesystems (e.g. /, /boot, /var etc) every two minutes. This runs on All Anvil and DSX nodes.

It is worth noting that the Prometheus exporters will also collect this information and it can be viewed graphically with Grafana.

Here is an example of the event text:

Example of event text
____
ID: d0719f4c-b1ca-11ef-9bd4-30e171639c80

Created: 2024-12-03 23:03:31 UTC

Modified: 2024-12-03 23:03:31 UTC

Type: DISK_USAGE_LIMIT_EXCEEDED

Description: Disk usage limit EXCEEDED on doxie-dsx1.lab.hammer.space:/boot with only 106M available.

Severity: ALERT
____

About the Prometheus Exporters

Hammerspace ships integrated Prometheus exporters. They export data about the system, the resource utilization of the system, and many counters related to the data services Hammerspace provides — particularly file-system protocols and mobility. The exporters are disabled by default.

Hammerspace does not provide a Prometheus server. You supply your own, and configure it to scrape the Hammerspace cluster periodically to collect the counters.

To visualize the data, Hammerspace publishes dashboards for Grafana. These dashboards read from a Prometheus data source and require no further configuration once that source is in place.

Enabling or Disabling Prometheus Exporters

Use the following command and options to either enable or disable Prometheus exporters in Hammerspace.

  1. Log in to the Admin CLI.

  2. Run the following command to enable Prometheus exporters:

    Command:

    cluster-config --prometheus-exporters-enable

    Expected output (the first lines of the cluster configuration; the Prometheus exporters field shows the new state):

    ID:                                 aeeadb7f-7258-41e3-aad9-bc57ec176769
    Name:                               AEL640TQCSS160
    State:                              Standalone
    Management IPs:                     [192.0.2.10/20]
    Data IPs:                           [192.0.2.10/20]
    Cluster floating IPs:               [192.0.2.10/20]
    Since:                              2026-09-22 13:11:34 UTC
    Created:                            2026-09-22 13:11:15 UTC
    Timezone:                           UTC
    NFS transport policy:               TLS forbidden
    GFS participant max suspected time: 30 minutes
    Metered billing:                    Disabled
    Prometheus exporters:               Enabled
    <Additional output removed for brevity>

Alternatively, to disable the exporters, use the same command but with the prometheus-exporters-disable option:

cluster-config --prometheus-exporters-disable

Either command prints the cluster configuration, and the Prometheus exporters field in that output shows the new state. To confirm the current state later, run cluster-config with no options and check the same field.

Prometheus Exporter Ports

Remote monitoring with Prometheus requires network connectivity between the Prometheus server and the Hammerspace Anvil and DSX nodes. The exporters listen on the following TCP ports:

Node Protocol Ports

Anvil

TCP

9100 – 9103

DSX

TCP

9100 and 9105

Port 9104 is not used by the Hammerspace exporters.

Storing and Viewing Prometheus Metrics

The Hammerspace exporters expose metrics; they do not store them. To retain and visualize metrics, run a Prometheus server that scrapes the cluster, and point Grafana at that server as a data source.

Hammerspace publishes a set of Grafana dashboards for this purpose:

The following third-party documentation is not published by Hammerspace, but is useful when setting up the surrounding components:

Monitoring S3 Activity

When an S3 server is created and used, each DSX node keeps an in-memory summary of the S3 requests it serves. This summary is read directly from the node and is separate from the Prometheus exporters.

To view the S3 summary metrics for a single DSX node:

  1. Obtain the hostname or IP address of the DSX node, using the node-list command or the Infrastructure > Storage Systems page in the Management GUI.

    node-list
  2. In a supported web browser with network access to the DSX node, go to <dsx-node>/s3/metrics, where <dsx-node> is that node’s hostname or IP address. For example:

    https://gfs-dads1.company.internal/s3/metrics
    http://10.20.43.201/s3/metrics

The summary reports, for each S3 action that has been requested:

  • the number of times the action was requested;

  • the client-side error count (HTTP 400–499) and the server-side error count;

  • the server-side duration of the action, grouped into predefined ranges;

  • the slowest request for that action since the S3 service last started, with the client IP address that request came from.

Distributions differ from node to node, depending on each node’s hardware and its share of the S3 workload.

These counters are held in memory and accumulate only while the S3 service is running on that node. They are cleared whenever the S3 service restarts — during a software update, for example, or when the node reboots.