Monitoring Cluster Health
The local drives on a product node are automatically monitored for space utilization. There is a WARNING event at 70% and an EXCEEDED event at 90%. This is done by a space monitoring service, which checks all the local filesystems (e.g. /, /boot, /var etc) every two minutes. This runs on All Anvil and DSX nodes.
It is worth noting that the Prometheus exporters will also collect this information and it can be viewed graphically with Grafana.
Here is an example of the event text:
____ ID: d0719f4c-b1ca-11ef-9bd4-30e171639c80 Created: 2024-12-03 23:03:31 UTC Modified: 2024-12-03 23:03:31 UTC Type: DISK_USAGE_LIMIT_EXCEEDED Description: Disk usage limit EXCEEDED on doxie-dsx1.lab.hammer.space:/boot with only 106M available. Severity: ALERT ____
About the Prometheus Exporters
Hammerspace ships integrated Prometheus exporters. They export data about the system, the resource utilization of the system, and many counters related to the data services Hammerspace provides — particularly file-system protocols and mobility. The exporters are disabled by default.
Hammerspace does not provide a Prometheus server. You supply your own, and configure it to scrape the Hammerspace cluster periodically to collect the counters.
To visualize the data, Hammerspace publishes dashboards for Grafana. These dashboards read from a Prometheus data source and require no further configuration once that source is in place.
Enabling or Disabling Prometheus Exporters
Use the following command and options to either enable or disable Prometheus exporters in Hammerspace.
-
Log in to the Admin CLI.
-
Run the following command to enable Prometheus exporters:
Command:
cluster-config --prometheus-exporters-enableExpected output (the first lines of the cluster configuration; the Prometheus exporters field shows the new state):
ID: f976411d-c581-4d56-a48a-2d8c70b13b60 Name: AEEPQ7Z5Z3487K State: Standalone Management IPs: [192.0.2.10/20] Data IPs: [192.0.2.10/20] Cluster floating IPs: [192.0.2.10/20] Since: 2026-09-23 21:17:03 UTC Created: 2026-09-23 21:16:42 UTC Timezone: UTC GFS participant max suspected time: 30 minutes Metered billing: Disabled Prometheus exporters: Enabled Evaluation expiration: 2026-10-23 21:16:42 UTC <Additional output removed for brevity>
Alternatively, to disable the exporters, use the same command but with the prometheus-exporters-disable option:
cluster-config --prometheus-exporters-disable
Either command prints the cluster configuration, and the Prometheus exporters field in that output shows the new state. To confirm the current state later, run cluster-config with no options and check the same field.
Prometheus Exporter Ports
Remote monitoring with Prometheus requires network connectivity between the Prometheus server and the Hammerspace Anvil and DSX nodes. The exporters listen on the following TCP ports:
| Node | Protocol | Ports |
|---|---|---|
Anvil |
TCP |
9100 – 9103 |
DSX |
TCP |
9100 and 9105 |
| Port 9104 is not used by the Hammerspace exporters. |
Storing and Viewing Prometheus Metrics
The Hammerspace exporters expose metrics; they do not store them. To retain and visualize metrics, run a Prometheus server that scrapes the cluster, and point Grafana at that server as a data source.
Hammerspace publishes a set of Grafana dashboards for this purpose:
-
Hammerspace Grafana Dashboards — downloading, installing, and configuring the dashboards and the related components.
The following third-party documentation is not published by Hammerspace, but is useful when setting up the surrounding components:
-
Monitoring Linux host metrics with the Node Exporter — installing Prometheus exporters on Linux nodes.
-
DCGM Exporter — using Prometheus and Grafana with NVIDIA GPU nodes.
Monitoring S3 Activity
When an S3 server is created and used, each DSX node keeps an in-memory summary of the S3 requests it serves. This summary is read directly from the node and is separate from the Prometheus exporters.
To view the S3 summary metrics for a single DSX node:
-
Obtain the hostname or IP address of the DSX node, using the
node-listcommand or the Infrastructure > Storage Systems page in the Management GUI.node-list -
In a supported web browser with network access to the DSX node, go to
<dsx-node>/s3/metrics, where<dsx-node>is that node’s hostname or IP address. For example:https://gfs-dads1.company.internal/s3/metrics http://10.20.43.201/s3/metrics
The summary reports, for each S3 action that has been requested:
-
the number of times the action was requested;
-
the client-side error count (HTTP 400–499) and the server-side error count;
-
the server-side duration of the action, grouped into predefined ranges;
-
the slowest request for that action since the S3 service last started, with the client IP address that request came from.
Distributions differ from node to node, depending on each node’s hardware and its share of the S3 workload.
| These counters are held in memory and accumulate only while the S3 service is running on that node. They are cleared whenever the S3 service restarts — during a software update, for example, or when the node reboots. |