Monitoring and Alerting
Because replication is eventually consistent, continuous monitoring is essential to keep performance and consistency predictable.
What to Watch
- Replication latency
-
Monitor replication-lag metrics (available via API, CLI, GUI, and Prometheus) so the metadata path stays responsive and lag does not consistently exceed your thresholds. In the Management GUI, the Global File System section shows replication-latency charts (1-hour, 1-day, 1-week) per share.
- Default alert threshold
-
By default, an alert is raised when replication lags beyond 300 seconds. Configure alert thresholds for replication latency in the GUI (per-share, on the Global File System tab of the Details view) or via the API/CLI.
Integrations
-
Grafana / Prometheus — track replication health in real time and detect delays and failures before they become critical with custom dashboards and alerts. To turn on the exporters and set up dashboards, see Monitoring Cluster Health in the Administration Guide.
-
API — everything shown in the GUI is available via API for custom monitoring.
-
SNMP / Syslog — forward Hammerspace alerts into existing enterprise monitoring.
A Healthy GFS at a Glance
-
Replication lag well under the alert threshold and not trending upward.
-
Participants showing Admin/Oper state Up across all shares.
-
Shared OSV capacity comfortably ahead of in-flight transfer (see the 14-day guideline in Prerequisites, Sizing, and Networking).
-
No stale object-volume reservations or sites stuck in a removing/cleaning state (see Troubleshooting: Removal and Decommissioning).