Search the docs

Monitoring and Alerting

Replication is eventually consistent; therefore, continuous monitoring is critical to ensure predictable performance and consistency.

  • Monitor Latency: Actively monitor Replication Lag metrics (available via the API/CLI/GUI/Prometheus) to ensure the Metadata Path remains responsive and that lag does not consistently exceed key thresholds. (Note that by default, an alert will be triggered if replication lags beyond 60-seconds, or 5x the desired replication interval, whichever comes first.) The Hammerspace GUI provides a "Global File System" section where administrators can view Replication latency charts (1-hour, 1-day, 1-week views) for each share.

  • Alert Thresholds: Configure alert thresholds within the GUI for Replication latency. If latency exceeds this threshold, an alert will be displayed in the UI.

  • Grafana: Visually track Replication health in real-time with metrics from Prometheus and proactively detect delays and failures before they become critical with customized dashboards and alerts. For enabling the exporters and installing the dashboards, see Monitoring Cluster Health in the Hammerspace Administration Guide.

  • API Access: All data displayed in the GUI is accessible via APIs, allowing for custom monitoring solutions.

  • SNMP/Syslog Integration: Leverage SNMP or Syslog to integrate Hammerspace alerts into existing enterprise monitoring systems.

Reference Appendix B to see common errors that may be encountered.