Search the docs

Ongoing Usage

General

  • Ensure that jobs are configured to use storage on Hammerspace shares, both in terms of mounting the shares on the GPU nodes, and placing data on them.

  • Ensure that node maintenance workflows do not take ownership of NVME drives and reformat or remount them. The Hammerspace Tier 0 Ansible playbooks include some scripts and a remount service to try to prevent GPU farm administrators and automation from unmounting volumes used by Hammerspace Tier 0.

  • Ensure that maintenance operations affect nodes within a single availability zone at a time.

  • For fleet updates (e.g., kernel patches), perform rolling updates at rack scale.

  • Disable "Availability Drop" during maintenance to prevent the system from prematurely evicting nodes that are merely rebooting.

  • Do not re-provision instances if a simple reboot suffices; re-provisioning can change IPs and wipe local storage volumes.

Ingesting Data

There are essentially three ways to get data into Hammerspace:

  • Copy in - Any data copy tool that can write to an NFS mount could be used to bring existing data into Hammerspace.

  • Create new - Workloads produce some kind of output or result. The Tier 0 architecture is designed to store that output.

  • Hammerspace Assimilation and Data Orchestration [Support article may require login to view.] - For existing data currently on other storage platforms, Hammerspace provides a way to "import" the metadata and hook a filesystem into a Hammerspace share, initially without moving the actual data. Objectives can be applied to the share into which the data is assimilated to move instances of the data onto Tier 0 or LSS nodes based on any metadata criteria such as time stamps, filename or extension, or path.

Snapshots

Hammerspace has a snapshot technology as one level of data protection. With supported storage systems, Hammerspace snapshots can use offloaded file cloning. This currently supports DSX nodes and some third-party storage systems - it does not currently support XFS file cloning on Linux servers including Tier 0 and LSS. When offloaded cloning is not available, the snapshots revert to full file copy on change. For this reason, snapshots should be avoided during checkpoint creation.