Search the docs

Hammerspace Multi-Site Fundamentals

Hammerspace is designed to provide an active-active, eventually consistent Global File System (GFS) across multiple geographically dispersed sites. It operates by replicating metadata and employs Hammerspace objectives to instantiate data at designated site participants, ensuring accessibility and resilience.

  • Eventually Consistent: Changes made on one site will eventually propagate to all other replicated sites.

  • Active-Active: All sites can be enabled to serve data and accept writes. Keep in mind that global accessibility can allow for active-active write patterns and requires consideration and management.

  • Metadata Replication: Checkpoints of metadata are taken at regular intervals (defaulting to 5 seconds per share/export) to replicate metadata diffs so that the GFS is consistent.

  • Data Replication: Data movement leverages a shared object storage bucket. Data is uploaded to the shared bucket by the owning site and pulled by requesting sites, either by objectives or on-demand.

Key Features

The Hammerspace Global File System offers the following key features:

  • Active-active data access.

  • Global access to data across multiple standard protocols (NFS 3, NFS 4.1, NFS 4.2, SMB 2.x/3.x, and S3). No proprietary client software is required.

  • Custom metadata, which can be manually or automatically applied.

  • Objective-based data policies, triggered by any combination of file system and/or custom metadata to automate a wide range of data services and data placement actions.

  • Optimized and transparent data orchestration across sites with integrated global deduplication and compression.

  • Secure connections and encryption of the data payload during transfer and at rest in cloud/object storage.

  • Global snapshots that can be stored anywhere.

  • Global undelete.

  • File-granular WORM, which can be part of an automated objective triggered by any metadata fragment.

  • Anti-virus scanning.

  • Access-based Enumeration (ABE) across NFS, SMB, and S3.

  • Conflict resolution with cross-site ownership versioning.

  • Per-site data management for optimal resource usage.

  • Heterogenous storage support. Each site can use one or more storage platforms of any type, from any vendor.

  • Capacity agnostic. Only requires enough capacity for active data.

  • Intelligent caching. Drive local data availability through automated objectives triggered by any combination of multiple metadata types.

  • Highly multi-threaded, cross-site data movement.

  • Tolerant of high-latency for site-to-site connectivity.

  • Supports up to 16 sites per share.

  • File-granular data management (tiering, archiving, cloning, mobility…).

  • Non-disruptive data mobility, even on live data.

  • Support for extremely high-performance AI and HPC use cases on commodity storage, with support for Parallel NFSv4.2 with FlexFiles.

Key Components

Component Description

Sites

Sites can be anything from geographically separated locations, all the way to fault domains within the same facility. Each "site" will host Anvils and can host some combination of DSX, Linux Storage Server (LSS) and 3rd-party NAS. Each site will also have its own set of objectives that dictate data placement for that site.

Anvils or Metadata Servers

Each site has metadata servers known as Anvils that are responsible for managing file system metadata (directory structure, file names, timestamps, etc.). These are typically deployed as HA pairs for site-level resiliency.

DSX Nodes

These servers have three roles: store, mover, and portal. The one we are mainly concerned with in the GFS context is mover, which handles the actual data transfer. Movers interact with cloud buckets and local storage to place file instances.

Shared Object Buckets

A central object storage location used for transferring file data between sites. This must be reachable by all sites participating in a given share.

Objectives

Policies defined by administrators to control data placement (e.g., "keep a copy in cloud," "keep a copy on SITE B").

Clients

User machines accessing the GFS via protocols like NFS, SMB, S3, and CSI.

Distributed Workflows: Different sites work on the same dataset but at different, non-overlapping times. It is crucial to ensure that Replication is fully caught up before the next site begins active writes.

Burst-to-Compute: Compute and data are often not co-local. Hammerspace will orchestrate data to compute resources regardless of location whether it is a remote data center and/or public cloud (cloud bursting / hybrid cloud). Output data should ideally be written to new, site-specific locations or otherwise managed carefully to avoid conflicts.

Multi-site Collaboration: Enable follow-the-sun workflows where the primary "active" site performing read/write operations shifts geographically as the workday progresses globally.

Site-Specific Data: Workloads where each site operates on its own distinct set of files/directories, with minimal or no concurrent modification of shared datasets.

Data Tiering to the Cloud and/or Object Storage: Tier data seamlessly in the background without interrupting user access. This enables active data to remain on Tier 0 and Tier 1 while orchestrating data to Cloud and/or Object storage for dormant data as an active archive.

Multi-site Consolidation: Hammerspace global file system provides a single logical system across multiple sites allowing customers to consolidate their storage regardless of physical location.

Unsupported Use Cases and Configurations

Concurrent Updates to the Same File: Applications that attempt to modify the exact same file simultaneously from multiple active sites are unsupported. Examples include:

  • Multiple batch schedulers writing to the same log file.

  • Concurrent pip install into a shared directory can lead to writing to the same file(s)

  • Shared home directories accessed concurrently from multiple VDI instances or workstations (e.g., Chrome profiles).

  • Version control systems (e.g., Git) used concurrently on the same repository from different sites without explicit workflow controls.

  • Engineering design applications (e.g., Windchill) or media editing software that expects synchronous, byte-range locking across sites for shared files.

  • Application logging directories that are not unique to each site

  • Large, read-write data files associated with databases or virtual machine images being concurrently run at multiple sites or changing sites with no time to allow data to move to the new site.

Considerations for the Following Configurations: The following setups can cause significant overhead and performance issues during snapshots due to the lack of offloaded cloning support.

  • A system with a single NAS instance and no offloaded clones is not recommended for Replication.

  • Systems not listed on the Hardware and Software Compatibility List require specific testing of their offloaded clone capabilities before they can be considered fully supported.

IMPORTANT: Hammerspace does not provide synchronous, cross-site locking. Any application or workflow that relies on this will encounter issues.

Data Orchestration Replication Flow Example - Staged Workflow

  1. File Modification (SITE A): A user modifies a file on SITE A.

  2. Metadata Update: SITE A updates the file’s metadata (e.g., modify time, last use time).

  3. Metadata Checkpoints: Regular checkpoints capture these metadata changes.

  4. Metadata Replication: Metadata changes are replicated to other sites (SITE B, SITE C, etc.). This involves sending diffs computed from metadata checkpoints.

  5. Proactive Upload: An objective on the owning site (SITE A) triggers an upload of the new dataset to the shared object volume(s).

  6. Data Download: An objective on a remote site (SITE B) identifies that it should have the latest copy of the file. It then downloads a copy of the file from the shared object volume.

  7. Local File Copy: The downloaded data from the shared object volume is used to create a local NAS copy of the file on SITE B.