Search the docs

Architecture

With the shared-OSV concept in place, the node roles are straightforward. Each site runs two kinds of Hammerspace node — Anvil and DSX — with very different responsibilities. Each site is designed to operate independently of the others.

gfs architecture diagram
Figure 1. Conceptual GFS architecture: two sites, one shared OSV

In words: each site — whether on-premises or in the cloud — runs its own Anvil and one or more DSX nodes, serving its own local users and applications directly over NFS, SMB, or S3. Anvil nodes handle all cross-site metadata replication, and DSX nodes coordinate directly with their counterparts at other sites about what data needs to replicate — but the file bytes themselves move only through the shared OSV, not directly between DSX nodes. A site that loses its connection to the others keeps serving its own local data uninterrupted.

The second line in the figure — data replication through the shared OSV — is the node-level view of the walkthrough in Core concepts: the shared Object Storage Volume, under "How a site gets a file it does not hold locally." That walkthrough describes what happens; this figure shows which nodes do it.

Anvil — Metadata and Management

Anvil nodes are responsible for all file-system and extensible metadata and for systems management, including the API and Management GUI. Anvil nodes handle all cross-site metadata replication, bi-directionally, as point-to-point connections between Anvil nodes.

  • Connections between Anvil nodes are encrypted with TLS (AES-256). Customers may install their own certificate.

  • Four ports must be open between all Anvil nodes across all participating sites: 443, 9097, 9298, 9299.

DSX — Data Services

DSX nodes transfer data chunks to and from cloud/object storage, use local storage to hold data, and present NFS, SMB, and S3 to site-local clients. DSX nodes coordinate directly with their counterparts at other sites about what data needs to replicate, but the file bytes themselves move only through the shared OSV — not directly between DSX nodes. DSX nodes must have full connectivity to the shared OSV(s).

Why Sites Can Briefly Differ

GFS keeps every site’s copy of the namespace consistent over time, not instantly. Two things follow directly from the Anvil/DSX split above:

  • Metadata and data replicate independently. A file, or a change to one, can show up in the namespace at a remote site — carried there by Anvil-to-Anvil metadata replication — before its actual bytes have finished moving through the shared OSV. That’s expected, not a fault: the namespace and the data behind it travel on separate paths, at separate speeds.

  • Changes propagate asynchronously. A write made at one site reaches the other sites soon afterward, not instantly and not necessarily in the order it was made. See How conflicts are handled for what this means when two sites change the same file.

This same separation is also why cleanup between sites has to be coordinated: because more than one site can reference the same object in the shared volume, removing it safely means the sites first have to agree nothing still needs it. See The shared Object Storage Volume in depth for how that coordinated garbage collection works.

Deduplication Chunk Tracking

Hammerspace stores file data in the shared OSV as deduplicated chunks, matched by content rather than by file. When a file changes, only the changed chunks need to move between sites, and chunks already present in the shared OSV are not re-uploaded. This is what keeps cross-site transfer small relative to total file size. See The shared Object Storage Volume in depth for chunk size, how deduplication is applied, and the other OSV features that reduce what has to move.