Search the docs

GFS Best Practices and Guardrails

Overall Capacity Planning: It is important to consider the overall capacity required for:

  • Data required to remain online at each location: Hammerspace will prefer NAS-based storage volumes, but will spill over to object storage tiers if capacity is needed. That is why it is essential to have spare capacity available if this occurs.

  • Any centralized, tiered data: If the desire is to treat each site as more of a "cache", size each site’s NAS volumes to accommodate the active working data, tiering to a centralized, shared object volume or bucket to optimize for capacity. Steps should be taken to ensure this object volume is highly durable and available.

  • All data that will be transferred between sites: A shared object volume will be used for site-to-site data transfers. Site-to-site transfers will stall if there is not enough capacity within the share object volume to accommodate the data being requested.

    • It is recommended that enough capacity be configured to host 14-days worth of data transfer.

    • It is also recommended that a second bucket be configured with a higher online-delay objective. This allows for all transfers to prefer one bucket, using the second only if the first is unavailable, ensuring that data transfers are not interrupted due to a single bucket outage..

Data Deduplication: Deduplication can significantly reduce storage consumption, with savings of at least 10-20% and sometimes up to 50% depending on the data set. Dedupe is on by default and it’s recommended to keep it on.

Data Compression: While enabling compression can reduce data significantly, it will also bring additional load to the DSX nodes performing said compression. Care should be given to sizing the DSX nodes appropriately to accommodate this additional load.

Coordinated Garbage Collection (CGC): CGC is a process that removes objects from buckets that no longer have references from any site, which helps to reduce bucket utilization and prevent object space from growing unchecked.

Consistency is Key: If using encryption, ensure that the passphrase or KMS key is consistent across all sites using the same bucket. Inconsistent encryption can prevent sites from decrypting files and will cause data orchestration to fail.

Addressing Specific Protocol Challenges

SMB workflows, particularly those involving file creation and renaming (e.g., "New Word Document" → rename), generate multiple metadata operations in quick succession. When combined with Replication delays, this can easily lead to race conditions and metadata conflicts (e.g., conflict files, lost and found entries). Implement very strict guardrails for SMB use in multi-SITE Active-active scenarios. It is generally best confined to active-passive or read-only access in replicated environments.

Handling Replication Delays and Disconnections

Replication Falling Behind: This can be caused by under-resourced hardware (slow disks, insufficient CPU/memory), network issues (latency, bandwidth), or workload patterns (high churn, concurrent writes).The goal is to keep Replication delay under an hour.

Extended Disconnection: While GFS can run indefinitely disconnected, reconnecting after a long period (especially if active changes occurred on multiple sites) significantly increases the risk of complex metadata conflicts.

For active-active scenarios, extended disconnections should be avoided. If a site is disconnected for a prolonged period, treat it as a potential source of data inconsistencies upon reconnection, and be prepared for manual reconciliation.

Metadata and Data Collisions

Metadata and data collisions occur when two or more sites modify the same file or directory simultaneously. Another possible scenario is during situations where sites are unable to reach one another, yet workers at each site continue to work on their local instances of the same data. In this case manual steps may be necessary to reconcile these updates.

Data Collisions: When a file is being actively written to on one site, it cannot be accessed or uploaded to another site. A remote site attempting to read or write to that file will hang until the file is closed by the primary writer. This can cause delays for the end user. Note that this should not be confused with snapshot activity in which any open files will still be protected in the snapshot. If files remain open and access is needed on other sites, these files can then be accessed from the snapshot.

Conflict Files: In cases of simultaneous writes to the same file, a conflict file is created. While this prevents data loss from a merge, it indicates a problematic workflow. Where possible, forwarding workloads to the same site or location can significantly reduce this phenomenon.

Handling Conflicts

When using data across multiple sites, for those very far apart from each other, there are several considerations that come into play that affect the end user experience:

  • How fast will I see new files?

  • What happens if I modify the same file on both sites at the same time?

  • What if the sites are disconnected from each other for an extended period?

This often comes down to how ownership and conflict resolution is handled. With Hammerspace, there are several options that can be applied to ensure various workflows can be supported.

As previously discussed, objectives are used to help manage data across sites. Objectives can place data locally and ensure that local copies are kept up to date when changes happen to files/objects on other sites. Objectives can also determine the behavior for such files/objects when conflicts occur:

  • log-xfer-<time> objective is used to ensure that a version of the file is retained before the file changes owner. File ownership is determined by the site that last wrote to it. File ownership is important if another site needs to use the data, but it does not have a local copy of the data, it will reach out to the site that owns the data and orchestrate mobility to the local site.

  • undelete-<time> objective can help save files for a specific period if they are deleted. Undelete enables a quick recovery when files are accidentally deleted.

The Global File System uses an eventually consistent model, which relies on conflict resolution logic for when there are conflicts. Cross-site file locking is not supported.

File Conflict Scenario #1: Creating the Same Filename Across Multiple Sites at the Same Time

If the same file is created or modified on multiple sites during the same replication interval (or when sites are disconnected from each other), then the Global File System will automatically create two or more versions of the file, a site-local file with the original name, and a second file representing the remote version of the file. If more than two sites collide at the same time, there will be versions from each site.

Example

In the example shown in the table below, we create a file with the same filename, but a different size file, from each site to help explain the behavior. The commands are run from a local Linux client located on each site, with the share mounted under /global.

SITE A (ID: 0) SITE B (ID: 1) SITE C (ID: 2)

# cd /global

# dd if=/dev/urandom of=file bs=1024k count=1 conv=notrunc

# cd /global

# dd if=/dev/urandom of=file bs=1024k count=5 conv=notrunc

# cd /global

# dd if=/dev/urandom of=file bs=1024k count=9 conv=notrunc

Wait until the replication interval has passed and view the directory with ls -l. Note that the NFS client will cache directory contents which could temporarily display information that is out of date. Keep in mind that some less-relevant columns have been removed due to space limitations.

Note the highlighted file across shares:

SITE A (ID: 0) SITE B (ID: 1) SITE C (ID: 2)

# ls -l

total 15360 1048576 Feb 28 00:44 file 5242880 Feb 28 00:44 file[#S=1] 9437184 Feb 28 00:44 file[#S=2]

# ls -l

total 15360 5242880 Feb 28 00:44 file 1048576 Feb 28 00:44 file[#S=0] 9437184 Feb 28 00:44 file[#S=2]

# ls -l

total 15360 9437184 Feb 28 00:44 file 1048576 Feb 28 00:44 file[#S=0] 5242880 Feb 28 00:44 file[#S=1]

As shown in the table, above, the same file name exists on all the sites however the conflict file version from the other sites is automatically renamed to FILENAME[#S=SITE ID]. This is done to ensure that the file being used locally does not suddenly change name and it also ensures that other sites do not overwrite that local file content.

To resolve a conflict, simply copy the renamed file to the original file name and delete the automatically named file if it is no longer needed.

File Conflict Scenario #2: Writing to an Existing File on Multiple Sites at the Same Time

While file metadata replicates to all sites, the actual file data (the file contents) is only replicated when requested, either explicitly through objectives or implicitly when the file is accessed. Objectives ensure that each site has a local copy of the data before writing to the file.

By default, the last writer will win when two sites write to the same file within the same replication interval. To help protect against user mistakes, the Global File System has the capability to preserve previous file content. This functionality must be turned on, which is done by using an objective.

Since the site-ownership ends up creating additional copies of files (very similar to Undelete and Versioning functionality), these extra copies have an expiration date, which is the time set by the objective. If snapshots are taken before files are deleted, then the files will remain within the snapshots until either the snapshot retention period or the file expiration period ends, whichever is the longest.

Site Ownership Objectives

The name of the objective indicates how long these files are retained in the live tree. When a retained file is copied, the retention is not copied with the file.

  • log-xfer-1-hour

  • log-xfer-1-day

  • log-xfer-1-week

  • log-xfer-1-month

The retained files are available in .snapshot/current,and each filename is made unique by both a timestamp and the originating site identifier being appended to the end of the files.

*# ls -lt .snapshot/current/*

[NOTE]

-rw-r--r-- 1 root root 217 Feb 28 07:33 file
-rw-r--r-- 1 root root 174 Feb 28 07:30 file[#M=20200916.212518.75035 #S=1-5]

In the example above, the file entry represents the live tree version, and the second entry represents the saved version of the file with the timestamp appended to the file. S=1-5 translates to Participant ID (site) 1 and 5 is the object ID in the file system.

The site ID can be determined from the mounted share itself if the Hammerscript and Hammerspace Toolkit (HSTK) [Support article may require login to view.] is installed, or it can be determined via the admin login using the share-list --name <SHARE NAME> output.

User Education for Workflows

End users often lack awareness of the underlying eventually consistent nature of a GFS. For workflows that involve shared data and multiple sites, administrators must implement:

  • Workflow Integration: For complex multi-site, multi-user scenarios (e.g., in the media & entertainment industry), integrate with third-party workflow management systems (e.g., ShotGrid via ShotHammer plugins) to enforce check-in/check-out, manage file access, and coordinate changes.

  • Site-Specific Workspaces: Encourage users to work within designated site-specific directories (e.g., /home/user/siteA_projects) when performing active writes to avoid conflicts.

  • Clear Guidance: Provide explicit instructions on how users should interact with shared data across sites or when to use site-specific copies.

  • Asynchronous Reservation (Advisory Locking): Implement advisory locking mechanisms to signal intent to modify a file. Users must understand that this is advisory and relies on Replication being in sync. It requires a manual process to acquire and release the "lock" and verify its status across sites.