The Shared Object Storage Volume in Depth
The shared Object Storage Volume (shared OSV) is where cross-site file data lives in transit and at rest. This chapter explains how data is stored in the shared bucket and how to read the retained-version filenames in the snapshot directory.
How Data Is Stored in the Shared Bucket
When the owning site needs to make data available to other sites, its DSX nodes write the file’s data into the shared OSV as objects. Several OSV features reduce how much data must move and how much capacity it consumes:
- File chunking
-
4 MB chunk size, on by default. It can be disabled by adding the OSV in native mode, but disabling is uncommon.
- Deduplication
-
Hammerspace deduplicates by matching SHA-256 sums of chunks. Because only changed chunks move, the bytes transferred between sites are often far smaller than the file. Deduplication cannot be disabled (there is no benefit to disabling it) and typically saves 10–20%, sometimes up to 50%, depending on the data.
- Compression
-
Optional per OSV. It can save capacity and bandwidth but adds CPU load on the DSX nodes performing it; size DSX accordingly.
- Encryption
-
If a Key Management Service (KMS) is configured, new objects placed in the OSV are encrypted unless the OSV is added with the no-encryption option.
| These four features must be configured identically on the shared OSV at every participating site. In particular, if the KMS is not the same across all clusters, clusters with a different or absent KMS cannot decrypt objects written by sites with encryption enabled, and data orchestration fails. |
Garbage Collection
Coordinated Garbage Collection (CGC) removes objects from buckets that no longer have references from any site, keeping bucket utilization from growing unchecked.
Reading the Snapshot Directory: ls -l Decode
.snapshot/current (below) holds retained versions created automatically by
the log-xfer and undelete objectives — not the same thing as a formal
point-in-time Snapshot, which is a separate mechanism covered later in
Protecting data: snapshots, versioning, and undelete.
The two live in similarly-named locations (.snapshot here vs. .fsnapshot for
real Snapshots) but work independently; don’t confuse one for the other.
|
Retained versions created by log-xfer, versioning and undelete all appear in
.snapshot/current and all use the same filename format. What differs is what
triggered the copy, not how the filename looks: log-xfer on a change of site
ownership, versioning when a file is written again after having been closed for
a period, undelete on deletion. Each filename is made unique by an appended
timestamp and the originating site identifier:
# ls -lt .snapshot/current/
-rw-r--r-- 1 root root 217 Feb 28 07:33 file
-rw-r--r-- 1 root root 174 Feb 28 07:30 file[#M=20200916.212518.75035 #S=1-5]
The first entry is the live-tree version. The second is a saved version: #M= is
the timestamp, and #S=1-5 indicates Participant (site) ID 1 with object ID 5 in
the file system. The site ID can also be found from a mounted share with the
Hammerspace Toolkit (hs eval -e participants <PATH>) or via the admin CLI
(share-list --name <SHARE NAME>).