Using the Hammerspace Toolkit
In this section, we will review all of the commands provided by the HSTK, and see common examples of how the commands can be used. The HSTK commands can be arranged into four sections based on their use:
-
Metadata Operations - Interact with file metadata
-
Reporting - Obtain detailed statistics and information about the filesystem
-
Filesystem Operations - Perform a variety of common filesystem operations, but offload the execution to Hammerspace
-
General - Other tasks that fall outside of the other categories
Keep in mind that this document only covers a small portion of the values you can query using Hammerscript. Consult the File Metadata Example (Full) section of this document for an example of the typical metadata generated for every file. Using the examples provided in this document, you can begin to build your own HSTK commands that provide whatever filesystem information that you require.
Running HS Commands from the Cluster Root Export
Mounting the cluster root export and running hs commands from there enables you to gather data on all cluster shares at once. Note that hs commands cannot natively traverse from the cluster root into share folders to summarize data. However, there are ways around this behavior that provide us with the simplest method for gathering data on all cluster shares at once.
In this guide, you will see examples where we use the Linux ls and find commands from the cluster root export to feed paths to hs commands, allowing them to directly target each share folder using one command execution. The collected output from these commands, whether viewed raw or exported into spreadsheet tools, enable us to analyze report data for all cluster shares using a single report.
Best Practice:
-
If you mount the cluster root export to use the HSTK, for safety reasons, it should be for gathering information only.
-
If you need to use the HSTK to make metadata or file changes, you should mount the share directly.
Commands - Metadata Operations
The most common use case for the HSTK is to query and manipulate file metadata for the purposes of reporting, data orchestration (via objectives), and data management. In this section, we will review the subset of HSTK commands that are used for these tasks. More advanced HSTK examples will be provided in the Advanced Filesystem Reporting section of this document.
| Most of the Hammerscript examples below can be used for both file system reporting and as conditions for objectives. Consult the How to Configure Hammerspace Objectives [Support article may require login to view.] for examples about how you use these values as objective conditions. |
Tag
The hs tag command enables you to view, set and remove Hammerspace custom metadata tags. Tags are a key-value pair, where the value can be specified or excluded. If no value is supplied, the tag will be set to TRUE. Tags can have a value added as a string (-s ‘some value') or a Hammescript expression (-e ‘some expression'). In this document, all examples used are strings.
Tags are set using the hs tag set <tag name> <file or directory name> command, checked using the hs tag list <file or directory name>, and removed with hs tag delete <tag name> <file or directory name> as shown in the following examples:
# Apply tag Project with value Apollo to file recovery.txt
$ hs tag set Project -s 'Apollo' recovery.txt
# List tags applied to file recovery.txt
$ hs tag list recovery.txt
TAGS_TABLE{
|NAME = "Project",
|VALUE = "Apollo"}
# Remove tag Project from file recovery.txt, you do not need to specify the value
$ hs tag delete Project recovery.txt
As stated in the example command, you do not need to specify the tag value when removing tags.
Best Practice: Where possible, implement controls or standards regarding the usage of tags. Given that tags are often targeted by objectives or even HSTK queries, it’s important that their spelling, case, and exact terms used are consistent.
You can also set tags using the Hammerspace Metadata Plugin for Microsoft Windows. Consult the Hammerspace Metadata Plugin for Windows File Explorer section of this document for information about using that tool to manage file or directory tags.
Using Numbers as Tag Values
In the previous section, we mentioned that tags can have a value, and that value can be set as an expression or a string. It is important to note that expressions include numbers, and there are reasons why you might want to set a numeric value as an expression rather than a string.
The following is an example of how a number set as an expression could be used. In this example, we will add a tag with a value of 3 set as an expression.
# Add tag launchphase with an expression of 3 to file launch-meco-tail.jpg
$ hs tag add launchphase -e 3 launch-meco-tail.jpg
# List tags for file launch-meco-tail.jpg
$ hs tag list launch-meco-tail.jpg
TAGS_TABLE{
|NAME = "launchphase",
|VALUE = 3;
Tag values set as strings have limited options as to how they can be evaluated, meaning the tag value either does or does not match the string provided. However, numeric tag values added as expressions can be evaluated using any available operator.
In this example, we see how we can use greater than(>) or less than or equal to(⇐) operators to find tagged files whose values meet the condition specified.
Specifically, we are using the label values to find photos (recursively, -r) that match certain phases of a rocket launch.
# Return files with tag launchphase whose value is greater than 3
$ hs eval -r -e 'GET_TAG("launchphase")>3'
launch-feco-tail.jpg
launch-orbit-side.jpg
# Return files with tag launchphase whose value is less than/equal to 3
$ hs eval -r -e 'GET_TAG("launchphase")<=3'
launch-meco-tail.jpg
launch-burn-tail.jpg
launch-liftoff-tail.jpg
launch-countdown-side.jpg
Best Practice: This example shows why it is important to consider not only the format of the tag names, but the values themselves. While the tag values can always be updated, you should consider how you plan to use them when setting the guidelines for tag use.
Integrating Tags into Other Workflows
Hammerspace has published example scripts that set tags on files based on embedded tags extracted from files or from cloud analytics tools, like Amazon Rekognition. Scripts available in GitHub include:
-
exif2hs - Uses the open-source exiftool to extract standard tags from a wide set of file formats, then sets them as Hammerspace tags.
-
rek2hs - Obtains the object reference for a file with an instance in an Amazon bucket, passes the object reference to Rekognition, then scrapes the Rekognition label and confidence, and sets them as Hammerspace tags.
Keyword
The hs keyword command can be used to apply keywords to files and directories, which allows them to be targeted by objectives or even HSTK queries. Keywords are set using the hs keyword add <keyword name> <file or directory name> command, checked using the hs keyword list, and removed removed with hs keyword delete <keyword name> <file or directory name> as shown in the following examples:
# Add keyword FAA to file telemetry.txt
$ hs keyword add FAA telemetry.txt
# List keywords added to file telemetry.txt
$ hs keyword list telemetry.txt
KEYWORDS_TABLE{
|KEYWORD = "FAA"}
# Remove keyword FAA from file telemetry.txt
$ hs keyword delete FAA telemetry.txt
Best Practice: Note that keywords do not use a schema, meaning they are free-form. Where possible, implement controls or standards regarding the usage of keywords. Given that keywords are often targeted by objectives or even HSTK queries, it’s important that their spelling, case, and exact terms used are consistent.
You can also set keywords using the Hammerspace Metadata Plugin for Microsoft Windows. Consult the Hammerspace Metadata Plugin for Windows File Explorer section of this document for information about using that tool to manage file or directory keywords.
Label
The HSTK can be used to assign existing labels to files and directories as metadata. Labels have an administrator-defined schema that is set up using the Hammerspace CLI or API. Users can only apply labels that are already defined.
This type of metadata only lives within the file system itself, and does not change the contents of the file. The metadata is also replicated when used together with the Global File System feature.
|
Apply a Label
As with other custom metadata, labels can be applied to a file or directory using the hs label add command:
# Add label Cape Canaveral to file telemetry.txt
$ hs label add "Cape Canaveral" telemetry.txt
Label "Cape Canaveral" was created to have an implied (parent) label of "US", so when we list the labels using hs label list, we see that both have been applied:
# List labels applied to file telemetry.txt
$ hs label list telemetry.txt
LABELS_TABLE{
|LABEL_ = LABEL('US'),
|COUNT = 1;
|LABEL_ = LABEL('Cape Canaveral'),
|COUNT = 1}
Remove a Label
The hs label delete command is used to remove labels. If the label removed has an implied (parent) label, that will also be removed.
# Delete the label Cape Canaveral from file telemetry.txt
$ hs label delete 'Cape Canaveral' telemetry.txt
# List labels applied to file telemetry.txt
$ hs label list telemetry.txt
LABELS_TABLE{}
Attribute
Attributes have a system-defined hierarchy. In other words, the set of attributes is predefined as part of the Hammerspace product. In this section, we will review some of the most basic uses of attributes. For a more advanced example that better illustrates how they might be used, refer to the Hammerspace and Data Lifecycle Management section of this document.
You can list the some of the defined attributes and their values for a file using the hs attribute list command:
# List attributes for file telemetry.txt
$ hs attribute list telemetry.txt
ITEM_ATTRIBUTES_TYPED_TABLE{
|VIRUS_SCAN = VIRUS_SCAN_STATE('UNSCANNED'),
MIME = MIME_TYPE_TABLE{
|STRING = "/"},
|DO_NOT_MOVE = FALSE,
|CONSIDER_VOLATILE = FALSE,
|RECENTLY_USED_DURATION = 00:05:00,
|SELECTED = FALSE,
|SELECTED2 = FALSE,
|PROMOTE = FALSE,
|DEMOTE = FALSE,
COLORS = COLORS_TABLE{},
CSI_DETAILS = CSI_DETAILS_TABLE{}}
While we won’t go into all the attributes and their uses in this document, we will use "COLORS" as an example. The COLORS attribute contains a set of colors applied to a file. The allowed list of colors is predefined in the Hammerspace product, similar to how colors are available in Apple MacOS. The hs attribute set colors command is used to set file colors:
# Set the color to RED for file telemetry.txt
$ hs attribute set colors -e 'COLORS_TABLE{COLOR_LOOKUP("RED")}' telemetry.txt
# List the attributes for file telemetry.txt
$ hs attribute list telemetry.txt
ITEM_ATTRIBUTES_TYPED_TABLE{
|VIRUS_SCAN = VIRUS_SCAN_STATE('UNSCANNED'),
MIME = MIME_TYPE_TABLE{
|STRING = "/"},
|DO_NOT_MOVE = FALSE,
|CONSIDER_VOLATILE = FALSE,
|RECENTLY_USED_DURATION = 00:05:00,
|SELECTED = FALSE,
|SELECTED2 = FALSE,
|PROMOTE = FALSE,
|DEMOTE = FALSE,
COLORS = COLORS_TABLE{
|COLOR = COLOR_LOOKUP('RED'),
|COUNT = 1},
CSI_DETAILS = CSI_DETAILS_TABLE{}}
| Consult the Hammerspace Metadata Plugin for Windows File Explorer section of this document for information about using that tool to manage the color attribute. |
You can set multiple colors by separating the lookups with a semicolon:
# Set the color to RED and PURPLE for file telemetry.txt
$ hs attribute set colors -e 'COLORS_TABLE{COLOR_LOOKUP("RED");COLOR_LOOKUP("PURPLE")}' telemetry.txt
Once set, you can use the hs eval command to query for files with a particular color selected:
# List all files that have the color set to GREEN, including path
$ hs eval -r -e 'HAS_COLOR("green")?PATH' .
"./telemetry.txt"
"./images/launch_pad.jpg"
Keep-On-Site
One of the default objectives applied to a Hammerspace share keeps files online that have the keep-on-site value set for that site. Assuming that the default objective is in place (or one similar to it), setting this value will keep the file (or contents of the directory) online using nothing more than the default share objectives.
When preparing to set this value, first list the sites using the hs keep-on-site available command:
# List all sites where the current share is available
$ hs keep-on-site available
NYC-HS-A
LA-HS-A
In this example, we see that the share is available on two sites - NYC and LA.
| If the data is in a share that is not a global file share, only the local cluster name will be returned. |
The value can then be set with the Hammerspace toolkit using the hs keep-on-site command specifying add or delete as needed:
# Add the keep-on-site designation for site NYC-HS-A to the current directory
$ hs keep-on-site add NYC-HS-A .
# Remove the keep-on-site designation for site NYC-HS-A to the current directory
$ hs keep-on-site delete NYC-HS-A .
In the previous example, we targeted the current directory using a period (.) - you can also specify a directory or file name.
| Consult the Hammerspace Administrators Guide and How to Configure Hammerspace Objectives [Support article may require login to view.] for additional information about the default share objectives, and the keep-on-site feature. |
Rekognition-Tag
The hs rekognition-tag command is also used to set file and directory tags, and follows the same format as the hs tag command. Consult the tag section of this document for the command syntax and examples.
Reporting
The HSTK and Hammerscript can be used for a wide variety of filesystem reporting needs, at whatever level of granularity that you require. In this section, we will review the subset of HSTK commands that are used for file system reporting.
More advanced examples are available in the Advanced Filesystem Reporting section of this document.
Frequently Used Metadata Values
In this section, we will outline some of the most commonly used metadata values within Hammerscript expressions, be it for objectives or while using the HSTK.
Many of the Hammerscript examples in this document leverage values such as SPACE_USED and LAST_USE_AGE. The former is self explanatory, while the latter represents the last time a file is closed and is the preferred value for measuring when a file was last used.
Additional metadata values are detailed in the Appendix A section Additional Metadata Objects And Functions. For a full list of metadata objects Hammerspace shares, refer to the File Metadata Example (Full) section of this document.
Timestamps
-
CREATE_AGE = (3 DAYS+17:02:20) - When a file was created
-
CHANGE_AGE = 2.50547 SECONDS - When a file was last modified, this timestamp also resets in response to a change in file permissions
-
ACCESS_AGE = 38.5719 SECONDS - When a file was last opened; note that clients may cache reads which can cause this value to be out of date
-
MODIFY_AGE = 38.5101 SECONDS - When a file was last modified, note that this timestamp does not reset if the file contents are not modified
File Details
-
NAME = "telemetry.txt" - The file name
-
PATH = "./mission-6/telemetry.txt" - The full path to the file, starts at the share root
-
OWNER = USER('jason@catalyst.local') - The object user owner
-
OWNER_GROUP = GROUP('Apollo_PrimeCrew@catalyst.local') - The object group owner
-
PARENT_SHARE = SHARE('gfstest') - The share the file is in
-
SPACE_USED = 4.096 KBYTES - The space used on disk, note that this may vary from SIZE as the default allocation unit is 4KB
-
SIZE = 73 BYTES - The file size as reported to the client, note that this does not reflect the actual space used on the volumes, nor the total space used by all file instances
-
IS_FILE = TRUE - TRUE indicates the object is a file, FALSE indicates it is either a symlink or directory
-
IS_SYMLINK = FALSE - TRUE indicates the object is a symlink, FALSE indicates it is either a file or a directory
-
IS_DIRECTORY = FALSE - TRUE indicates the object is a directory, FALSE indicates it is either a file or symlink
| Tags, Labels, and Keywords are also commonly used values, though the expressions used to match them are unique. Consult the Commands - Metadata Operations section of this document for an overview of how to write expressions that target these values. |
File Status
-
OVERALL_ALIGNMENT = ALIGNMENT('ALIGNED') - The file alignment status
-
IS_ONLINE = TRUE - TRUE indicates the file is available in a local online volume, FALSE indicates it is located on offline (object) storage or must be requested from another site
-
IS_OFFLINE = FALSE - TRUE indicates the file is not available on a local online volume, FALSE indicates it is located in a local online volume
-
DATA_ORIGIN_LOCAL = TRUE - TRUE indicates the current site was the last to edit the file and would be considered the owner, FALSE indicates another site is currently the owner
-
DATA_ORIGIN_REMOTE = FALSE - FALSE indicates the current site was the last to edit the file and would be considered the owner, TRUE indicates another site is currently the owner
-
ORIGIN_SITE = SITE('NYC-HS-A') - Indicates which site currently owns the file, meaning the file was last written to there
| Directories and 0 byte files are stored entirely in metadata, which will cause them to always show as online. |
Eval
The hs eval command allows us to evaluate Hammerscript expressions on a file, or simply return a specified metadata value. This can be used for a wide range of purposes such as reporting, testing objective conditions, or advanced searches.
Eval commands are not recursive by default. They will be performed on the current directory OR the path if one is provided. To make the command recursive, add the -r switch.
Best Practice: If you are looking for output suitable for later analysis in spreadsheet tools, hs eval is likely to require the least amount of effort to parse. That said, hs eval does not summarize output, but rather displays results for each file individually, so you may find that hs sum (described in the next section) is a more suitable command for some use cases. The same queries will often work for both commands, though it’s important to note that hs sum is recursive, so it is easy enough to try both and compare the output.
HS Eval Command Examples
In this section we will review some common hs eval commands. More examples are provided in the Advanced Filesystem Reporting section of this document.
List All File Instance Volume Locations for the Specified File
The following example command will list all cluster volumes that have instances of a file (telemetry.txt) using the default instances.volume expression:
# List all volumes with instances of the telemetry.txt file.
$ hs eval -e instances.volume telemetry.txt
INSTANCES_TABLE{
|VOLUME = STORAGE_VOLUME('Minio-NFS::/vols/NYC1');
|VOLUME = STORAGE_VOLUME('Minio-S3::gfs')}
Display All Alignment Details for a File
The following example command will provide a detailed output of the alignment status for a file (telemetry.txt) using the default alignment_details expression (partial output shown):
# List the file alignment details for the telemetry.txt file
$ hs eval -e alignment_details telemetry.txt
ALIGNMENT_DETAILS_TABLE{
|CHANGE_TIME = LOCAL_TIME('1969-12-31 19:00:00'),
|LAST_CHECKED_TIME = LOCAL_TIME('1969-12-31 19:00:00'),
|OVERALL = ALIGNMENT('PARTIALLY ALIGNED'),
|PLACE_ON = ALIGNMENT('PARTIALLY ALIGNED'),
|EXCLUDE_FROM = ALIGNMENT('ALIGNED'),
|CONFINE_TO = ALIGNMENT('ALIGNED'),
Note that this file is not fully aligned. Based on the output, it would seem that all applicable place-on objectives have not yet been met.
Display the Site That Currently Owns a File
Hammerspace global file systems use a term called owner to denote which site was the last one to write to a file. While every participant site has the ability to keep an online instance of the latest version of the file, only one of those sites is designated as an owner.
Features such as the optional log-xfer objective, referred to as File conflict resolution in the Hammerspace share GUI, use the owner value to determine if a file version needs to be created when site ownership changes.
THe following example command queries the metadata value DATA_ORIGIN_SITE.SITE_NAME to return the current owning site.
# Display the current origin site for file telemetry.txt
$ hs eval -e 'DATA_ORIGIN_SITE.SITE_NAME' telemetry.txt
"NYC-HS-A"
| In the Total Number of Files Owned By Each Global File Share Site section of this guide, we will show how to summarize ownership details for a collection of files. |
Sum
The hs sum command is used to perform fast calculations on a set of files. Typically, a Hammerscript expression used with hs eval can be used with hs sum, but note that hs sum is recursive while hs eval is not (unless specified).
The hs eval command differs from hs sum in that you can sum and sort data using the HSTK itself, rather than needing to import it into spreadsheet tools to perform that task. Plus, hs sum also grants you access to built in reports for the top 10, 100, or 1000 examples for supported metadata objects.
|
HS Sum Command Examples
In this section we will review some common hs sum commands. More examples are provided in the Advanced Filesystem Reporting section of this document.
Total Number of Files and Space Used - Specified Directory
The following command will return the file count and total amount of space used (\{1FILE/FILE,SPACE_USED/BYTES}) by all files within the specified directory (telemetry_dump).
# Return the file count and space used for all files in the telemetry_dump directory
$ hs sum -e 'IS_FILE?{1FILE/FILE,SPACE_USED/BYTES}' telemetry_dump
{4313, 17661952}
Commands like this can also use wildcards as a target (*). Note that if there were files and directories within the directory where the command was executed, both would be returned as shown in the following example where apollo_telemetry, telemetry_dump2, and telemetry_dump3 are directories.
# Return the file count and space used for all files and directories within the current directory
$ hs sum -e 'IS_FILE?{1FILE/FILE,SPACE_USED/BYTES}' *
##### apollo_telemetry
{51, 208896}
##### data_gemini.aaaa
{1, 4096}
##### data_gemini.aaab
{1, 4096}
##### telemetry_dump2
{5001, 20484096}
##### telemetry_dump3
{235, 962560}
In the previous example, we see output for both files and directories. Files are identified by the first value being a 1 (for 1 file), while directories that contain more than 1 file would return the count of files within and their total space used.
Best Practice - Leveraging Linux to Enhance Your HSTK Commands
You can use ordinary Linux commands to narrow the scope of some commands, like in this case with directories. The following examples will reuse the command we learned in the previous section, but you will see how easy it is to customize how the HSTK command is executed.
The following example uses a Linux ls -d command to list only directories, and then pipes that output to the hs sum command:
# Use ls -d */ | xargs to pipe directory names to HSTK for execution
$ ls -d */ | xargs hs sum -e 'IS_FILE?{1FILE/FILE,SPACE_USED/BYTES}'
##### apollo_telemetry
{51, 208896}
##### mission-6
{106, 425984}
##### mission_cancel
{4314, 17661952}
##### prj_apollo
{196, 42594}
. . .
Note our output now excludes files, so there is less cleanup you would need to do before analyzing this data further. This technique can be applied to similar situations where your HSTK target is directories, not files.
In this example, a recursive operation is enabled. The command will again only execute against directories, but we are using a Linux find . -type d command to do a recursive search for directories. The resulting output is again piped to the same hs sum command, and the output now includes subdirectories:
# Execute the hs sum command shown for all folders within the current path
$ find . -type d | xargs hs sum -e 'IS_FILE?{1FILE/FILE,SPACE_USED/BYTES}'
##### .
{14247, 113399520}
##### prj_apollo
{235, 962560}
##### mission-6
{106, 425984}
##### mission-6/launch_details
{53, 212992}
##### telemetry_dump3
{73, 55375584}
##### telemetry_dump3/archive
{15, 55142112}
While running this report against a very large directory report will take some time, it provides a very deep level of insight into where your files are located.
Best Practice: Use a recursion limit to control how deep into the directory tree the find command will traverse. The following example limits find to 3 levels (level 1 is the current directory).
# Set Linux find command to a max depth of 3
$ find . -maxdepth 3 -type d | xargs hs sum -e 'IS_FILE?{1FILE/FILE,SPACE_USED/BYTES}'
Combining Linux tools with HSTK commands provides you with the maximum flexibility when it comes to performing tasks or gathering the data that you need.
Total Number of Files and Space Used - Optional Top 10 Table
The following example would return the total number of files (not directories, symlinks, or devices) and space used. The second command also includes a top 10 table, while the third includes only the table:
# Return the file count and space used for the contents of the current directory
$ hs sum -e 'IS_FILE?{1FILE/FILE,SPACE_USED/BYTES}'
{9711, 39772160}
# Return the file count and space used for the contents of the current directory, plus a top 10 file list
$ hs sum -e 'IS_FILE?{1FILE/FILE,SPACE_USED/BYTES,TOP10_TABLE{{SPACE_USED/BYTES,PATH}}}'
{9711, 39772160, TOP10_TABLE{
|KEY = {12288, "./.DS_Store"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaiz"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaiy"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaix"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaiw"};
. . .
# Return the space used and file name with path for the top 10 files under the current directory
$ hs sum -e 'TOP10_TABLE{{SPACE_USED/BYTES,PATH}}'
TOP10_TABLE{
|KEY = {12288, "./.DS_Store"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaiz"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaiy"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaix"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaiw"};
. . .
|
Number of Files with the Specified Alignment Status
The following provides two different examples of commands that output stats on file alignment. The first example shows the total number of files that are aligned (overall_alignment=ALIGNMENT("ALIGNED")), while the second displays the count of files that are in any state other than aligned (overall_alignment!=ALIGNMENT("ALIGNED")).
# Return the count of aligned files for the content of the current directory
$ hs sum -e 'IS_FILE&&OVERALL_ALIGNMENT=ALIGNMENT("ALIGNED")' .
14247
0
# Display the count of files that are not in the aligned state (an ! was added)
$ hs sum -e 'IS_FILE&&OVERALL_ALIGNMENT!=ALIGNMENT("ALIGNED")' .
53
The possible alignment statuses include:
-
ALIGNED: file instances are all where they should be
-
MISALIGNED: all file instances are not where they should be
-
PARTIALLY ALIGNED: some file instances are where they should be
If you want a count of files per alignment status, you can simply use the default hs collsum command discussed in the next section.
File Instance Count and Space Used - per Volume
The following hs sum command will return a per cluster volume (KEY=INSTANCES[ROW].VOLUME) report with file instance count and space used (VALUE=\{1FILE/FILE,SPACE_USED/BYTES}):
# Return the file instance count and space used in bytes for each cluster volume
$ hs sum -e 'IS_FILE?ROWS(INSTANCES)?SUMS_TABLE{|::KEY=INSTANCES[ROW].VOLUME,|::VALUE={1FILE/FILE,SPACE_USED/BYTES}}[ROWS(INSTANCES)]'
SUMS_TABLE{
|KEY = STORAGE_VOLUME('NYC-dsx-1.catalyst.local::/hsvol0'),
|VALUE = {2425, 9932800};
|KEY = STORAGE_VOLUME('NYC-dsx-1.catalyst.local::/hsvol1'),
|VALUE = {2431, 9957376};
|KEY = STORAGE_VOLUME('NYC-dsx-2.catalyst.local::/hsvol0'),
|VALUE = {2429, 9957376};
|KEY = STORAGE_VOLUME('NYC-dsx-2.catalyst.local::/hsvol1'),
|VALUE = {2423, 9924608};
|KEY = STORAGE_VOLUME('Bucket-NYC'),
|VALUE = {9708, 39772160}}
|
You can return more values in the output by adding more metadata values within the brackets of the VALUE portion of the command (VALUE=\{1FILE/FILE,SPACE_USED/BYTES}); the following example would include a count of files that are considered a threat (IS_THREAT) by an external ICAP antivirus scan:
# Return the file instance count, count of files marked as a threat, and space used in bytes for each cluster volume
$ hs sum -e 'IS_FILE?ROWS(INSTANCES)?SUMS_TABLE{|::KEY=INSTANCES[ROW].VOLUME,|::VALUE={1FILE,IS_THREAT,SPACE_USED}}[ROWS(INSTANCES)]'
SUMS_TABLE{
|KEY = STORAGE_VOLUME('NYC-dsx-1.catalyst.local::/hsvol0'),
|VALUE = {2.425 KFILES, 35, 9.9328 MBYTES};
. . .
| The data is outputted in the order of the fields you provided. In this example, the IS_THREAT value is 35. |
File Instance Count, Total Space Used, and 10 Largest Files - per Volume
If you want the same report used in a previous example (file count and space used per volume), but also need a top 10 table detailing space used and file name and path (TOP10_TABLE\{\{SPACE_USED/BYTES,PATH}}) for each of the volumes, you can easily expand the command:
# Return the file instance count and space used in bytes for each cluster volume
$ hs sum -e 'IS_FILE?ROWS(INSTANCES)?SUMS_TABLE{|::KEY=INSTANCES[ROW].VOLUME,|::VALUE={1FILE/FILE,SPACE_USED/BYTES,TOP10_TABLE{{SPACE_USED/BYTES,PATH}}}}[ROWS(INSTANCES)]'
SUMS_TABLE{
|KEY = STORAGE_VOLUME('NYC-dsx-1.catalyst.local::/hsvol0'),
VALUE = {2425, 9932800, TOP10_TABLE{
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaiy"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaiw"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaip"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aail"};
|KEY = {4096, "./telemetry_dump3/telemetry_apollo.aaih"};
. . .
| You can also use TOP100 or TOP1000 if you need a longer list of files. |
Total Number of Files Owned by Each Global File Share Site
The Display The Site That Currently Owns A File section of this guide showed how to use hs eval to query a file to see which site currently owns it. In this example, we will summarize the count of files owned by each site in the global file share.
The following command will summarize the data for the current share recursively from the current folder:
# Return number of files owned for each site that is a member of the current global file share
$ hs sum -e 'IS_FILE?ROWS(DATA_ORIGIN_SITE.SITE_NAME)?SUMS_TABLE{|::KEY=DATA_ORIGIN_SITE[ROW].SITE_NAME,|::VALUE={1FILE/FILE}}[ROWS(SITE_NAME)]' .
SUMS_TABLE{
|KEY = "NYC-HS-A",
|VALUE = {34453};
|KEY = "LA-HS-A",
|VALUE = {14245}}
It is important to note that ownership can change at any moment. All that is required is for some other site to edit the file (opening a file does not change the owner). So while this report will give you some insight into where the latest versions of your files currently live, know that those stats can change frequently.
Collsum
The hs collsum command provides usage details about one or more of the default share collation summaries. While it is possible to create an unlimited number of custom summaries using hs eval or hs sum, the default collsums are automatically generated and cover many of the most common storage reports.
The following command leverages grep to obtain the standard list of collsums. A description of each is provided to the right of the name:
# List all standard collsums
$ hs collsum | grep SUMMATION
SUMMATION('BASIC'), {9.662 KFILES, 39.551 MBYTES, 0 OPERATIONS /
SECOND}; #Standard filesystem stats
SUMMATION('BY-HEAT'), { #File by access age and IOPS
SUMMATION('BY-ACCESS-AGE'), { #By access time
SUMMATION('BY-MODIFY-AGE'), { #By modify time
SUMMATION('BY-CHANGE-AGE'), { #By change time
SUMMATION('BY-CREATE-AGE'), { #By create time
SUMMATION('BY-SPACE-USED'), { #By space used
SUMMATION('BY-ACTUAL-MODIFY-AGE'), { #By modify time
SUMMATION('BY-VOLUME'), { #By instance location
SUMMATION('BY-ACTIVE-OBJECTIVE'), { #By active objective
SUMMATION('BY-ALIGNMENT'), INDEXED_TABLE{ #By alignment type
SUMMATION('BY-ALIGNMENT-STATE'), { #By alignment state
SUMMATION('BY-ERRORS'), { #By active error
SUMMATION('BY-TYPE'), { #File and directory count
SUMMATION('BY-VERSION'), { #Per snapshot version count
SUMMATION('BY-MIME'), {"/", {9.662 KFILES, 39.551 MBYTES, 0 OPERATIONS
/ SECOND}}; #Standard fs stats by mime type
SUMMATION('BY-VIRUS-SCAN'), {VIRUS_SCAN_STATE('UNSCANNED'), {9.662
KFILES, 39.551 MBYTES, 0 OPERATIONS / SECOND}}; #By scan state
SUMMATION('TOP-FILES'), TOP1000_TABLE{ #Top 100 files by size
Collsums are continually updated, and are the fastest method to obtain filesystem statistics. They will always return data for the entire share, regardless of where the command was actually executed from within the share.
Many collsums provide data similar to what you find in the Hammerspace GUI per-share dashboards, though if you are looking to prepare your own custom reports, the HSTK is the preferred method for obtaining this information.
Best Practice: Exploring the output of the various collsums is an easy way to learn about the different metadata values that you might want to use to create your own custom reports.
HS Collsum Examples
In this section we will review some of the available collsums. Many are self-explanatory, such as those based on space used or timestamps. If you are curious about ones not mentioned here, simply execute the command and observe what information is returned. If you execute just the hs collsum command, all summations for the share will be returned.
Basic
The basic collsum returns file instance count, total space used, and the IOPS.
# Use collsum to list the file instance count, total space used, and total IOPS
$ hs collsum --collation BASIC
INDEXED_TABLE{SUMMATION('BASIC'), {9.662 KFILES, 39.551 MBYTES, 0 OPERATIONS / SECOND}}
By-Alignment-State
The by-alignment-state collsum shows the number of and space consumed by files, broken down according to alignment. In the following example, you will see that a large number of files are waiting for a volume with sufficient space to place the files.
# Return the by-alignment-state collsum data
$ hs collsum --collation BY-ALIGNMENT-STATE
INDEXED_TABLE{SUMMATION('BY-ALIGNMENT-STATE'), {
ALIGNMENT_STATE('UNKNOWN'), {0 FILES, 0 BYTES, 0 OPERATIONS / SECOND};
ALIGNMENT_STATE('ALIGNED'), {7 FILES, 4.096 KBYTES, 0 OPERATIONS / SECOND};
ALIGNMENT_STATE('WAITING_FOR_MOBILITY'), {0 FILES, 0 BYTES, 0 OPERATIONS / SECOND};
ALIGNMENT_STATE('WAITING_FOR_TARGET_WITH_SPACE'), {9.655 KFILES, 39.547 MBYTES, 0 OPERATIONS / SECOND}}}
| In this example, the Hammerspace cluster has a volume with available space to write files to, but it is not the volume that meets the share objective requirements. Rather than prevent the client from writing data, the cluster places files wherever it can, marks them as unaligned (and why), and will monitor the cluster to see if space eventually becomes available on a volume that meets the objective requirements. |
By-Modify-Age
The following is a partial output of the by-modify-age collsum command. As with similar collsums, the full output extends as far as needed to break apart the values into smaller groups.
# Return the by-modify-age collsum data
$ hs collsum --collation BY-MODIFY-AGE
INDEXED_TABLE{SUMMATION('BY-MODIFY-AGE'), {
TIMESPAN_LEVEL('UNDER 1 MINUTE OLD'), {0 FILES, 0 BYTES, 0 OPERATIONS / SECOND};
TIMESPAN_LEVEL('1 TO 5 MINUTES OLD'), {0 FILES, 0 BYTES, 0 OPERATIONS / SECOND};
TIMESPAN_LEVEL('5 TO 15 MINUTES OLD'), {0 FILES, 0 BYTES, 0 OPERATIONS / SECOND};
TIMESPAN_LEVEL('15 TO 30 MINUTES OLD'), {0 FILES, 0 BYTES, 0 OPERATIONS / SECOND};
TIMESPAN_LEVEL('30 MINUTES TO 1 HOUR OLD'), {0 FILES, 0 BYTES, 0 OPERATIONS / SECOND};
TIMESPAN_LEVEL('1 TO 3 HOURS OLD'), {0 FILES, 0 BYTES, 0 OPERATIONS / SECOND};
TIMESPAN_LEVEL('3 TO 6 HOURS OLD'), {0 FILES, 0 BYTES, 0 OPERATIONS / SECOND};
Local-Objectives
The local-objectives collsum will return up to 1000 files that have objectives applied directly to them. Objectives applied in this way may be overriding those applied at the share root, so it is important to review where this has been done from time to time.
In the following example, we can see that the final line of the output starts the list of files that have objectives locally applied. Only 1 file is impacted in this example (log_mercury.aaab):
# Return the local-objectives collsum data
$ hs collsum local-objectives
INDEXED_TABLE{
SUMMATION('BASIC'), {1 FILE, 4.096 KBYTES, 0 OPERATIONS / SECOND};
SUMMATION('BY-HEAT'), {TEMPERATURE_LEVEL('1 TO 2 WEEKS OLD'), {1 FILE, 4.096 KBYTES, 0 OPERATIONS / SECOND}};
SUMMATION('BY-ACCESS-AGE'), {TIMESPAN_LEVEL('1 TO 2 WEEKS OLD'), {1 FILE, 4.096 KBYTES, 0 OPERATIONS / SECOND}};
SUMMATION('BY-MODIFY-AGE'), {TIMESPAN_LEVEL('1 TO 2 WEEKS OLD'), {1 FILE, 4.096 KBYTES, 0 OPERATIONS / SECOND}};
SUMMATION('BY-CHANGE-AGE'), {TIMESPAN_LEVEL('1 TO 2 WEEKS OLD'), {1 FILE, 4.096 KBYTES, 0 OPERATIONS / SECOND}};
SUMMATION('BY-CREATE-AGE'), {TIMESPAN_LEVEL('1 TO 2 WEEKS OLD'), {1 FILE, 4.096 KBYTES, 0 OPERATIONS / SECOND}};
SUMMATION('BY-SPACE-USED'), {SIZE_LEVEL('4 TO 32 KBYTES'), {1 FILE, 4.096 KBYTES, 0 OPERATIONS / SECOND}};
SUMMATION('TOP-FILES'), TOP1000_TABLE{|KEY = {4.096 KBYTES, "./log_mercury.aaab"}}}
Usage
The hs usage command shows information about file instance count, capacity consumption, volume usage, and more about files in a given path.
# Help page for hs usage
$ hs usage --help
Usage: hs usage [OPTIONS] COMMAND [ARGS]...
Options:
--help Show this message and exit.
Commands:
alignment Alignment state of files each file(s) of files in dir(s)
virus-scan Virus scan state of files each file(s) of files in dir(s)
owner Owner state of files each file(s) of files in dir(s)
online Summary of files on NAS volumes in the dir
volume Usage for each volume backing each dir(s)
user Users consuming the most capacity in each dir(s)
objectives Objectives applied and capacity managed by dir(s)
mime_tags All tags added by mime discovery on dir(s)
rekognition_tags All tags added by Rekognition on dir(s)
dirs Number of subdirectories under specified directory(ies), not including that directory
HS Usage Examples
In this section we will review some of the hs usage command options. Many are self explanatory.
If you are curious about ones not mentioned here, just execute the command and observe what information is returned.
| The hs usage command returns results based on where it was executed and is recursive. As such, it is a useful tool to get quick summary reports of the values it supports. |
Volume
The hs usage volume command gives the file instance count and space usage of the specified path (in this case the current directory) on each volume used:
# List file instances per cluster volume for the current directory/subdirectories
$ hs usage volume
SUMS_TABLE{
|KEY = #EMPTY,
|VALUE = 3;
|KEY = STORAGE_VOLUME('NYC-dsx-1.catalyst.local::/hsvol0'),
|VALUE = 2425;
|KEY = STORAGE_VOLUME('NYC-dsx-1.catalyst.local::/hsvol1'),
|VALUE = 2431;
|KEY = STORAGE_VOLUME('NYC-dsx-2.catalyst.local::/hsvol0'),
|VALUE = 2429;
|KEY = STORAGE_VOLUME('NYC-dsx-2.catalyst.local::/hsvol1'),
|VALUE = 2423;
|KEY = STORAGE_VOLUME('Bucket-NYC'),
|VALUE = 9708}
| In the above example, there are 3 files with an #EMPTY storage volume. These files have instances on another site in a global file system. The local site still knows about them, they are visible in the share, and users can still access them. Once a user attempts to open one of these files, a site to site mobility will occur and the files will now have a local instance. |
Always remember that the number of instances is not necessarily equal to the number of files. If your objectives cause multiple instances of a file to be created for data protection or other reasons, then each instance would show up in this report.
Owner
The hs usage owner command gives the file count by file owner within the specified path (in this case the current directory):
# List file count by file owner by user for the current directory/subdirectories
$ hs usage owner
SUMS_TABLE{
|KEY = USER('neil@catalyst.local'),
|VALUE = 9709;
|KEY = USER('buzz@catalyst.local'),
|VALUE = 2}
Alignment
The hs usage alignment command gives the file count by alignment state within the specified path (in this case, the current directory):
# List file count by alignment status for the current directory/subdirectories
$ hs usage alignment
SUMS_TABLE{
|KEY = ALIGNMENT('ALIGNED'),
|VALUE = 9711}
Perf
The hs perf command provides performance and operation stats on a per-share basis, and is out of scope for this document. Hammerspace recommends using the integrated Prometheus exporters to collect cluster performance statistics and display them using the custom Hammerspace dashboards for Grafana.
For information regarding the use of Grafana and Prometheus to gather Hammerspace cluster performance data, see Monitoring Cluster Health in the Hammerspace Administration Guide.
Advanced Filesystem Reporting
One of the most common reporting requests is for information regarding file ownership and capacity utilization. In this section, we will review some examples of HSTK commands used to create these reports.
When determining how to build your reports, you may wonder which of the commands that we have discussed is the best option.
For example, when using the same expression, the following commands will return the same data, but in two different formats:
-
The hs eval command will return the path, name, and size of all files that meet the condition specified. The command description will explain exactly what the condition text is looking for.
-
The hs sum command uses the same expression as hs-eval (though the output fields are modified: \{1FILE/FILES,SIZE/BYTES}), but instead returns a summary of the same information.
In the next section, we will review the results for various expressions. The only difference will be which command was used to perform the query.
Choosing Between HS Eval and HS Sum
The following three examples show the same Hammerscript expression being used with hs eval and then hs sum. When choosing which command, know that it will ultimately depend on the question you are trying to answer. That being said, it is usually quite easy to insert your expression into either one of those commands. While only hs sum can build the TOP tables, hs eval can do thorough reporting, and then analyze the data how you see fit using any number of tools.
The following examples show the difference in output between hs eval and hs sum for three different expressions.
# Return file name with path and size in bytes for all files less than 10kB
$ hs eval -r -e 'IS_FILE?SIZE<10KBYTES?{PATH,SIZE/BYTES}' .
{"./telemetry.txt", 535}
{"./log_mercury.aaaa", 1024}
{"./log_mercury.aaab", 1024}
{"./log_mercury.aaac", 1024}
{"./log_mercury.aaad", 1024}
. . .
# Return file count and total space used in bytes for all files less than 10kB
$ hs sum -e 'IS_FILE?SIZE<10KBYTES?{1FILE/FILES,SIZE/BYTES}' .
{4289, 4386138}
# Return file name with path and size in bytes for all files with the extension iso; update the value as needed to match other file extensions or names
$ hs eval -r -e 'IS_FILE?FNMATCH("*.iso",NAME)?{PATH,SIZE/BYTES}' .
{"./Rocky_v9.iso", 20}
{"./telemetry_dump/Boot_Recovery.iso", 15}
. . .
# Return file count and total space used in bytes for all files with the extension iso; update the value as needed to match other file extensions or names
$ hs sum -e 'IS_FILE?FNMATCH("*.iso",NAME)?{1FILE/FILES,SIZE/BYTES}' .
{2, 35}
# Return file name with path and size in bytes for all files contained within directories whose name matches tele* (case sensitive)
$ hs eval -r -e 'IS_FILE?FNMATCH("*/tele*/*",PATH)?{PATH,SIZE/BYTES}' . | more
{"./telemetry_dump3/telemetry_apollo.aaaa", 1024}
{"./telemetry_dump6/telemetry_mercury.aaab", 1024}
{"./telemetry_dump11/telemetry_gemini.aaac", 1024}
. . .
# Return file count and total space used in bytes for all files contained within directories whose name matches tele* (case sensitive)
$ hs sum -e 'IS_FILE?FNMATCH("*/tele*/*",PATH)?{1FILES/FILE,SIZE/BYTES}' .
{9525, 9743416}
Commands similar to this would make it easy to find those files which are consuming the most space.
| Remember that hs eval is not recursive by default. You must add the -r switch to make it scan all files and directories within your current path. The hs sum command is recursive, so the -r switch is not required. |
User Quota Reports
Hammerspace does not implement user or group quotas, though you can set a size limit on each share which functions as a share quota. However, Hammerspace does provide the ability to generate sophisticated usage reports as you might expect from a quota tool or system.
For these examples, we are mounted to a Hammerspace share, and the Python virtual environment is in the command search path. The commands work identically on Windows and Linux.
File Count and Total Space Used by File Owner
The command in the following example creates a simple quota report for all users with files under the current directory.
Note that unit conversions were not specified for files and space used, so the output for files and capacity use a variety of units (FILES vs KFILES, BYTES vs KBYTES vs GBYTES).
|
# Return file count and total space used by file owner
$ hs sum -e 'IS_FILE?SUMS_TABLE{|KEY={OWNER},|VALUE={1FILE,SPACE_USED}}'
SUMS_TABLE{
|KEY = {USER('neil@catalyst.local')},
|VALUE = {10.285 KFILES, 13.38 KBYTES};
|KEY = {USER('buzz@catalyst.local')},
|VALUE = {14.245 KFILES, 163.38 GBYTES};
|KEY = {USER('michael@catalyst.local')},
|VALUE = {9.115 KFILES, 113.38 KBYTES};
File Count and Total Space Used by File Owner per Volume
The following more sophisticated quota report breaks down each user’s (based on file owner) usage under the current directory, broken down by storage volume. This includes volumes on DSX and NFS storage, as well as those located on object and cloud storage.
# Return the file count and total space used on a per-volume basis for each user
$ hs sum -e 'IS_FILE?SUMS_TABLE{|::KEY={OWNER, OWNER_GROUP,INSTANCES[PARENT.ROW].VOLUME},|::VALUE={1FILE,SPACE_USED}}[ROWS(INSTANCES)]' .
SUMS_TABLE{
|KEY = {USER('neil@catalyst.local|neil@catalyst.local'), GROUP('Apollo_Astronauts@catalyst.local'), STORAGE_VOLUME('missionhq-dsx1.catalyst.local::/hsvol0')},
|VALUE = {13 FILES, 276.36 MBYTES};
|KEY = {USER('neil@catalyst.local|neil@catalyst.local'), GROUP('Apollo_Astronauts@catalyst.local'), STORAGE_VOLUME('missionhq-dsx2.catalyst.local::/hsvol0')},
|VALUE = {12 FILES, 287.88 MBYTES};
|KEY = {USER('neil@catalyst.local|neil@catalyst.local'), GROUP('Apollo_Astronauts@catalyst.local'), STORAGE_VOLUME('ntap-sim-7m::/vol/assim_vol_unix/q_uu3')},
|VALUE = {12 FILES, 12.792 MBYTES};
|KEY = {USER('neil@catalyst.local|neil@catalyst.local'), GROUP('Apollo_Astronauts@catalyst.local'), STORAGE_VOLUME('AWS::peter-demo2')},
|VALUE = {2 FILES, 268.51 MBYTES};
|KEY = {USER('neil@catalyst.local|neil@catalyst.local'), GROUP('Apollo_Astronauts@catalyst.local'), STORAGE_VOLUME('AWS::hspeterbucket1')},
|VALUE = {5 FILES, 319.56 MBYTES};
|KEY = {USER('neil@catalyst.local|neil@catalyst.local'), GROUP('Apollo_Astronauts@catalyst.local'), STORAGE_VOLUME('AWS::peter3')},
|VALUE = {3 FILES, 17.216 MBYTES};
|KEY = {USER('peter@catalyst.local|peter@catalyst.local'), GROUP('Apollo_Astronauts@catalyst.local'), STORAGE_VOLUME('missionhq-dsx2.catalyst.local::/hsvol0')},
|VALUE = {1 FILE, 4.096 KBYTES};
|KEY = {USER('ford@catalyst.local|ford@catalyst.local'), GROUP('Apollo_Astronauts@catalyst.local'), STORAGE_VOLUME('ntap-sim-7m::/vol/assim_vol_unix/q_uu3')},
|VALUE = {1 FILE, 0 BYTES};
|KEY = {USER('ford@catalyst.local|ford@catalyst.local'), GROUP('nerds@catalyst.local'), STORAGE_VOLUME('missionhq-dsx1.catalyst.local::/hsvol0')},
|VALUE = {1 FILE, 4.096 KBYTES};
General File and Directory Reporting
In this section, we will review a variety of reports used to report on file data for the entire cluster, find specific file types, or files that meet certain characteristics.
Top Files - All Cluster Shares
The following command, designed to be executed from the Hammerspace cluster root export, leverages a Linux ls command to pipe a list of share directories to a hs sum command.
This command returns the Top 10 files with for every share, their parent share, and the file size in bytes (\{PARENT_SHARE,SPACE_USED/BYTES,PATH}).
| A csv-formatted version of this report can be generated using the HSTK_Report_Builder.sh script. |
# Return the top 10 table by file size for all folders under the current path
$ ls -d */ | xargs hs sum -e 'TOP10_TABLE{{PARENT_SHARE,SPACE_USED/BYTES,PATH}}'
##### gfstest
TOP10_TABLE{
|KEY = {SHARE('gfstest'), 13783920, "./telemetry_dump3/archive/librocksdbjni568385507767029494.so"};
|KEY = {SHARE('gfstest'), 13783920, "./telemetry_dump3/archive/librocksdbjni1540166046508206384.so"};
|KEY = {SHARE('gfstest'), 13762560, "./telemetry_dump3/archive/librocksdbjni8354332257291138511.so"};
. . .
##### OnlineDelay
TOP10_TABLE{
|KEY = {SHARE('OnlineDelay'), 13721600, "./more.file7"};
|KEY = {SHARE('OnlineDelay'), 13721600, "./more.file6"};
|KEY = {SHARE('OnlineDelay'), 13721600, "./more.file5"};
. . .
File Count and Total Space Used by Directory
In this command, useful for quickly assessing a collection of user home directories, we will use find . -type d command to get a list of directories, and then pipe that list into a xargs hs sum command to obtain the file count and total space used for each directory (\{,1FILES/FILE,SPACE_USED/BYTES,}), including the current directory.
# Return the path, file count, and space used for every directory recursively
$ find . -type d | xargs hs sum -e 'IS_FILE?{,1FILES/FILE,SPACE_USED/BYTES,}'
##### .
{, 14055, 57544704, }
##### .Lost+Found
{, 1, 4096, }
##### .Lost+Found/missing_dir__ino=102605
#EMPTY
##### telemetry_dump
{, 4288, 17555456, }
##### telemetry_dump2
{, 5001, 20484096, }
##### telemetry_dump3
{, 235, 962560, }
##### apollo_telemetry
{, 51, 208896, }
##### mission-6
{, 106, 425984, }
##### mission-6/launch_details
{, 53, 212992, }
##### mission_cancel
{, 4314, 17661952, }
In the example above, we are combining ordinary Linux commands with HSTK as described in the Best Practice - Leveraging Linux To Enhance Your HSTK Commands section of this document. Note that an additional comma was added to the command output string, directly in front of 1FILES and after BYTES. This was intentional, and makes it easier to parse the data later. Parsing techniques are explained in greater detail in the Options For Processing HSTK Output of this document.
| A csv-formatted version of this report can be generated using the HSTK_Report_Builder.sh script. |
Matching Multiple File Name or Directory Path Values
While it is possible to string together multiple FNMATCH expressions into a single conditional statement, you can also build a summary expression as shown in the following example.
The format is fairly straightforward, you only need to edit a couple of options based on what you are trying to accomplish. Lets look at two examples, one for directory paths, another for files names:
SUM({||#A=FNMATCH(({"*.txt";"*.aaak";"*apollo*.*";"*.iso";"telemetry.*"})[ROW],NAME)}[5])==1
SUM({||#A=FNMATCH(({"*/*dump2*/*";"*/*apollo*/*"})[ROW],PATH)}[2])==0
The expressions should be updated as follows:
-
The quoted values inside the inner braces (for example
".txt"or"/dump2/“) are the file names or directory paths to match. For paths, you must include wildcards before and after the target path as shown (”/dump/*"). -
The
NAMEorPATHkeyword following[ROW],indicates whether you are matching a file NAME or a directory PATH. -
The count in square brackets near the end of the expression (
[5]and[2]in these examples) should equal the number of file or directory name examples provided within the expression. -
The final comparison value indicates either a positive (
==1) or negative (==0) match.
# Return file name with path and size in bytes for all files that match the 5 names provided
$ hs eval -r -e 'IS_FILE?SUM({||#A=FNMATCH(({"*.txt";"*.aaak";"*apollo*.*";"*.iso";"telemetry.*"})[ROW],NAME)}[5])==1?{PATH,SPACE_USED/BYTES}'
{"./log_mercury.aaak", 4096}
{"./sensor7_apollo.aaaa", 4096}
{"./1char.txt", 4096}
{"./new_output.txt", 4096}
{"./telemetry_dump3/0byte.txt", 0}
{"./newfile2.txt", 4096}
{"./telemetry_dump3/1char.txt", 4096}
{"./telemetry_dump3/log_mercury.aaak", 4096}
{"./telemetry_dump3/newfile2.txt", 4096}
{"./telemetry_dump3/Rocky_v9.iso", 4096}
. . .
# Return file name with path and size in bytes for all files that are not located in the paths provided
$ hs eval -r -e 'SUM({||#A=FNMATCH(({"*/*dump2*/*";"*/*apollo*/*"})[ROW],PATH)}[2])==0?{PATH,SPACE_USED/BYTES}' | more
{"./.Lost+Found", 0}
{"./prj_mercury", 0}
{"./prj_gemini", 0}
{"./telemetry_dump3/data_gemini.aaab", 4096}
{"./telemetry_dump3/data_gemini.aaac", 4096}
{"./mission_cancel/file.aaaj", 4096}
{"./mission_cancel/file.aaak", 4096}
{"./telemetry_dump/data_gemini.aaad", 4096}
{"./telemetry_dump/data_gemini.aaac", 4096}
{"./mission-6/launch_details/sensor7_apollo.aaal", 4096}
{"./mission-6/launch_details/sensor7_apollo.aaam", 4096}
. . .
| The equivalent hs sum command has not been provided, though the format is the same as the examples provided in the previous section. |
Find Files or Directories Based on Tag, Label, or Keyword
The following example command will return the path, name, and size of all files that have the TAG, TAG with value, LABEL, or KEYWORD indicated. Consult the Metadata Operations section of this document for examples of how to apply and remove these from files and directories.
In the examples provided below, note the first two commands that query for a tag (HAS_TAG("mission_specs")), and then for a tag and value (GET_TAG("mission_specs")=="launch"). You will see that all files with that tag are returned for the first query, but the second query returns fewer files. This indicates that not all files with the tag have the same tag value.
# Return file name with path and size for all files that have the tag mission_specs
$ hs eval -r -e 'IS_FILE?HAS_TAG("mission_specs")?{PATH,SIZE}' .
{"./telemetry_dump/file.acsl", 1.024 KBYTES}
{"./telemetry_dump/file.adrw", 1.024 KBYTES}
{"./telemetry_dump/telemetry.txt", 11 BYTES}
{"./telemetry_dump/qa_summary.txt", 11 BYTES}
# Return file name with path and size for all files that have the tag mission_specs with value launch
$ hs eval -r -e 'IS_FILE?GET_TAG("mission_specs")=="launch"?{PATH,SIZE}' .
{"./telemetry_dump/file.acsl", 1.024 KBYTES}
{"./telemetry_dump/file.adrw", 1.024 KBYTES}
# Return file name with path and size for all files that have the label launch_report
$ hs eval -r -e 'IS_FILE? HAS_LABEL("launch_report")?{PATH,SIZE}' .
{"./telemetry_dump/file.aalr", 1.024 KBYTES}
{"./telemetry_dump/file.aecq", 1.024 KBYTES}
# Return file name with path and size for all files that have the keyword mission_prep
$ hs eval -r -e 'IS_FILE? HAS_KEYWORD("mission_prep")?{PATH,SIZE}' .
{"./telemetry_dump/file.aeop", 1.024 KBYTES}
{"./telemetry_dump/file.afym", 1.024 KBYTES}
|
Find Files Based on Usage
The following example will find all files that have been accessed in the last 7 days, and list the number of those files and space used.
# Return the total file count and space used for files used within the last 7 days
$ hs sum -e 'IS_FILE&&ACCESS_AGE<=7DAYS?{1FILE,SPACE_USED}' .
{264 FILES, 56.137 MBYTES}
The following example will find all files accessed more than 2 days ago within the mission-6 folder, list the number of those files, space used, and print the top 10 table, which lists the 10 largest files.
# Return the file count and space used for files in the mission-6 folder that were used 2 days or more ago, include a top 10 table of files by space used
$ hs sum -e 'IS_FILE&&ACCESS_AGE>=2DAYS?{1FILE,SPACE_USED,TOP10_TABLE{{SPACE_USED,DPATH}}}' mission-6
{106 FILES, 425.98 KBYTES, TOP10_TABLE{
|KEY = {4.096 KBYTES, "./mission-6/telemetry.txt"};
|KEY = {4.096 KBYTES, "./mission-6/sensor7_apollo.aaao"};
|KEY = {4.096 KBYTES, "./mission-6/sensor7_apollo.aaan"};
|KEY = {4.096 KBYTES, "./mission-6/sensor7_apollo.aaam"};
Commands - Filesystem Operations
The HSTK provides many useful options for performing offloaded filesystem operations. If you have a need to perform one of the commands that it supports, using it to run the command offloads the operation to the cluster itself, rather than requiring the client to perform the task.By offloading typically resource intensive filesystem operations, you eliminate the client and network overhead they typically require.
CP
The hs cp command performs a fast offloaded recursive copy using a clone operation. The following example shows how it is used:
# Perform an offloaded clone of file mission-template.db
$ hs cp -a mission-template.db mission.apollo.db
| The clone operation is only offloaded for online cluster volumes that support offloaded cloning, such as DSX volumes. While Hammerspace clusters can use external NFS storage as an online volume, not all NFS servers support offloaded cloning. |
RM
The hs rm command is a multithreaded version of the Linux rm -rf command. This command performs the operation directly on the Hammerspace cluster, which provides the most efficient way to delete large directories from Hammerspace shares.
The command should be executed using the same syntax as the Linux equivalent (rm -rf) as shown in the following example.
# Perform a recursive delete of the mission_cancel directory
$ hs rm -rf mission_cancel
{
"status":"PASSED",
"dirents_found":55,
"started":168722881205540,
"finished":168723202884660,
"rate":341.314503
}
RSYNC
The hs rsync command is an offloaded version of the Linux recursive directory equalizer. Note that like rsync, when updating a directory, files are both added and removed based on the state of the source directory when the command was run. By offloading the operation, this command ensures the best performance when syncing data from within locations on a Hammerspace cluster.
The following is an example of the command syntax, note that the --delete and --archive options are required:
# Rsync the contents of the telemetry_dump directory into the mission_cancel directory
$ hs rsync --delete --archive telemetry_dump/ mission_cancel
{"assim_id":3}
| The rsync command uses the Hammerspace assimilation feature to copy process requests, which explains the status returned of "assim_id":3. |
RM-ST
The hs rm-st is identical to the hs rm command, though it is not multi-threaded. The command syntax and options are identical. Refer to the RM section of this document for an example of how the command is used.
Commands - General
The HSTK provides a number of commands that can be used to provide the status of various aspects of the cluster and filesystem. In this section, we will review these various commands.
Status
The hs status command provides views into the status of various processes, tasks, and objects in the Hammerspace environment.
# Help page for hs status
$ hs status --help
Usage: hs status [OPTIONS] COMMAND [ARGS]...
[sub] System, component, task status
Options:
--help Show this message and exit.
Commands:
assimilation State of current assimilations
csi Details about the kubernetes CSI
collections Collections present in the share
errors Files in the share with errors
open Files open each dir(s)
replication Replication progress for the share(s)
sweeper Progress of sweeper (checks file placement) for each...
volume Health of volumes backing the share(s)
These commands generally require no options. Simply execute them from within a Hammerspace share similar to the following example (partial command output shown):
# Show the replication status for the current share
$ hs status replication
REPLICATION_DETAILS_TABLE{
|SITE = SITE('LA-HS-A'),
|INTERVAL = 5 SECONDS,
|SEND_TIME = LOCAL_TIME('2024-11-19 14:53:58'),
|RECV_TIME = LOCAL_TIME('2024-11-19 14:53:50'),
SNAPS = REPLICATION_DETAILS_SUB_TABLE{
|SENT_REQUESTS = 15732,
|SENT_BYTES = 0 BYTES,
|SUCCEEDED_REQUESTS = 15732,
|SUCCEEDED_BYTES = 0 BYTES,
|ERRORED_REQUESTS = 0,
. . .
# Show the status of all cluster volumes
$ hs status volume
{
"NYC-dsx-2.catalyst.local::/hsvol0", STORAGE_VOLUME_STATUS('REMOVED'), OPERATIONAL_STATUS('DOWN');
"NYC-dsx-2.catalyst.local::/hsvol1", STORAGE_VOLUME_STATUS('REMOVED'), OPERATIONAL_STATUS('DOWN');
"NYC-dsx-1.catalyst.local::/hsvol0", STORAGE_VOLUME_STATUS('REMOVED'), OPERATIONAL_STATUS('DOWN');
"NYC-dsx-1.catalyst.local::/hsvol1", STORAGE_VOLUME_STATUS('REMOVED'), OPERATIONAL_STATUS('DOWN');
"Minio-NFS::/vols/NYC1", STORAGE_VOLUME_STATUS('OK'), OPERATIONAL_STATUS('UP');
"NYC-dsx-1.catalyst.local::/hsvol0", STORAGE_VOLUME_STATUS('OK'), OPERATIONAL_STATUS('UP');
"NYC-dsx-1.catalyst.local::/hsvol1", STORAGE_VOLUME_STATUS('OK'), OPERATIONAL_STATUS('UP');
"NYC-dsx-2.catalyst.local::/hsvol0", STORAGE_VOLUME_STATUS('OK'), OPERATIONAL_STATUS('UP');
"NYC-dsx-2.catalyst.local::/hsvol1", STORAGE_VOLUME_STATUS('OK'), OPERATIONAL_STATUS('UP');
"Bucket-NYC", STORAGE_VOLUME_STATUS('OK'), OPERATIONAL_STATUS('UP');
"Bucket-LA", STORAGE_VOLUME_STATUS('OK'), OPERATIONAL_STATUS('UP');
. . .
Objective
In this section, we will review the HSTK commands used to list and manage objectives.
Objectives are defined by the administrator using the Hammerspace CLI, GUI, or API. However, they can be applied to files and directories (with inheritance) by users with the HTSK. The HSTK also provides a way to list the inherited and directly applied objectives on files and directories.
# Help page for hs objective
$ hs objective --help
Usage: hs objective [OPTIONS] COMMAND [ARGS]...
Options:
--help Show this message and exit.
Commands:
list list all (objective,expression) pairs assigned
has Get/list objective assignments
delete remove (objective,expression) pair from inode(s)
add Add (objective,expression) pair to inode(s)
Best Practice:
-
If you intend to apply objectives directly to individual files or directories, you should determine if there are any objectives applied directly to the share that you don’t want overridden.
-
If an upstream objectives expression is set to TRUE, a downstream objective (like one applied directly to a file) could potentially override it. If you want to prevent that from happening, ensure the critical objective expression is set to ALWAYS.
-
For more information regarding Hammerspace objectives, including their design and how they are applied, refer to the How to Configure Hammerspace Objectives [Support article may require login to view.].
List Inherited or Directly Applied Objectives
The base hs objective list command will list objectives that are directly applied or inherited to the specified file or directory. Note that just because an objective is listed, it does not mean it actually applies at the time the command is run.
# List all objectives that are applied to the file telemetry.txt
$ hs objective list telemetry.txt
SLOS_TABLE{
|OBJECTIVE = SLO('durability-1-nine'),
|COUNT = EXPRESSION(DATA_ORIGIN_LOCAL?ALWAYS);
|OBJECTIVE = SLO('durability-3-nines'),
|COUNT = EXPRESSION(DATA_ORIGIN_LOCAL AND IS_DURABLE);
|OBJECTIVE = SLO('availability-1-nine'),
|COUNT = EXPRESSION(DATA_ORIGIN_LOCAL?ALWAYS);
|OBJECTIVE = SLO('keep-online'),
|COUNT = EXPRESSION(IS_BEING_CREATED OR HAS_KEEP_ON_THIS_SITE);
|OBJECTIVE = SLO('optimize-for-capacity'),
|COUNT = TRUE;
|OBJECTIVE = SLO('delegate-on-open'),
|COUNT = TRUE;
|OBJECTIVE = SLO('compress-on-object'),
|COUNT = TRUE;
|OBJECTIVE = SLO('content-based-chunk-on-object'),
|COUNT = TRUE;
|OBJECTIVE = SLO('sync-metadata'),
|COUNT = TRUE;
|OBJECTIVE = SLO('keep-online'),
|COUNT = EXPRESSION(IS_BEING_CREATED OR HAS_ONLINE_INSTANCE AND IS_RECENTLY_USED?ALWAYS);
|OBJECTIVE = SLO('layout-get-on-open'),
|COUNT = EXPRESSION(IS_BEING_CREATED OR HAS_ONLINE_INSTANCE)}
Use the --effective switch to list objectives that (based on the conditions) actually apply to the file, or use --active lists those currently affecting the file in some way. Active objectives are subject to change over time as the state of the file changes, which is why the active and effective objectives do not always match.
# List all effective objectives for file telemetry.txt
$ hs objective list --effective telemetry.txt
{
SLO('keep-online'), 0;
SLO('durability-1-nine'), ALWAYS;
SLO('compress-on-object'), 1;
SLO('content-based-chunk-on-object'), 1;
SLO('sync-metadata'), 1;
SLO('optimize-for-capacity'), 1;
SLO('place-on-Bucket-NYC'), 1;
SLO('versioning-1-day'), 1;
SLO('undelete-1-day'), 1;
SLO('place-on-DSX'), 1}
# List all active objectives for telemetry.txt
$ hs objective list --active telemetry.txt
{
SLO('durability-1-nine');
SLO('compress-on-object');
SLO('content-based-chunk-on-object');
SLO('sync-metadata');
SLO('optimize-for-capacity');
SLO('place-on-Bucket-NYC');
SLO('versioning-1-day');
SLO('undelete-1-day');
SLO('place-on-DSX')}
In the above example, we see that the keep-online objective is effective but not active. The file is currently online so no action needs to be taken.
Add or Delete Directory or File Objectives
The hs objective add command can be used to add objectives to files or directories, and can be run recursively. Alternatively, the hs objective delete command can remove them.
# Help page for hs objective add
$ hs objective add --help
Usage: hs objective add [OPTIONS] NAME paths
Add (objective,expression) pair to inode(s)
Options:
-r, --recursive Apply recursively
--nonfiles Apply recursively to non files
-j, --json Expect --exp to be JSON input
-i, --exp-stdin Read expression from stdin
-e, --exp TEXT Value as hammerscript expression
-s, --string TEXT Value as a string
--help Show this message and exit.
The following example shows an objective being added to a file, the objectives checked to verify the objective was added, and then the objective is deleted.
# Add the keep-online objective to the telemetry.txt file
$ hs objective add keep-online telemetry.txt
# List the current objectives applied to the telemetry.txt file
$ hs objective list telemetry.txt
SLOS_TABLE{
|OBJECTIVE = SLO('keep-online'),
|COUNT = TRUE;
. . .
# Delete the keep-online objective from the telemetry.txt file
$ hs objective delete keep-online telemetry.txt
| If you need a list of objectives that you can apply, use the hs dump objectives command detailed in the next section. |
You can also add objectives with expressions, which act as conditions. This example shows two different commonly used objective expressions, ALWAYS and TRUE. ALWAYS indicates that the objective cannot be overridden by other objectives, while TRUE indicates that it can.
That being said, if an objective applied at the share root is marked as ALWAYS, it cannot be overridden by ones that are directly applied to files or directories within.
# Add the keep-online objective to the telemetry.txt file, effective always
$ hs objective add -e 'always' keep-online telemetry.txt
# List the current objectives applied to the telemetry.txt file
$ hs objective list telemetry.txt
SLOS_TABLE{
|OBJECTIVE = SLO('keep-online'),
|COUNT = EXPRESSION(ALWAYS);
. . .
# Add the keep-online objective to the newout.txt file, effective true
$ hs objective add -e 'true' keep-online newout.txt
# List the current objectives applied to the newout.txt file
$ hs objective list newout.txt
SLOS_TABLE{
|OBJECTIVE = SLO('keep-online'),
|COUNT = TRUE;
. . .
Dump
The hs dump command provides a number of subcommands used to obtain information about the cluster or the files and directories that it contains. The information obtained from subcommands such as misaligned, threat, files_on_volume can also be obtained using hs sum or hs eval commands, while others such as inode are typically only used by Hammerspace support.
A common use case for the hs dump command is to list the objectives whose names you may need to know. The hs dump objectives command lists all objectives present on the cluster. The following example is a shortened version of the command output.
# List all objectives available on this Hammerspace cluster
$ hs dump objectives
performance-fast
performance-medium
performance-slow
keep-online
do-not-move
optimize-for-capacity
delegate-on-open
Remember that objectives are created per-cluster. While some are created by default, all others are created when you add volumes, create volume groups, or create your own custom objectives.
| This document focuses primarily on metadata operations and reporting, so we will not go into further detail regarding the hs dump command. The following hs dump help page provides details about the various subcommands. |
# Help page for hs dump
$ hs dump --help
Usage: hs dump [OPTIONS] COMMAND [ARGS]...
[sub] Dump info about various items
Options:
--help Show this message and exit.
Commands:
inode inode metadata
iinfo Alternative inode details, always in JSON format
share Full share(s) metadata
misaligned Dump details about misaligned files on the share(s)
threat Dump details about files that are a virus threat on the...
map_file_to_obj For --native object volumes, dump a mapping between file...
files_on_volume List all files that have data on the specified volume
per...
volumes List available volumes in the cluster