Search the docs

Architecture

Servers

Tier 0 Node

A Tier 0 node consists of a server platform with CPU, memory, disk, and networking as the baseline, with some specialized or additional compute which may be in the form of GPUs or extra CPU, along with additional storage beyond the boot device. It also often includes additional networking, some of which may be dedicated or intended for node to node compute related communication.

tier0 tier 0 node image1
Figure 1. Tier 0 node

Linux Storage Server

LSS typically has more storage devices than a Tier 0 node (24, 32, or more depending on the system selected). An LSS can also provide additional services such as data movers and quorum. An LSS is typically not configured as a client, except as necessary for services it provides.

Networking

Tier 0 deployments should use recent high speed networking technology supporting 200 - 800Gbps speeds.

Flat Networks

Hammerspace recommends flat networks as much as practicable in Tier 0 environments. The goal is to minimize network hops between nodes. With multiple availability zones, each zone could be a flat network (subnet).

Jumbo Frames (MTU)

  • It is recommended to use the customer’s typical MTU (Maximum Transmission Unit) settings. MTU values should generally be in the range of 1500 to 9000; anything outside of this range should be discussed.

  • Switches should typically be set to an MTU of 9216.

  • Modern switches often ship with Jumbo frames enabled by default. While Jumbo frames may not significantly increase wire speed, they can reduce CPU consumption and interrupt handling due to fewer packets.

  • Test network connectivity with ping using a specific packet size (e.g., 8972 for a 9000 MTU) and the "do not fragment" flag is advised to confirm end-to-end connectivity.

  • General Recommendation (Tier 0): If used on Tier 0 nodes, bonding must be configured carefully due to the NUMA architecture of some server platforms. On some dual socket platforms, it is critical that bonds consist of interfaces all in the same NUMA domain to avoid a performance penalty due to I/O traversing domains. If you plan to use RDMA, bonds must consist of ports on the same physical NIC. In many cases, using an IP per NIC may be simpler and more reliable. Bonding is supported on Hammerspace Anvil and DSX nodes (bonding mode 4, also known as 802.3ad). This requires configuration of port channels with LACP of the ports on the connected switches.

  • Packet Distribution: The default packet distribution algorithm in Linux (referred to as xmit_hash_policy) is typically layer2 or layer2+3. Depending on the network topology, you may see better packet outbound packet distribution using layer3+4. These parameters are described in detail in RedHat’s documentation.

You can change the xmit_hash_policy by editing the BONDING_OPTS= parameter in /etc/sysconfig/network-scripts/ifcfg-bond0 (or whatever the bond number is).
  • Caveat: Some customer environments (e.g., cloud, pre-existing infrastructure) may require bonding for redundancy or throughput. In such cases, professional services engagement is necessary to ensure proper configuration.

IP per NIC

  • Recommendation: For systems with multiple NICs, assign a unique IP address to each NIC.

  • Rationale: Provides multiple independent network paths, improving parallelism and resilience.

  • Configuration: Requires specific sysctl and network configuration to ensure proper routing and load balancing without introducing routing conflicts.

Availability Zones

An availability zone is a part of an IT infrastructure system that is designed to be completely independent of other availability zones. This means it doesn’t share any critical components, like power, cooling, or network, with other zones.

Components to consider may include:

  • Power, whether circuit, rail, panel or similar

  • Cooling

  • Fire suppression zones

  • Separate rooms or floors within a datacenter

  • Racks or rows

  • Network equipment and architecture

  • Security access zones

While many discussions focus on cloud provider AZs, in this context, we focus on AZs within a datacenter. Hammerspace supports defining AZs for storage resources using a naming convention with a prefix indicating the AZ.

Hammerspace recommends that availability zones be designed approximately symmetrical in that they each have the same number of storage nodes (Tier 0 or LSS) and volumes. You should also place redundant Hammerspace nodes, including Anvil, DSX (if used), and mover nodes in different AZs. If Anvil nodes are in different AZs, and each AZ is a different subnet (or set of subnets), you will need to implement the ExaBGP (Border Gateway Protocol service) on the Anvils and configure the switches to allow Anvils to trigger the movement of the subnet containing the cluster management IP to the switch port to which the primary Anvil is connected.

When data is stored only on Tier 0, a minimum of 4 AZs is required, and 6 AZs is recommended. With 4 AZs, one can be taken offline for maintenance, leaving 3 providing full 3-way redundancy. But that doesn’t allow for multiple failures during maintenance. Having additional AZs (5 or 6 total) allows for failures with automatic resilvering across three surviving AZs when an AZ is offline for maintenance.