Search the docs

Appendix 2 - Use Tier 0 for Checkpointing

Checkpointing Workflow Example

  1. Checkpoint Initiation: The application triggers a checkpoint, at which point each node creates a file on the Hammerspace global shared storage system. Based on Service Level Objectives set up in Hammerspace, the storage node selected to back each file is the NVMe that is local to each node, thus making it possible to bypass the NFS protocol and networking stack as described above.

  2. Local Write Completion: Once the write is complete, the GPU resumes computation without delay.

  3. Asynchronous Replication: Hammerspace detects the new checkpoint file and begins replicating it to designated storage tiers based on policies.

  4. Data Availability: The checkpoint is now safely stored in multiple locations, ensuring it can be used for recovery if needed.

Tier 0 Checkpoint Analysis

This analysis quantifies the benefits of using Tier 0 storage. Specifically, leveraging local NVMe storage within compute nodes—for checkpointing, compared to traditional methods that rely on external networked storage systems, even those connected via high-speed 800Gb Ethernet (800GbE).

We will use the NVIDIA A100 GPU for our calculations, consider a 1,000-node cluster with 8,000 GPUs, 100 Petabytes of storage capacity, and 1 TB/sec of aggregate throughput.

Estimating Checkpoint Data Size

System Configuration:

  • Compute Node: NVIDIA DGX A100 or HGX system

  • GPUs per Node: 8 NVIDIA A100 GPUs

  • GPU Memory per GPU: 80 GB (also available in 40 GB variants)

  • Total GPU Memory per Node: 8 GPUs × 80 GB = 640 GB

  • CPU Memory: Assume 256 GB per node (can vary)

  • Total Memory to Checkpoint: GPU Memory + Relevant CPU Memory

Example Checkpoint Data Size Estimate
600 GB

Estimating Checkpoint Time When Writing to External Storage

Considerations when Writing Checkpoints to External Storage:

  • Network Overheads: Protocol overheads, congestion, and latency reduce effective bandwidth.

  • Shared Infrastructure: Multiple nodes checkpointing simultaneously can saturate the network and storage array.

  • Effective Bandwidth per Node: Often significantly less than theoretical maximum. Let’s conservatively estimate 1 GB/s per node.

Example Time to write a 600 GB checkpoint to external storage
600 GB/1 GB per second = 600 seconds

Estimating Checkpoint Time When Writing to Tier 0 Local NVMe Storage

Local NVMe Storage Specifications:

  • NVMe Devices per Node: 8 NVMe drives

  • NVMe Interface: PCIe Gen5

  • Bandwidth per NVMe Device: Approximately 14 GB/s (real-world write performance)

  • Total Aggregate Bandwidth: 112 GB/s per node

  • Effective Bandwidth Assuming 90% Efficiency: 100.8 GB/s per node

Example Time to write a 600 GB checkpoint to Tier 0 storage
600 GB/100.8 GB per second = 5.95 seconds (6)