Date of Award
2026-08-01
Degree Name
Master of Science
Department
Computer Science
Advisor(s)
Deepak K. Tosh
Abstract
Modern deep learning workloads increasingly rely on distributed computation, and High Performance Computing systems can provide the necessary resources through large GPU allocations across interconnected nodes. Despite this, most distributed training frameworks operate under static resource assignments once a job is deployed. Research on NERSC Perlmutter has shown that 50% of GPU-enabled jobs use 25% or less of available GPU memory, and elastic training can reduce this underutilization by dynamically adjusting active workers. However, existing elastic systems have been developed mainly for cloud environments where fault tolerance and cost optimization are the primary concerns. Applying elastic training to HPC environments as a tool for intelligent resource sharing and analyzing its accuracy impact on machine learning models has not been addressed in existing literature.
This thesis presents an elastic distributed training framework built on Elastic Horovodand validated on NERSC Perlmutter. HERD enables dynamic GPU reallocation across concurrent training jobs through a centralized TCP server. A pluggable policy architecture separates scaling logic from elastic infrastructure. Two default policies are included: a memory-based policy using aggregate GPU memory utilization as a scaling trigger, and an elastic resizing policy for controlled evaluation. HERD is deployed on eight nodes and validated through concurrent elastic workloads sharing a fixed GPU resource pool.
Experimental results demonstrate that elastic GPU reallocation during training does not introduce accuracy degradation. The maximum accuracy gap across all configurations is 4.6%, with elastic downscaling achieving the highest accuracy. Scaling direction predicts generalization behavior through the large-batch generalization effect. Downscaling reduces the effective batch size and narrows the generalization gap, while upscaling increases it. GPU memory utilization was also confirmed as a reliable per-epoch scaling signal. These findings show that elastic distributed training on HPC systems is technically feasible, does not compromise model accuracy, and reduces GPU underutilization through elastic resizing.
Language
en
Provenance
Received from ProQuest
Copyright Date
2026-08
File Size
90 p.
File Format
application/pdf
Rights Holder
Alejandro Guerrero Rodriguez
Recommended Citation
Guerrero Rodriguez, Alejandro, "HERD: A Policy-Driven Elastic Resource Distribution Framework for HPC Deep Learning" (2026). Open Access Theses & Dissertations. 4693.
https://scholarworks.utep.edu/open_etd/4693