Methodology for GPU Frequency Switching Latency Measurement
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61989100%3A27740%2F25%3A10259702" target="_blank" >RIV/61989100:27740/25:10259702 - isvavai.cz</a>
Výsledek na webu
<a href="https://ieeexplore.ieee.org/document/11105888" target="_blank" >https://ieeexplore.ieee.org/document/11105888</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/IPDPSW66978.2025.00133" target="_blank" >10.1109/IPDPSW66978.2025.00133</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Methodology for GPU Frequency Switching Latency Measurement
Popis výsledku v původním jazyce
The push towards exascale and post-exascale computing in HPC and AI brings together thousands of CPUs and specialized accelerator hardware, making energy optimization crucial as power costs rival system purchase prices. Energy efficiency techniques based on frequency and voltage scaling have been developed and fine-tuned for CPUs, which led to deep understanding of how the CPU hardware behaves under frequency adjustments. In contrast, accelerators, particularly GPUs, have not yet been studied to the same extent in this context.We introduce a methodology to evaluate the latency coupled with accelerator frequency scaling driven by the control CPU (GPU switching latency). The approach employs a minimal, iterative workload that allows statistically distinguishing runtime differences between frequency pairs. It first measures execution times for each frequency and then determines the latency of switching from an initial to a target frequency by tracking runtime changes and repeating measurements to ensure statistical robustness. Finally, the methodology filters out outliers from external factors like driver management or system interruptions. The methodology is implemented in the tool LATEST with support for CUDA accelerators. Evaluated on three Nvidia GPUs - GH200, A100-SXM4, and RTX Quadro 6000 - the analysis reveals significant differences in the switching latency, evaluates optimal frequency change rates, and identifies frequency pairs to avoid due to high overhead. © 2025 Elsevier B.V., All rights reserved.
Název v anglickém jazyce
Methodology for GPU Frequency Switching Latency Measurement
Popis výsledku anglicky
The push towards exascale and post-exascale computing in HPC and AI brings together thousands of CPUs and specialized accelerator hardware, making energy optimization crucial as power costs rival system purchase prices. Energy efficiency techniques based on frequency and voltage scaling have been developed and fine-tuned for CPUs, which led to deep understanding of how the CPU hardware behaves under frequency adjustments. In contrast, accelerators, particularly GPUs, have not yet been studied to the same extent in this context.We introduce a methodology to evaluate the latency coupled with accelerator frequency scaling driven by the control CPU (GPU switching latency). The approach employs a minimal, iterative workload that allows statistically distinguishing runtime differences between frequency pairs. It first measures execution times for each frequency and then determines the latency of switching from an initial to a target frequency by tracking runtime changes and repeating measurements to ensure statistical robustness. Finally, the methodology filters out outliers from external factors like driver management or system interruptions. The methodology is implemented in the tool LATEST with support for CUDA accelerators. Evaluated on three Nvidia GPUs - GH200, A100-SXM4, and RTX Quadro 6000 - the analysis reveals significant differences in the switching latency, evaluates optimal frequency change rates, and identifies frequency pairs to avoid due to high overhead. © 2025 Elsevier B.V., All rights reserved.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
2025 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025 : proceedings : 3-7 June 2025 Milan, Italy
ISBN
979-8-3315-2643-6
ISSN
2639-3867
e-ISSN
2995-066X
Počet stran výsledku
10
Strana od-do
830-839
Název nakladatele
IEEE
Místo vydání
Piscataway
Místo konání akce
Milán
Datum konání akce
3. 6. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—