All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Methodology for GPU Frequency Switching Latency Measurement

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61989100%3A27740%2F25%3A10259702" target="_blank" >RIV/61989100:27740/25:10259702 - isvavai.cz</a>

  • Result on the web

    <a href="https://ieeexplore.ieee.org/document/11105888" target="_blank" >https://ieeexplore.ieee.org/document/11105888</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1109/IPDPSW66978.2025.00133" target="_blank" >10.1109/IPDPSW66978.2025.00133</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Methodology for GPU Frequency Switching Latency Measurement

  • Original language description

    The push towards exascale and post-exascale computing in HPC and AI brings together thousands of CPUs and specialized accelerator hardware, making energy optimization crucial as power costs rival system purchase prices. Energy efficiency techniques based on frequency and voltage scaling have been developed and fine-tuned for CPUs, which led to deep understanding of how the CPU hardware behaves under frequency adjustments. In contrast, accelerators, particularly GPUs, have not yet been studied to the same extent in this context.We introduce a methodology to evaluate the latency coupled with accelerator frequency scaling driven by the control CPU (GPU switching latency). The approach employs a minimal, iterative workload that allows statistically distinguishing runtime differences between frequency pairs. It first measures execution times for each frequency and then determines the latency of switching from an initial to a target frequency by tracking runtime changes and repeating measurements to ensure statistical robustness. Finally, the methodology filters out outliers from external factors like driver management or system interruptions. The methodology is implemented in the tool LATEST with support for CUDA accelerators. Evaluated on three Nvidia GPUs - GH200, A100-SXM4, and RTX Quadro 6000 - the analysis reveals significant differences in the switching latency, evaluates optimal frequency change rates, and identifies frequency pairs to avoid due to high overhead. © 2025 Elsevier B.V., All rights reserved.

  • Czech name

  • Czech description

Classification

  • Type

    D - Article in proceedings

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Article name in the collection

    2025 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025 : proceedings : 3-7 June 2025 Milan, Italy

  • ISBN

    979-8-3315-2643-6

  • ISSN

    2639-3867

  • e-ISSN

    2995-066X

  • Number of pages

    10

  • Pages from-to

    830-839

  • Publisher name

    IEEE

  • Place of publication

    Piscataway

  • Event location

    Milán

  • Event date

    Jun 3, 2025

  • Type of event by nationality

    WRD - Celosvětová akce

  • UT code for WoS article