Abstract
dc:description.abstractThe fast broader adoption of ML applications has caused a surge in their global energy usage, necessitating a comprehensive understanding of the tradeoffs between execution speed and energy consumption. While previous work was focused on time-only or inference-only studies, we provide a more complete picture by covering a wider space of parameters. We contribute: (1) time and energy profiling across inference and training of DNN operations, (2) operations-level and full DNN performance predictions trained on our profiling results and (3) graphical evaluation and validation of profiling and predictor results. Profiling results of Nvidia A30 performance across several core clocks reveal an energy optimum of 900 MHz, aligning with the manufacturer base clock of 930 MHz. This work provides the tools which enable the correct choice of target GPU and clock speed for existing and future models.
Degree
thesis:*- Level thesis:degree_level
- master
- Grantor dc:publisher
- Universität Heidelberg
- Year
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Nicolai, Constantin
- Contributors dc:contributor
-
- Fröning, Holger
Identifiers
dc:identifier.*- Repository record source_url
- http://www.ub.uni-heidelberg.de/archiv/37234
- OAI identifier oai:identifier
- oai:archiv.ub.uni-heidelberg.de:37234