{"id":{"repo_id":"unr","oai_identifier":"oai:scholarwolf.unr.edu:11714/8129"},"canonical_url":"https://search.dev.ndltd.org/etd/unr/oai:scholarwolf.unr.edu:11714/8129","repository":{"repo_id":"unr","name":"University of Nevada - Reno","base_url":"https://scholarwolf.unr.edu/server/oai/request"},"display":{"title":"I/O Throughput Prediction for HPC Applications Using Darshan Logs","abstract":"As most High Performance Computing (HPC) applications deal with large volumes of data, I/O performance is of critical importance to optimize application performance. Despite having large-scale, high-performance parallel file systems, many applications still suffer from poor I/O performance. Although existing system monitoring tools gather performance statistics, it can be challenging to interpret multidimensional data, thereby distinguishing normal behavior from abnormal ones. Therefore, it is important to derive models that can process I/O statistics gathered by existing monitoring tools. In this thesis, I develop machine learning (ML) models to process file system statistics as reported by Darshan monitoring tool to predict I/O throughput of HPC applications, which then can be compared against the observed I/O throughput to identify performance issues. By processing Darshan logs of BlueWaters supercomputer, I trained several ML models including Decision Tree, Random Forest, Gradient Boosting Tree, and Deep Neural Network (DNN) using different feature scaling methods. I found that the DNN model outperformed other solutions as it can estimate the throughput of I/O operations within the 16 MB/s range. We believe that this work makes an important contribution to the field by deriving accurate models to process Darshan logs to detect file system performance anomalies (e.g., overloaded metadata server, high resource interference, etc.) that can be tackled in a timely manner to minimize interruptions.","abstract_html":"As most High Performance Computing (HPC) applications deal with large volumes of data, I/O performance is of critical importance to optimize application performance. Despite having large-scale, high-performance parallel file systems, many applications still suffer from poor I/O performance. Although existing system monitoring tools gather performance statistics, it can be challenging to interpret multidimensional data, thereby distinguishing normal behavior from abnormal ones. Therefore, it is important to derive models that can process I/O statistics gathered by existing monitoring tools. In this thesis, I develop machine learning (ML) models to process file system statistics as reported by Darshan monitoring tool to predict I/O throughput of HPC applications, which then can be compared against the observed I/O throughput to identify performance issues. By processing Darshan logs of BlueWaters supercomputer, I trained several ML models including Decision Tree, Random Forest, Gradient Boosting Tree, and Deep Neural Network (DNN) using different feature scaling methods. I found that the DNN model outperformed other solutions as it can estimate the throughput of I/O operations within the 16 MB/s range. We believe that this work makes an important contribution to the field by deriving accurate models to process Darshan logs to detect file system performance anomalies (e.g., overloaded metadata server, high resource interference, etc.) that can be tackled in a timely manner to minimize interruptions.","abstract_has_math":false,"creators":["Gabriel, David James"],"institution":null,"degree_name":null,"degree_level":"Master's Degree","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Arslan, Engin"],"committee_chairs":[],"committee_members":["Nguyen, Tin","Cheol Yim, Won","Zeh, David W."],"year":2022,"date_issued":"2022","date_published":"2022","updated_at":"2026-07-27T21:46:52Z","subjects":["High Performance Computing","I/O","Machine Learning"],"languages":[],"rights":["Creative Commons Attribution 4.0 United States"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/11714/8129","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Arslan, Engin"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Nguyen, Tin","Cheol Yim, Won","Zeh, David W."]},{"key":"dc:creator","label":"Author","values":["Gabriel, David James"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2022-06-28T01:06:41Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2022-06-28T01:06:41Z"]},{"key":"dc:date.issued","label":"Date","values":["2022"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Master's Degree"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["High Performance Computing","I/O","Machine Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["Creative Commons Attribution 4.0 United States"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/11714/8129"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["As most High Performance Computing (HPC) applications deal with large volumes of data, I/O performance is of critical importance to optimize application performance. Despite having large-scale, high-performance parallel file systems, many applications still suffer from poor I/O performance. Although existing system monitoring tools gather performance statistics, it can be challenging to interpret multidimensional data, thereby distinguishing normal behavior from abnormal ones. Therefore, it is important to derive models that can process I/O statistics gathered by existing monitoring tools. In this thesis, I develop machine learning (ML) models to process file system statistics as reported by Darshan monitoring tool to predict I/O throughput of HPC applications, which then can be compared against the observed I/O throughput to identify performance issues. By processing Darshan logs of BlueWaters supercomputer, I trained several ML models including Decision Tree, Random Forest, Gradient Boosting Tree, and Deep Neural Network (DNN) using different feature scaling methods. I found that the DNN model outperformed other solutions as it can estimate the throughput of I/O operations within the 16 MB/s range. We believe that this work makes an important contribution to the field by deriving accurate models to process Darshan logs to detect file system performance anomalies (e.g., overloaded metadata server, high resource interference, etc.) that can be tackled in a timely manner to minimize interruptions."]},{"key":"dc:format","label":"Dc Format","values":["PDF"]},{"key":"dc:title","label":"Title","values":["I/O Throughput Prediction for HPC Applications Using Darshan Logs"]}]}],"canonical_facts":{"dc:contributor.advisor":["Arslan, Engin"],"dc:contributor.committeemember":["Nguyen, Tin","Cheol Yim, Won","Zeh, David W."],"dc:creator":["Gabriel, David James"],"dc:date.accessioned":["2022-06-28T01:06:41Z"],"dc:date.available":["2022-06-28T01:06:41Z"],"dc:date.issued":["2022"],"dc:description.abstract":["As most High Performance Computing (HPC) applications deal with large volumes of data, I/O performance is of critical importance to optimize application performance. Despite having large-scale, high-performance parallel file systems, many applications still suffer from poor I/O performance. Although existing system monitoring tools gather performance statistics, it can be challenging to interpret multidimensional data, thereby distinguishing normal behavior from abnormal ones. Therefore, it is important to derive models that can process I/O statistics gathered by existing monitoring tools. In this thesis, I develop machine learning (ML) models to process file system statistics as reported by Darshan monitoring tool to predict I/O throughput of HPC applications, which then can be compared against the observed I/O throughput to identify performance issues. By processing Darshan logs of BlueWaters supercomputer, I trained several ML models including Decision Tree, Random Forest, Gradient Boosting Tree, and Deep Neural Network (DNN) using different feature scaling methods. I found that the DNN model outperformed other solutions as it can estimate the throughput of I/O operations within the 16 MB/s range. We believe that this work makes an important contribution to the field by deriving accurate models to process Darshan logs to detect file system performance anomalies (e.g., overloaded metadata server, high resource interference, etc.) that can be tackled in a timely manner to minimize interruptions."],"dc:format":["PDF"],"dc:identifier.uri":["http://hdl.handle.net/11714/8129"],"dc:rights":["Creative Commons Attribution 4.0 United States"],"dc:subject":["High Performance Computing","I/O","Machine Learning"],"dc:title":["I/O Throughput Prediction for HPC Applications Using Darshan Logs"],"dc:type":["Thesis"],"thesis:degree_level":["Master's Degree"]},"updated_at":"2026-07-27T21:46:52Z"}