{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/101588"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/101588","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Machine learning for selecting parallel I/O benchmark applications","abstract":"I/O is one of the main performance bottlenecks for many data-intensive scientific applications. Accurate I/O performance benchmarking, which can help us better understand the causes of these bottlenecks and to guide the performance optimization of poor performing applications, is therefore an important problem. We investigate the use of submodular function maximization as a way to select a set of I/O benchmark applications using measures of similarities between applications computed from I/O statistics obtained from the Darshan logs of their jobs. Our optimization problem simultaneously seeks a set of applications that are representative of the applications running on the HPC platform they are chosen from while simultaneously encouraging them to possess diverse I/O behavior between them. We evaluate the quality of the selected applications by training classifiers using features extracted from the jobs of these applications to predict the I/O performance of other jobs that were ran on the platform. Our experiments indicate that the trained classifiers can achieve a fair level of accuracy, thereby lending credence to the feasibility of our optimization approach for selecting I/O benchmark applications.","abstract_html":"I/O is one of the main performance bottlenecks for many data-intensive scientific applications. Accurate I/O performance benchmarking, which can help us better understand the causes of these bottlenecks and to guide the performance optimization of poor performing applications, is therefore an important problem. We investigate the use of submodular function maximization as a way to select a set of I/O benchmark applications using measures of similarities between applications computed from I/O statistics obtained from the Darshan logs of their jobs. Our optimization problem simultaneously seeks a set of applications that are representative of the applications running on the HPC platform they are chosen from while simultaneously encouraging them to possess diverse I/O behavior between them. We evaluate the quality of the selected applications by training classifiers using features extracted from the jobs of these applications to predict the I/O performance of other jobs that were ran on the platform. Our experiments indicate that the trained classifiers can achieve a fair level of accuracy, thereby lending credence to the feasibility of our optimization approach for selecting I/O benchmark applications.","abstract_has_math":false,"creators":["Ng, Hong Wei"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Winslett, Marianne S"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-09-27T16:17:53Z","date_published":"2018-09-27T16:17:53Z","updated_at":"2026-07-22T22:24:40Z","subjects":["Machine Learning","Submodular Function Optimization","Parallel I/O"],"languages":["en"],"rights":["Copyright 2018 Hong Wei Ng"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/101588","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Winslett, Marianne S"]},{"key":"dc:creator","label":"Author","values":["Ng, Hong Wei"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-09-27T16:17:53Z","2018-07-16","2018-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","Submodular Function Optimization","Parallel I/O"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Hong Wei Ng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/101588"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["I/O is one of the main performance bottlenecks for many data-intensive scientific applications. Accurate I/O performance benchmarking, which can help us better understand the causes of these bottlenecks and to guide the performance optimization of poor performing applications, is therefore an important problem. We investigate the use of submodular function maximization as a way to select a set of I/O benchmark applications using measures of similarities between applications computed from I/O statistics obtained from the Darshan logs of their jobs. Our optimization problem simultaneously seeks a set of applications that are representative of the applications running on the HPC platform they are chosen from while simultaneously encouraging them to possess diverse I/O behavior between them. We evaluate the quality of the selected applications by training classifiers using features extracted from the jobs of these applications to predict the I/O performance of other jobs that were ran on the platform. Our experiments indicate that the trained classifiers can achieve a fair level of accuracy, thereby lending credence to the feasibility of our optimization approach for selecting I/O benchmark applications.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2018-09-27 without embargo terms","The student, Hong Wei Ng, accepted the attached license on 2018-07-15 at 13:30.","The student, Hong Wei Ng, submitted this Thesis for approval on 2018-07-15 at 13:53.","This Thesis was approved for publication on 2018-07-16 at 10:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12879 on 2018-09-27 at 10:48:39","Made available in DSpace on 2018-09-27T16:17:53Z (GMT). No. of bitstreams: 3 NG-THESIS-2018.pdf: 463885 bytes, checksum: 51e79a0f823b85093d2b61ffab3b0cb3 (MD5) submitted-thesis.zip: 2143285 bytes, checksum: 397f286c16e43f4f6c4aab5ebdc894c8 (MD5) LICENSE.txt: 4208 bytes, checksum: d10358093a4bdaa3085481d150b44896 (MD5) Previous issue date: 2018-07-16"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Machine learning for selecting parallel I/O benchmark applications"]}]}],"canonical_facts":{"dc:contributor":["Winslett, Marianne S"],"dc:creator":["Ng, Hong Wei"],"dc:date":["2018-09-27T16:17:53Z","2018-07-16","2018-08"],"dc:description":["I/O is one of the main performance bottlenecks for many data-intensive scientific applications. Accurate I/O performance benchmarking, which can help us better understand the causes of these bottlenecks and to guide the performance optimization of poor performing applications, is therefore an important problem. We investigate the use of submodular function maximization as a way to select a set of I/O benchmark applications using measures of similarities between applications computed from I/O statistics obtained from the Darshan logs of their jobs. Our optimization problem simultaneously seeks a set of applications that are representative of the applications running on the HPC platform they are chosen from while simultaneously encouraging them to possess diverse I/O behavior between them. We evaluate the quality of the selected applications by training classifiers using features extracted from the jobs of these applications to predict the I/O performance of other jobs that were ran on the platform. Our experiments indicate that the trained classifiers can achieve a fair level of accuracy, thereby lending credence to the feasibility of our optimization approach for selecting I/O benchmark applications.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2018-09-27 without embargo terms","The student, Hong Wei Ng, accepted the attached license on 2018-07-15 at 13:30.","The student, Hong Wei Ng, submitted this Thesis for approval on 2018-07-15 at 13:53.","This Thesis was approved for publication on 2018-07-16 at 10:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12879 on 2018-09-27 at 10:48:39","Made available in DSpace on 2018-09-27T16:17:53Z (GMT). No. of bitstreams: 3 NG-THESIS-2018.pdf: 463885 bytes, checksum: 51e79a0f823b85093d2b61ffab3b0cb3 (MD5) submitted-thesis.zip: 2143285 bytes, checksum: 397f286c16e43f4f6c4aab5ebdc894c8 (MD5) LICENSE.txt: 4208 bytes, checksum: d10358093a4bdaa3085481d150b44896 (MD5) Previous issue date: 2018-07-16"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/101588"],"dc:language":["en"],"dc:rights":["Copyright 2018 Hong Wei Ng"],"dc:subject":["Machine Learning","Submodular Function Optimization","Parallel I/O"],"dc:title":["Machine learning for selecting parallel I/O benchmark applications"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:40Z"}