{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/106208"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/106208","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Scalable non-blocking Krylov solvers for extreme-scale computing","abstract":"Krylov solvers are key kernels in many large-scale science and engineering applications for solving sparse linear systems. Extreme-scale systems have many factors that increase communication costs and cause performance variation across cores that can reduce performance at scale. Many Krylov solvers require frequent blocking allreduce collective operations that can limit performance at scale due to the increasing cost of this collective as the node count increases and the cost of synchronizing all processes. This thesis investigates non-blocking Krylov solver variations designed to reduce communication costs by overlapping communication and computation using non-blocking allreduces. These variations can allow us to hide most of the allreduce cost and avoiding synchronizing all processes to produce better performance at scale. This work builds on gaps in the literature to help us gain a more thorough understanding of the performance and robustness of these solvers and how we can use them to efficiently solve linear systems at scale in practice. A variety of blocking and non-blocking Krylov solvers are analyzed in detail with multiple different preconditioners on multiple leadership-class supercomputers. Performance analysis tools and performance models are developed to provide deeper insight into the performance barriers encountered by these algorithms and show how they relate to observed performance. These tools guide us to a variety of optimizations to further improve solver performance. The Nek5000 and Quda applications are used to analyze the effectiveness of these solvers in practice. Both applications are designed to perform well at scale, however they need further improvements to reach their desired performance. The resulting tools and analysis provide us with a better understanding of how to improve performance at scale that can benefit a wider range of applications.","abstract_html":"Krylov solvers are key kernels in many large-scale science and engineering applications for solving sparse linear systems. Extreme-scale systems have many factors that increase communication costs and cause performance variation across cores that can reduce performance at scale. Many Krylov solvers require frequent blocking allreduce collective operations that can limit performance at scale due to the increasing cost of this collective as the node count increases and the cost of synchronizing all processes. This thesis investigates non-blocking Krylov solver variations designed to reduce communication costs by overlapping communication and computation using non-blocking allreduces. These variations can allow us to hide most of the allreduce cost and avoiding synchronizing all processes to produce better performance at scale. This work builds on gaps in the literature to help us gain a more thorough understanding of the performance and robustness of these solvers and how we can use them to efficiently solve linear systems at scale in practice. A variety of blocking and non-blocking Krylov solvers are analyzed in detail with multiple different preconditioners on multiple leadership-class supercomputers. Performance analysis tools and performance models are developed to provide deeper insight into the performance barriers encountered by these algorithms and show how they relate to observed performance. These tools guide us to a variety of optimizations to further improve solver performance. The Nek5000 and Quda applications are used to analyze the effectiveness of these solvers in practice. Both applications are designed to perform well at scale, however they need further improvements to reach their desired performance. The resulting tools and analysis provide us with a better understanding of how to improve performance at scale that can benefit a wider range of applications.","abstract_has_math":false,"creators":["Eller, Paul R."],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Gropp, William D.","Olson, Luke N.","Fischer, Paul F.","Hoemmen, Mark"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-03-02T21:58:15Z","date_published":"2020-03-02T21:58:15Z","updated_at":"2026-07-22T22:24:45Z","subjects":["Scientific Computing","High Performance Computing","Linear Solvers","Krylov Solvers","Scalable Solvers","Performance Modeling","Performance Analysis","Non-blocking Communication"],"languages":["en"],"rights":["Copyright 2019 Paul R. Eller"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/106208","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Gropp, William D.","Olson, Luke N.","Fischer, Paul F.","Hoemmen, Mark"]},{"key":"dc:creator","label":"Author","values":["Eller, Paul R."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-03-02T21:58:15Z","2019-11-25","2019-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Scientific Computing","High Performance Computing","Linear Solvers","Krylov Solvers","Scalable Solvers","Performance Modeling","Performance Analysis","Non-blocking Communication"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Paul R. Eller"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/106208"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Krylov solvers are key kernels in many large-scale science and engineering applications for solving sparse linear systems. Extreme-scale systems have many factors that increase communication costs and cause performance variation across cores that can reduce performance at scale. Many Krylov solvers require frequent blocking allreduce collective operations that can limit performance at scale due to the increasing cost of this collective as the node count increases and the cost of synchronizing all processes. This thesis investigates non-blocking Krylov solver variations designed to reduce communication costs by overlapping communication and computation using non-blocking allreduces. These variations can allow us to hide most of the allreduce cost and avoiding synchronizing all processes to produce better performance at scale. This work builds on gaps in the literature to help us gain a more thorough understanding of the performance and robustness of these solvers and how we can use them to efficiently solve linear systems at scale in practice. A variety of blocking and non-blocking Krylov solvers are analyzed in detail with multiple different preconditioners on multiple leadership-class supercomputers. Performance analysis tools and performance models are developed to provide deeper insight into the performance barriers encountered by these algorithms and show how they relate to observed performance. These tools guide us to a variety of optimizations to further improve solver performance. The Nek5000 and Quda applications are used to analyze the effectiveness of these solvers in practice. Both applications are designed to perform well at scale, however they need further improvements to reach their desired performance. The resulting tools and analysis provide us with a better understanding of how to improve performance at scale that can benefit a wider range of applications.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-02-28 without embargo terms","The student, Paul Eller, accepted the attached license on 2019-11-22 at 18:55.","The student, Paul Eller, submitted this Dissertation for approval on 2019-11-22 at 19:10.","This Dissertation was approved for publication on 2019-11-25 at 14:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14598 on 2020-02-28 at 17:14:06","Made available in DSpace on 2020-03-02T21:58:15Z (GMT). No. of bitstreams: 2 ELLER-DISSERTATION-2019.pdf: 17559363 bytes, checksum: 7608e046e58f3670a08d7b15c4ba5319 (MD5) LICENSE.txt: 4207 bytes, checksum: 51e7a5e0c8e70cae4709d754d46517ce (MD5) Previous issue date: 2019-11-25"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Scalable non-blocking Krylov solvers for extreme-scale computing"]}]}],"canonical_facts":{"dc:contributor":["Gropp, William D.","Olson, Luke N.","Fischer, Paul F.","Hoemmen, Mark"],"dc:creator":["Eller, Paul R."],"dc:date":["2020-03-02T21:58:15Z","2019-11-25","2019-12"],"dc:description":["Krylov solvers are key kernels in many large-scale science and engineering applications for solving sparse linear systems. Extreme-scale systems have many factors that increase communication costs and cause performance variation across cores that can reduce performance at scale. Many Krylov solvers require frequent blocking allreduce collective operations that can limit performance at scale due to the increasing cost of this collective as the node count increases and the cost of synchronizing all processes. This thesis investigates non-blocking Krylov solver variations designed to reduce communication costs by overlapping communication and computation using non-blocking allreduces. These variations can allow us to hide most of the allreduce cost and avoiding synchronizing all processes to produce better performance at scale. This work builds on gaps in the literature to help us gain a more thorough understanding of the performance and robustness of these solvers and how we can use them to efficiently solve linear systems at scale in practice. A variety of blocking and non-blocking Krylov solvers are analyzed in detail with multiple different preconditioners on multiple leadership-class supercomputers. Performance analysis tools and performance models are developed to provide deeper insight into the performance barriers encountered by these algorithms and show how they relate to observed performance. These tools guide us to a variety of optimizations to further improve solver performance. The Nek5000 and Quda applications are used to analyze the effectiveness of these solvers in practice. Both applications are designed to perform well at scale, however they need further improvements to reach their desired performance. The resulting tools and analysis provide us with a better understanding of how to improve performance at scale that can benefit a wider range of applications.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-02-28 without embargo terms","The student, Paul Eller, accepted the attached license on 2019-11-22 at 18:55.","The student, Paul Eller, submitted this Dissertation for approval on 2019-11-22 at 19:10.","This Dissertation was approved for publication on 2019-11-25 at 14:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14598 on 2020-02-28 at 17:14:06","Made available in DSpace on 2020-03-02T21:58:15Z (GMT). No. of bitstreams: 2 ELLER-DISSERTATION-2019.pdf: 17559363 bytes, checksum: 7608e046e58f3670a08d7b15c4ba5319 (MD5) LICENSE.txt: 4207 bytes, checksum: 51e7a5e0c8e70cae4709d754d46517ce (MD5) Previous issue date: 2019-11-25"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/106208"],"dc:language":["en"],"dc:rights":["Copyright 2019 Paul R. Eller"],"dc:subject":["Scientific Computing","High Performance Computing","Linear Solvers","Krylov Solvers","Scalable Solvers","Performance Modeling","Performance Analysis","Non-blocking Communication"],"dc:title":["Scalable non-blocking Krylov solvers for extreme-scale computing"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:45Z"}