{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113011"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113011","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Program transformation and code generation for developing, modeling, and optimizing GPU programs","abstract":"The ability to efficiently optimize or re-optimize an algorithm for high performance on a particular processor architecture is crucial in a wide spectrum of engineering and scientific applications. To help bridge the gap between fully automated optimization, which does not reliably produce optimal code, and fully manual optimization, which is inefficient and error-prone, we present three mechanisms facilitating a human-guided, partially automated process for production, optimization, and analysis of high-performance code: (1) formal mechanisms for describing and reasoning about program dependencies and loop nest structure within a programming system providing transformation-based code generation for GPUs and CPUs; (2) a system for constructing customizable, black-box performance models for GPUs; and (3) a visual user interface for program optimization and analysis. We demonstrate how our human-accessible dependency and loop-nest semantics enable reasoning about program correctness in a transformation and optimization procedure for two example scientific computing computations. We show how models created using our hardware-agnostic performance modeling approach can predict execution times accurately enough to select the best-performing of multiple implementation variants of a given computation, demonstrating results for three computations on five different GPUs. We show how our program transformation interface enables access to and navigation of a comprehensive search space of implementation variants for a given computation, as well as comparisons of multiple optimization paths. In doing so, we demonstrate how these architecture-agnostic tools, some of which are already being utilized by a persistent user base, further the goal of facilitating production, optimization, and analysis of high-performance code.","abstract_html":"The ability to efficiently optimize or re-optimize an algorithm for high performance on a particular processor architecture is crucial in a wide spectrum of engineering and scientific applications. To help bridge the gap between fully automated optimization, which does not reliably produce optimal code, and fully manual optimization, which is inefficient and error-prone, we present three mechanisms facilitating a human-guided, partially automated process for production, optimization, and analysis of high-performance code: (1) formal mechanisms for describing and reasoning about program dependencies and loop nest structure within a programming system providing transformation-based code generation for GPUs and CPUs; (2) a system for constructing customizable, black-box performance models for GPUs; and (3) a visual user interface for program optimization and analysis. We demonstrate how our human-accessible dependency and loop-nest semantics enable reasoning about program correctness in a transformation and optimization procedure for two example scientific computing computations. We show how models created using our hardware-agnostic performance modeling approach can predict execution times accurately enough to select the best-performing of multiple implementation variants of a given computation, demonstrating results for three computations on five different GPUs. We show how our program transformation interface enables access to and navigation of a comprehensive search space of implementation variants for a given computation, as well as comparisons of multiple optimization paths. In doing so, we demonstrate how these architecture-agnostic tools, some of which are already being utilized by a persistent user base, further the goal of facilitating production, optimization, and analysis of high-performance code.","abstract_has_math":false,"creators":["Stevens, James D"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Klöckner, Andreas","Gropp, William","Solomonik, Edgar","Owens, John"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T21:45:33Z","date_published":"2022-01-12T21:45:33Z","updated_at":"2026-07-22T22:24:52Z","subjects":["Code generation","OpenCL","Program transformation","Polyhedral model","Dependencies","Performance model","GPU","Microbenchmark","Black box","Performance optimization","Visual programming interface","Human-guided transformation"],"languages":["en"],"rights":["Copyright 2021 James D. Stevens"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113011","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Klöckner, Andreas","Gropp, William","Solomonik, Edgar","Owens, John"]},{"key":"dc:creator","label":"Author","values":["Stevens, James D"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T21:45:33Z","2021-07-13","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Code generation","OpenCL","Program transformation","Polyhedral model","Dependencies","Performance model","GPU","Microbenchmark","Black box","Performance optimization","Visual programming interface","Human-guided transformation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 James D. Stevens"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113011"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The ability to efficiently optimize or re-optimize an algorithm for high performance on a particular processor architecture is crucial in a wide spectrum of engineering and scientific applications. To help bridge the gap between fully automated optimization, which does not reliably produce optimal code, and fully manual optimization, which is inefficient and error-prone, we present three mechanisms facilitating a human-guided, partially automated process for production, optimization, and analysis of high-performance code: (1) formal mechanisms for describing and reasoning about program dependencies and loop nest structure within a programming system providing transformation-based code generation for GPUs and CPUs; (2) a system for constructing customizable, black-box performance models for GPUs; and (3) a visual user interface for program optimization and analysis. We demonstrate how our human-accessible dependency and loop-nest semantics enable reasoning about program correctness in a transformation and optimization procedure for two example scientific computing computations. We show how models created using our hardware-agnostic performance modeling approach can predict execution times accurately enough to select the best-performing of multiple implementation variants of a given computation, demonstrating results for three computations on five different GPUs. We show how our program transformation interface enables access to and navigation of a comprehensive search space of implementation variants for a given computation, as well as comparisons of multiple optimization paths. In doing so, we demonstrate how these architecture-agnostic tools, some of which are already being utilized by a persistent user base, further the goal of facilitating production, optimization, and analysis of high-performance code.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, James Stevens, accepted the attached license on 2021-07-09 at 17:53.","The student, James Stevens, submitted this Dissertation for approval on 2021-07-09 at 18:04.","This Dissertation was approved for publication on 2021-07-13 at 16:54.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16841 on 2022-01-12 at 12:44:47","Made available in DSpace on 2022-01-12T21:45:33Z (GMT). No. of bitstreams: 3 STEVENS-DISSERTATION-2021.pdf: 3352357 bytes, checksum: 721e2a1475481a44decff7f8fc35dacd (MD5) LICENSE.txt: 4210 bytes, checksum: ccfc8c390b3af61ac9e890a8d3d7b59f (MD5) PROQUEST_LICENSE.txt: 4556 bytes, checksum: 8cfbce349698e523f9a7e7dca7dfae1d (MD5) Previous issue date: 2021-07-13"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Program transformation and code generation for developing, modeling, and optimizing GPU programs"]}]}],"canonical_facts":{"dc:contributor":["Klöckner, Andreas","Gropp, William","Solomonik, Edgar","Owens, John"],"dc:creator":["Stevens, James D"],"dc:date":["2022-01-12T21:45:33Z","2021-07-13","2021-08"],"dc:description":["The ability to efficiently optimize or re-optimize an algorithm for high performance on a particular processor architecture is crucial in a wide spectrum of engineering and scientific applications. To help bridge the gap between fully automated optimization, which does not reliably produce optimal code, and fully manual optimization, which is inefficient and error-prone, we present three mechanisms facilitating a human-guided, partially automated process for production, optimization, and analysis of high-performance code: (1) formal mechanisms for describing and reasoning about program dependencies and loop nest structure within a programming system providing transformation-based code generation for GPUs and CPUs; (2) a system for constructing customizable, black-box performance models for GPUs; and (3) a visual user interface for program optimization and analysis. We demonstrate how our human-accessible dependency and loop-nest semantics enable reasoning about program correctness in a transformation and optimization procedure for two example scientific computing computations. We show how models created using our hardware-agnostic performance modeling approach can predict execution times accurately enough to select the best-performing of multiple implementation variants of a given computation, demonstrating results for three computations on five different GPUs. We show how our program transformation interface enables access to and navigation of a comprehensive search space of implementation variants for a given computation, as well as comparisons of multiple optimization paths. In doing so, we demonstrate how these architecture-agnostic tools, some of which are already being utilized by a persistent user base, further the goal of facilitating production, optimization, and analysis of high-performance code.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, James Stevens, accepted the attached license on 2021-07-09 at 17:53.","The student, James Stevens, submitted this Dissertation for approval on 2021-07-09 at 18:04.","This Dissertation was approved for publication on 2021-07-13 at 16:54.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16841 on 2022-01-12 at 12:44:47","Made available in DSpace on 2022-01-12T21:45:33Z (GMT). No. of bitstreams: 3 STEVENS-DISSERTATION-2021.pdf: 3352357 bytes, checksum: 721e2a1475481a44decff7f8fc35dacd (MD5) LICENSE.txt: 4210 bytes, checksum: ccfc8c390b3af61ac9e890a8d3d7b59f (MD5) PROQUEST_LICENSE.txt: 4556 bytes, checksum: 8cfbce349698e523f9a7e7dca7dfae1d (MD5) Previous issue date: 2021-07-13"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113011"],"dc:language":["en"],"dc:rights":["Copyright 2021 James D. Stevens"],"dc:subject":["Code generation","OpenCL","Program transformation","Polyhedral model","Dependencies","Performance model","GPU","Microbenchmark","Black box","Performance optimization","Visual programming interface","Human-guided transformation"],"dc:title":["Program transformation and code generation for developing, modeling, and optimizing GPU programs"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}