{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/98331"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/98331","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Toward performance portability for CPUS and GPUS through algorithmic compositions","abstract":"This Dissertation was approved for publication on 2017-07-05 at 16:09.","abstract_html":"This Dissertation was approved for publication on 2017-07-05 at 16:09.","abstract_has_math":false,"creators":["Chang, Li-Wen"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hwu, Wen-Mei W.","Chen, Deming","Kim, Nam Sung","Lumetta, Steven S."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-09-29T17:56:23Z","date_published":"2017-09-29T17:56:23Z","updated_at":"2026-07-22T22:24:35Z","subjects":["Performance portability","Algorithmic composition","Parallel programming","TANGRAM","Programming language","Compiler","Graphics processing units (GPUs)","Central processing units (CPUs)","Open Computing Language (OpenCL)","Open Multi-Processing (OpenMP)","Open Accelerators (OpenACC)","C++ Accelerated Massive Parallelism (C++AMP)"],"languages":["en"],"rights":["Copyright 2017 Li-Wen Chang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/98331","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hwu, Wen-Mei W.","Chen, Deming","Kim, Nam Sung","Lumetta, Steven S."]},{"key":"dc:creator","label":"Author","values":["Chang, Li-Wen"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-09-29T17:56:23Z","2017-07-05","2017-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Performance portability","Algorithmic composition","Parallel programming","TANGRAM","Programming language","Compiler","Graphics processing units (GPUs)","Central processing units (CPUs)","Open Computing Language (OpenCL)","Open Multi-Processing (OpenMP)","Open Accelerators (OpenACC)","C++ Accelerated Massive Parallelism (C++AMP)"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Li-Wen Chang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/98331"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This Dissertation was approved for publication on 2017-07-05 at 16:09.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11312 on 2017-09-29 at 11:27:52","Made available in DSpace on 2017-09-29T17:56:23Z (GMT). No. of bitstreams: 4 CHANG-DISSERTATION-2017.pdf: 1657939 bytes, checksum: 7048c555db6c106ffefb431bfb4e5b47 (MD5) LICENSE.txt: 4209 bytes, checksum: 7a40d214e9a508b94657f974dfe31dd1 (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: 42ba911c1d5422ad15633f1bf4daefa6 (MD5) reprint_micro16.pdf: 134473 bytes, checksum: 67e346917d59814ec994f697ae909d19 (MD5) Previous issue date: 2017-07-05","The diversity of microarchitecture designs in heterogeneous computing systems allows programs to achieve high performance and energy efficiency, but results in substantial software redevelopment cost for each type or generation of hardware. To mitigate this cost, a performance portable programming system is required. This work presents my solution to the performance portability problem. I argue that a new language is required for replacing the current practices of programming systems to achieve practical performance portability. To support my argument, I first demonstrate the limited performance portability of the current practices by showing quantitative and qualitative evidences. I identify the main limiting issues of conventional programming languages. To overcome the issues, I propose a new modular, composition-based programming language that can effectively express an algorithmic design space with functional polymorphism, and a compiler that can effectively explore the design space and facilitate many high-level optimization techniques. This proposed approach achieves no less than 70% of the performance of highly optimized vendor libraries such as Intel MKL and NVIDIA CUBLAS/CUSPARSE on an Intel i7-3820 Sandy Bridge CPU, an NVIDIA C2050 Fermi GPU, and an NVIDIA K20c Kepler GPU.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-09-29 without embargo terms","The student, Li-Wen Chang, accepted the attached license on 2017-07-05 at 13:48.","The student, Li-Wen Chang, submitted this Dissertation for approval on 2017-07-05 at 14:09."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Toward performance portability for CPUS and GPUS through algorithmic compositions"]}]}],"canonical_facts":{"dc:contributor":["Hwu, Wen-Mei W.","Chen, Deming","Kim, Nam Sung","Lumetta, Steven S."],"dc:creator":["Chang, Li-Wen"],"dc:date":["2017-09-29T17:56:23Z","2017-07-05","2017-08"],"dc:description":["This Dissertation was approved for publication on 2017-07-05 at 16:09.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11312 on 2017-09-29 at 11:27:52","Made available in DSpace on 2017-09-29T17:56:23Z (GMT). No. of bitstreams: 4 CHANG-DISSERTATION-2017.pdf: 1657939 bytes, checksum: 7048c555db6c106ffefb431bfb4e5b47 (MD5) LICENSE.txt: 4209 bytes, checksum: 7a40d214e9a508b94657f974dfe31dd1 (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: 42ba911c1d5422ad15633f1bf4daefa6 (MD5) reprint_micro16.pdf: 134473 bytes, checksum: 67e346917d59814ec994f697ae909d19 (MD5) Previous issue date: 2017-07-05","The diversity of microarchitecture designs in heterogeneous computing systems allows programs to achieve high performance and energy efficiency, but results in substantial software redevelopment cost for each type or generation of hardware. To mitigate this cost, a performance portable programming system is required. This work presents my solution to the performance portability problem. I argue that a new language is required for replacing the current practices of programming systems to achieve practical performance portability. To support my argument, I first demonstrate the limited performance portability of the current practices by showing quantitative and qualitative evidences. I identify the main limiting issues of conventional programming languages. To overcome the issues, I propose a new modular, composition-based programming language that can effectively express an algorithmic design space with functional polymorphism, and a compiler that can effectively explore the design space and facilitate many high-level optimization techniques. This proposed approach achieves no less than 70% of the performance of highly optimized vendor libraries such as Intel MKL and NVIDIA CUBLAS/CUSPARSE on an Intel i7-3820 Sandy Bridge CPU, an NVIDIA C2050 Fermi GPU, and an NVIDIA K20c Kepler GPU.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-09-29 without embargo terms","The student, Li-Wen Chang, accepted the attached license on 2017-07-05 at 13:48.","The student, Li-Wen Chang, submitted this Dissertation for approval on 2017-07-05 at 14:09."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/98331"],"dc:language":["en"],"dc:rights":["Copyright 2017 Li-Wen Chang"],"dc:subject":["Performance portability","Algorithmic composition","Parallel programming","TANGRAM","Programming language","Compiler","Graphics processing units (GPUs)","Central processing units (CPUs)","Open Computing Language (OpenCL)","Open Multi-Processing (OpenMP)","Open Accelerators (OpenACC)","C++ Accelerated Massive Parallelism (C++AMP)"],"dc:title":["Toward performance portability for CPUS and GPUS through algorithmic compositions"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:35Z"}