{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/34589"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/34589","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Energy-efficient latency tolerance for 1000-core data parallel processors with decoupled strands","abstract":"This dissertation presents a novel decoupled latency tolerance technique for 1000-core data parallel processors. The approach focuses on developing instruction latency tolerance to improve performance for a single thread. The main idea behind the approach is to leverage the compiler to split the original thread into separate memory-accessing and memory-consuming instruction streams. The goal is to provide latency tolerance similar to high-performance techniques such as out-of-order execution while leveraging low hardware complexity similar to an in-order execution core. The research in this dissertation supports the following thesis: Pipeline stalls due to long exposed instruction latency are the main performance limiter for cached 1000-core data parallel processors. Leveraging natural decoupling of memory-access and memory-consumption, a serial thread of execution can be partitioned into strands providing energy-efficient latency tolerance. This dissertation motivates the need for latency tolerance in 1000-core data parallel processors and presents decoupled core architectures as an alternative to currently used techniques. This dissertation discusses the limitations of prior decoupled architectures, and proposes techniques to improve both latency tolerance and energy-efficiency. Finally, the success of the proposed decoupled architecture is demonstrated against other approaches by performing an exhaustive design space exploration of energy, area, and performance using high-fidelity performance and physical design models.","abstract_html":"This dissertation presents a novel decoupled latency tolerance technique for 1000-core data parallel processors. The approach focuses on developing instruction latency tolerance to improve performance for a single thread. The main idea behind the approach is to leverage the compiler to split the original thread into separate memory-accessing and memory-consuming instruction streams. The goal is to provide latency tolerance similar to high-performance techniques such as out-of-order execution while leveraging low hardware complexity similar to an in-order execution core. The research in this dissertation supports the following thesis: Pipeline stalls due to long exposed instruction latency are the main performance limiter for cached 1000-core data parallel processors. Leveraging natural decoupling of memory-access and memory-consumption, a serial thread of execution can be partitioned into strands providing energy-efficient latency tolerance. This dissertation motivates the need for latency tolerance in 1000-core data parallel processors and presents decoupled core architectures as an alternative to currently used techniques. This dissertation discusses the limitations of prior decoupled architectures, and proposes techniques to improve both latency tolerance and energy-efficiency. Finally, the success of the proposed decoupled architecture is demonstrated against other approaches by performing an exhaustive design space exploration of energy, area, and performance using high-fidelity performance and physical design models.","abstract_has_math":false,"creators":["Crago, Neal"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Patel, Sanjay J.","Hwu, Wen-Mei W.","Lumetta, Steven S.","Chen, Deming"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-09-18T21:27:03Z","date_published":"2012-09-18T21:27:03Z","updated_at":"2026-07-22T22:25:31Z","subjects":["Parallel Processing","Data-parallel","Graphics processing unit (GPU)","General-purpose computing on graphics processing units (GPGPU)","manycore","latency tolerance","decoupled architecture","compiler technique","energy-efficiency","power-efficiency","high-performance","low power","low energy"],"languages":["en"],"rights":["Copyright 2012 Neal Crago"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/34589","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Patel, Sanjay J.","Hwu, Wen-Mei W.","Lumetta, Steven S.","Chen, Deming"]},{"key":"dc:creator","label":"Author","values":["Crago, Neal"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-09-18T21:27:03Z","2014-09-18T10:00:37Z","2012-08"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Parallel Processing","Data-parallel","Graphics processing unit (GPU)","General-purpose computing on graphics processing units (GPGPU)","manycore","latency tolerance","decoupled architecture","compiler technique","energy-efficiency","power-efficiency","high-performance","low power","low energy"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2012 Neal Crago"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/34589"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This dissertation presents a novel decoupled latency tolerance technique for 1000-core data parallel processors. The approach focuses on developing instruction latency tolerance to improve performance for a single thread. The main idea behind the approach is to leverage the compiler to split the original thread into separate memory-accessing and memory-consuming instruction streams. The goal is to provide latency tolerance similar to high-performance techniques such as out-of-order execution while leveraging low hardware complexity similar to an in-order execution core. The research in this dissertation supports the following thesis: Pipeline stalls due to long exposed instruction latency are the main performance limiter for cached 1000-core data parallel processors. Leveraging natural decoupling of memory-access and memory-consumption, a serial thread of execution can be partitioned into strands providing energy-efficient latency tolerance. This dissertation motivates the need for latency tolerance in 1000-core data parallel processors and presents decoupled core architectures as an alternative to currently used techniques. This dissertation discusses the limitations of prior decoupled architectures, and proposes techniques to improve both latency tolerance and energy-efficiency. Finally, the success of the proposed decoupled architecture is demonstrated against other approaches by performing an exhaustive design space exploration of energy, area, and performance using high-fidelity performance and physical design models.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-06-19T19:03:41Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Crago_Neal.pdf: 3338678 bytes, checksum: 635196bb43d9f3d84f06799bed3f906b (MD5)","Made available in DSpace on 2012-09-18T21:27:03Z (GMT). No. of bitstreams: 2 Crago_Neal.pdf: 3338678 bytes, checksum: 635196bb43d9f3d84f06799bed3f906b (MD5) license.txt: 4057 bytes, checksum: 86fbe008458fbf43d3e59fb54092c9c3 (MD5)","Item marked as restricted to the 'Administrator' Group (id=1) by Seth Robbins (srobbins@illinois.edu) on 2012-09-18T21:27:38Z Item is restricted until 2014-09-18T21:27:16Z","Restriction data tranferred 2014-07-01T11:35:56-05:00 Original Data Group with Access Administrator Release Date: 2014-09-18 16:27:16 UTC Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 34863 on 2014-09-18T10:00:37Z."]},{"key":"dc:title","label":"Title","values":["Energy-efficient latency tolerance for 1000-core data parallel processors with decoupled strands"]}]}],"canonical_facts":{"dc:contributor":["Patel, Sanjay J.","Hwu, Wen-Mei W.","Lumetta, Steven S.","Chen, Deming"],"dc:creator":["Crago, Neal"],"dc:date":["2012-09-18T21:27:03Z","2014-09-18T10:00:37Z","2012-08"],"dc:description":["This dissertation presents a novel decoupled latency tolerance technique for 1000-core data parallel processors. The approach focuses on developing instruction latency tolerance to improve performance for a single thread. The main idea behind the approach is to leverage the compiler to split the original thread into separate memory-accessing and memory-consuming instruction streams. The goal is to provide latency tolerance similar to high-performance techniques such as out-of-order execution while leveraging low hardware complexity similar to an in-order execution core. The research in this dissertation supports the following thesis: Pipeline stalls due to long exposed instruction latency are the main performance limiter for cached 1000-core data parallel processors. Leveraging natural decoupling of memory-access and memory-consumption, a serial thread of execution can be partitioned into strands providing energy-efficient latency tolerance. This dissertation motivates the need for latency tolerance in 1000-core data parallel processors and presents decoupled core architectures as an alternative to currently used techniques. This dissertation discusses the limitations of prior decoupled architectures, and proposes techniques to improve both latency tolerance and energy-efficiency. Finally, the success of the proposed decoupled architecture is demonstrated against other approaches by performing an exhaustive design space exploration of energy, area, and performance using high-fidelity performance and physical design models.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-06-19T19:03:41Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Crago_Neal.pdf: 3338678 bytes, checksum: 635196bb43d9f3d84f06799bed3f906b (MD5)","Made available in DSpace on 2012-09-18T21:27:03Z (GMT). No. of bitstreams: 2 Crago_Neal.pdf: 3338678 bytes, checksum: 635196bb43d9f3d84f06799bed3f906b (MD5) license.txt: 4057 bytes, checksum: 86fbe008458fbf43d3e59fb54092c9c3 (MD5)","Item marked as restricted to the 'Administrator' Group (id=1) by Seth Robbins (srobbins@illinois.edu) on 2012-09-18T21:27:38Z Item is restricted until 2014-09-18T21:27:16Z","Restriction data tranferred 2014-07-01T11:35:56-05:00 Original Data Group with Access Administrator Release Date: 2014-09-18 16:27:16 UTC Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 34863 on 2014-09-18T10:00:37Z."],"dc:identifier":["http://hdl.handle.net/2142/34589"],"dc:language":["en"],"dc:rights":["Copyright 2012 Neal Crago"],"dc:subject":["Parallel Processing","Data-parallel","Graphics processing unit (GPU)","General-purpose computing on graphics processing units (GPGPU)","manycore","latency tolerance","decoupled architecture","compiler technique","energy-efficiency","power-efficiency","high-performance","low power","low energy"],"dc:title":["Energy-efficient latency tolerance for 1000-core data parallel processors with decoupled strands"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:31Z"}