{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/24372"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/24372","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Efficient memory-level parallelism extraction with decoupled strands","abstract":"We present Outrider, an architecture for throughput-oriented processors that exploits intra-thread memory-level parallelism (MLP) to improve performance efficiency on highly threaded workloads. Outrider enables a single thread of execution to be presented to the architecture as multiple decoupled instruction streams, consisting of either memory accessing or memory consuming instructions. The key insight is that by decoupling the instruction streams, the processor pipeline can expose MLP in a way similar to out-of-order designs while relying on a low-complexity in-order micro-architecture. Instead of adding more threads as is done in modern GPUs, Outrider can expose the same MLP with fewer threads and reduced contention for resources shared among threads. We demonstrate that Outrider can outperform single-threaded cores by 23-131% and a 4-way simultaneous multi-threaded core by up to 87% in data parallel applications in a 1024-core system. Outrider achieves these performance gains without incurring the overhead of additional hardware thread contexts, which results in improved efficiency compared to a multi-threaded core.","abstract_html":"We present Outrider, an architecture for throughput-oriented processors that exploits intra-thread memory-level parallelism (MLP) to improve performance efficiency on highly threaded workloads. Outrider enables a single thread of execution to be presented to the architecture as multiple decoupled instruction streams, consisting of either memory accessing or memory consuming instructions. The key insight is that by decoupling the instruction streams, the processor pipeline can expose MLP in a way similar to out-of-order designs while relying on a low-complexity in-order micro-architecture. Instead of adding more threads as is done in modern GPUs, Outrider can expose the same MLP with fewer threads and reduced contention for resources shared among threads. We demonstrate that Outrider can outperform single-threaded cores by 23-131% and a 4-way simultaneous multi-threaded core by up to 87% in data parallel applications in a 1024-core system. Outrider achieves these performance gains without incurring the overhead of additional hardware thread contexts, which results in improved efficiency compared to a multi-threaded core.","abstract_has_math":false,"creators":["Crago, Neal"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Patel, Sanjay J."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-05-25T14:39:25Z","date_published":"2011-05-25T14:39:25Z","updated_at":"2026-07-22T22:25:23Z","subjects":["Memory Latency Tolerance","Accelerators","Processors","Decoupled"],"languages":["en"],"rights":["Copyright 2011 Neal C. Crago"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/24372","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Patel, Sanjay J."]},{"key":"dc:creator","label":"Author","values":["Crago, Neal"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2011-05-25T14:39:25Z","2013-05-26T10:00:23Z","2011-05"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Memory Latency Tolerance","Accelerators","Processors","Decoupled"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2011 Neal C. Crago"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/24372"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We present Outrider, an architecture for throughput-oriented processors that exploits intra-thread memory-level parallelism (MLP) to improve performance efficiency on highly threaded workloads. Outrider enables a single thread of execution to be presented to the architecture as multiple decoupled instruction streams, consisting of either memory accessing or memory consuming instructions. The key insight is that by decoupling the instruction streams, the processor pipeline can expose MLP in a way similar to out-of-order designs while relying on a low-complexity in-order micro-architecture. Instead of adding more threads as is done in modern GPUs, Outrider can expose the same MLP with fewer threads and reduced contention for resources shared among threads. We demonstrate that Outrider can outperform single-threaded cores by 23-131% and a 4-way simultaneous multi-threaded core by up to 87% in data parallel applications in a 1024-core system. Outrider achieves these performance gains without incurring the overhead of additional hardware thread contexts, which results in improved efficiency compared to a multi-threaded core.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-02-28T18:29:41Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Crago_Neal.pdf: 604130 bytes, checksum: 126971f17aa43fb45cb1a2183fd4135e (MD5)","Made available in DSpace on 2011-05-25T14:39:25Z (GMT). No. of bitstreams: 2 Crago_Neal.pdf: 604132 bytes, checksum: bb5add8430d6506cee10ad12d8e0ada1 (MD5) license.txt: 4057 bytes, checksum: da69b4ea0637c38148299bb3f705c009 (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by William Ingram (wingram2@illinois.edu) on 2011-05-25T14:41:40Z Item is restricted until 2013-05-25T14:41:28Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2013-05-26T10:00:23Z Item was in collections: University of Illinois Dissertations and Theses (ID: 204) Dissertations and Theses - Electrical and Computer Engineering (ID: 446) No. of bitstreams: 3 Crago_Neal.pdf.txt: 101227 bytes, checksum: 5c6bbb5af7eb7a54490b35961e925c42 (MD5) Crago_Neal.pdf: 604132 bytes, checksum: bb5add8430d6506cee10ad12d8e0ada1 (MD5) license.txt: 4057 bytes, checksum: da69b4ea0637c38148299bb3f705c009 (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2013-05-26T10:00:23Z"]},{"key":"dc:title","label":"Title","values":["Efficient memory-level parallelism extraction with decoupled strands"]}]}],"canonical_facts":{"dc:contributor":["Patel, Sanjay J."],"dc:creator":["Crago, Neal"],"dc:date":["2011-05-25T14:39:25Z","2013-05-26T10:00:23Z","2011-05"],"dc:description":["We present Outrider, an architecture for throughput-oriented processors that exploits intra-thread memory-level parallelism (MLP) to improve performance efficiency on highly threaded workloads. Outrider enables a single thread of execution to be presented to the architecture as multiple decoupled instruction streams, consisting of either memory accessing or memory consuming instructions. The key insight is that by decoupling the instruction streams, the processor pipeline can expose MLP in a way similar to out-of-order designs while relying on a low-complexity in-order micro-architecture. Instead of adding more threads as is done in modern GPUs, Outrider can expose the same MLP with fewer threads and reduced contention for resources shared among threads. We demonstrate that Outrider can outperform single-threaded cores by 23-131% and a 4-way simultaneous multi-threaded core by up to 87% in data parallel applications in a 1024-core system. Outrider achieves these performance gains without incurring the overhead of additional hardware thread contexts, which results in improved efficiency compared to a multi-threaded core.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-02-28T18:29:41Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Crago_Neal.pdf: 604130 bytes, checksum: 126971f17aa43fb45cb1a2183fd4135e (MD5)","Made available in DSpace on 2011-05-25T14:39:25Z (GMT). No. of bitstreams: 2 Crago_Neal.pdf: 604132 bytes, checksum: bb5add8430d6506cee10ad12d8e0ada1 (MD5) license.txt: 4057 bytes, checksum: da69b4ea0637c38148299bb3f705c009 (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by William Ingram (wingram2@illinois.edu) on 2011-05-25T14:41:40Z Item is restricted until 2013-05-25T14:41:28Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2013-05-26T10:00:23Z Item was in collections: University of Illinois Dissertations and Theses (ID: 204) Dissertations and Theses - Electrical and Computer Engineering (ID: 446) No. of bitstreams: 3 Crago_Neal.pdf.txt: 101227 bytes, checksum: 5c6bbb5af7eb7a54490b35961e925c42 (MD5) Crago_Neal.pdf: 604132 bytes, checksum: bb5add8430d6506cee10ad12d8e0ada1 (MD5) license.txt: 4057 bytes, checksum: da69b4ea0637c38148299bb3f705c009 (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2013-05-26T10:00:23Z"],"dc:identifier":["http://hdl.handle.net/2142/24372"],"dc:language":["en"],"dc:rights":["Copyright 2011 Neal C. Crago"],"dc:subject":["Memory Latency Tolerance","Accelerators","Processors","Decoupled"],"dc:title":["Efficient memory-level parallelism extraction with decoupled strands"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:23Z"}