{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/22386"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/22386","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Processor parallelism considerations and memory latency reduction in shared memory multiprocessors","abstract":"A wide variety of computer architectures have been proposed to exploit parallelism at different granularities. These architectures have significant differences in instruction scheduling constraints, memory latencies, and synchronization overhead, making it difficult to determine which architecture can achieve the best performance on a given program. Trace-driven simulations and analytic models are used to compare the instruction-level parallelism of a superscalar processor and a pipelined processor with the loop-level parallelism of a shared memory multiprocessor. It is shown that the maximum speedup for a loop with a cyclic dependence graph is limited by its critical dependence ratio, independent of the number of iterations in the loop. The fine-grained processors are better suited for executing these loops with cyclic dependence graphs, while the multiprocessor has better performance on the very parallel loops with acyclic dependence graphs. When executing programs with a variety of loops and sequential code, the best performance is obtained using a multiprocessor architecture in which each individual processor has a fine-grained parallelism of two to four.","abstract_html":"A wide variety of computer architectures have been proposed to exploit parallelism at different granularities. These architectures have significant differences in instruction scheduling constraints, memory latencies, and synchronization overhead, making it difficult to determine which architecture can achieve the best performance on a given program. Trace-driven simulations and analytic models are used to compare the instruction-level parallelism of a superscalar processor and a pipelined processor with the loop-level parallelism of a shared memory multiprocessor. It is shown that the maximum speedup for a loop with a cyclic dependence graph is limited by its critical dependence ratio, independent of the number of iterations in the loop. The fine-grained processors are better suited for executing these loops with cyclic dependence graphs, while the multiprocessor has better performance on the very parallel loops with acyclic dependence graphs. When executing programs with a variety of loops and sequential code, the best performance is obtained using a multiprocessor architecture in which each individual processor has a fine-grained parallelism of two to four.","abstract_has_math":false,"creators":["Lilja, David John"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical and Computer Engineering","degree_department":null,"school":null,"contributors":["Yew, Pen-Chung"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-05-07T13:38:12Z","date_published":"2011-05-07T13:38:12Z","updated_at":"2026-07-22T22:25:19Z","subjects":["Engineering, Electronics and Electrical","Computer Science"],"languages":["eng"],"rights":["Copyright 1991 Lilja, David John"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["AAI9210894","(UMI)AAI9210894"],"render_values":[{"text":"AAI9210894","href":null,"code":true},{"text":"(UMI)AAI9210894","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/22386","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Yew, Pen-Chung"]},{"key":"dc:creator","label":"Author","values":["Lilja, David John"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2011-05-07T13:38:12Z","10000-01-01","1991"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical and Computer Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Engineering, Electronics and Electrical","Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 1991 Lilja, David John"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/22386","AAI9210894","(UMI)AAI9210894"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["A wide variety of computer architectures have been proposed to exploit parallelism at different granularities. These architectures have significant differences in instruction scheduling constraints, memory latencies, and synchronization overhead, making it difficult to determine which architecture can achieve the best performance on a given program. Trace-driven simulations and analytic models are used to compare the instruction-level parallelism of a superscalar processor and a pipelined processor with the loop-level parallelism of a shared memory multiprocessor. It is shown that the maximum speedup for a loop with a cyclic dependence graph is limited by its critical dependence ratio, independent of the number of iterations in the loop. The fine-grained processors are better suited for executing these loops with cyclic dependence graphs, while the multiprocessor has better performance on the very parallel loops with acyclic dependence graphs. When executing programs with a variety of loops and sequential code, the best performance is obtained using a multiprocessor architecture in which each individual processor has a fine-grained parallelism of two to four.","A major problem with this type of shared memory multiprocessor architecture is the long latency in fetching operands from the shared memory. Private data caches are an effective means of reducing this latency, but they introduce the complexity of a cache coherence mechanism. Both hardware and software schemes have been proposed for maintaining coherence in these systems. Unfortunately, hardware schemes have very high memory requirements, and software schemes rely on imprecise compile-time memory disambiguation. A new compiler-assisted directory coherence mechanism is proposed that combines the best aspects of the hardware and software approaches while eliminating many of their disadvantages. The pointer cache directory significantly reduces the size of a hardware directory by dynamically binding pointers to cache blocks only when the blocks are actually referenced. Compiler optimizations can further reduce the size of the directory by signaling the hardware to allocate pointers only when they are needed. Detailed trace-driven simulations show that the performance of this new approach is comparable to other coherence schemes, but with significantly lower memory requirements.","Made available in DSpace on 2011-05-07T13:38:12Z (GMT). No. of bitstreams: 2 license.txt: 4922 bytes, checksum: 910b249b4beec47e7ab768910c8f966f (MD5) 9210894.pdf: 6249112 bytes, checksum: 5b133d530944eee8da5cd90fbba4582c (MD5) Previous issue date: 1991","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by Howard Ding (hding2@illinois.edu) on 2011-05-07T14:57:16Z Item is restricted indefinitely.","Restriction data tranferred 2014-07-01T11:26:51-05:00 Original Data Group with Access UIUC Users [automated] Release Date: none Reason: ETDs are only available to UIUC Users without author permission","ETDs are only available to UIUC Users without author permission","U of I Only"]},{"key":"dc:title","label":"Title","values":["Processor parallelism considerations and memory latency reduction in shared memory multiprocessors"]}]}],"canonical_facts":{"dc:contributor":["Yew, Pen-Chung"],"dc:creator":["Lilja, David John"],"dc:date":["2011-05-07T13:38:12Z","10000-01-01","1991"],"dc:description":["A wide variety of computer architectures have been proposed to exploit parallelism at different granularities. These architectures have significant differences in instruction scheduling constraints, memory latencies, and synchronization overhead, making it difficult to determine which architecture can achieve the best performance on a given program. Trace-driven simulations and analytic models are used to compare the instruction-level parallelism of a superscalar processor and a pipelined processor with the loop-level parallelism of a shared memory multiprocessor. It is shown that the maximum speedup for a loop with a cyclic dependence graph is limited by its critical dependence ratio, independent of the number of iterations in the loop. The fine-grained processors are better suited for executing these loops with cyclic dependence graphs, while the multiprocessor has better performance on the very parallel loops with acyclic dependence graphs. When executing programs with a variety of loops and sequential code, the best performance is obtained using a multiprocessor architecture in which each individual processor has a fine-grained parallelism of two to four.","A major problem with this type of shared memory multiprocessor architecture is the long latency in fetching operands from the shared memory. Private data caches are an effective means of reducing this latency, but they introduce the complexity of a cache coherence mechanism. Both hardware and software schemes have been proposed for maintaining coherence in these systems. Unfortunately, hardware schemes have very high memory requirements, and software schemes rely on imprecise compile-time memory disambiguation. A new compiler-assisted directory coherence mechanism is proposed that combines the best aspects of the hardware and software approaches while eliminating many of their disadvantages. The pointer cache directory significantly reduces the size of a hardware directory by dynamically binding pointers to cache blocks only when the blocks are actually referenced. Compiler optimizations can further reduce the size of the directory by signaling the hardware to allocate pointers only when they are needed. Detailed trace-driven simulations show that the performance of this new approach is comparable to other coherence schemes, but with significantly lower memory requirements.","Made available in DSpace on 2011-05-07T13:38:12Z (GMT). No. of bitstreams: 2 license.txt: 4922 bytes, checksum: 910b249b4beec47e7ab768910c8f966f (MD5) 9210894.pdf: 6249112 bytes, checksum: 5b133d530944eee8da5cd90fbba4582c (MD5) Previous issue date: 1991","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by Howard Ding (hding2@illinois.edu) on 2011-05-07T14:57:16Z Item is restricted indefinitely.","Restriction data tranferred 2014-07-01T11:26:51-05:00 Original Data Group with Access UIUC Users [automated] Release Date: none Reason: ETDs are only available to UIUC Users without author permission","ETDs are only available to UIUC Users without author permission","U of I Only"],"dc:identifier":["http://hdl.handle.net/2142/22386","AAI9210894","(UMI)AAI9210894"],"dc:language":["eng"],"dc:rights":["Copyright 1991 Lilja, David John"],"dc:subject":["Engineering, Electronics and Electrical","Computer Science"],"dc:title":["Processor parallelism considerations and memory latency reduction in shared memory multiprocessors"],"dc:type":["text"],"thesis:degree_discipline":["Electrical and Computer Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:19Z"}