{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/71991"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/71991","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Cache Design and Performance in a Large-Scale Shared-Memory Multiprocessor System","abstract":"The use of a private cache in each processor of large-scale shared-memory multiprocessor systems can reduce long global memory latency but also introduces the cache coherence problem. Cache design and performance in a large-scale multiprocessor are affected by the cache coherence problem and the cache coherence scheme implemented. The behavior of a parallel program usually differs from that of the same program executed sequentially. Consequently, the cache behaves differently and may not perform as well as the cache in a uniprocessor system. Some results of previous cache studies for a uniprocessor system are less applicable to multiprocessor caches. In this thesis, the cache design and performance using a directory and a software coherence scheme in multistage-interconnection-network-based multiprocessor systems are studied using trace-driven timing simulation of numerical benchmarks. Design complexity and performance trade-offs for both schemes are studied. Their performance problems are analyzed in detail, and several improvements are proposed and evaluated and are shown to be effective in improving the performance. Next, the performance of the directory and software schemes are compared; the simple software scheme is shown to have better performance for numerical programs. The performance advantages and disadvantages of the two schemes are analyzed, and a new coherence scheme combining the best of both schemes is proposed. This new scheme is shown to achieve higher hit ratios. Overall, the global memory remains one of the major performance bottlenecks for a multiprocessor system even though private caches are being used. The effectiveness of memory caches to reduce global memory access latency is demonstrated.","abstract_html":"The use of a private cache in each processor of large-scale shared-memory multiprocessor systems can reduce long global memory latency but also introduces the cache coherence problem. Cache design and performance in a large-scale multiprocessor are affected by the cache coherence problem and the cache coherence scheme implemented. The behavior of a parallel program usually differs from that of the same program executed sequentially. Consequently, the cache behaves differently and may not perform as well as the cache in a uniprocessor system. Some results of previous cache studies for a uniprocessor system are less applicable to multiprocessor caches. In this thesis, the cache design and performance using a directory and a software coherence scheme in multistage-interconnection-network-based multiprocessor systems are studied using trace-driven timing simulation of numerical benchmarks. Design complexity and performance trade-offs for both schemes are studied. Their performance problems are analyzed in detail, and several improvements are proposed and evaluated and are shown to be effective in improving the performance. Next, the performance of the directory and software schemes are compared; the simple software scheme is shown to have better performance for numerical programs. The performance advantages and disadvantages of the two schemes are analyzed, and a new coherence scheme combining the best of both schemes is proposed. This new scheme is shown to achieve higher hit ratios. Overall, the global memory remains one of the major performance bottlenecks for a multiprocessor system even though private caches are being used. The effectiveness of memory caches to reduce global memory access latency is demonstrated.","abstract_has_math":false,"creators":["Chen, Yung-Chin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical Engineering","degree_department":null,"school":null,"contributors":["Veidenbaum, Alexander V."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-12-16T22:23:06Z","date_published":"2014-12-16T22:23:06Z","updated_at":"2026-07-22T22:26:05Z","subjects":["Engineering, Electronics and Electrical"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["(UMI)AAI9314852"],"render_values":[{"text":"(UMI)AAI9314852","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/71991","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Veidenbaum, Alexander V."]},{"key":"dc:creator","label":"Author","values":["Chen, Yung-Chin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014-12-16T22:23:06Z","10000-01-01","1993"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Engineering, Electronics and Electrical"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/71991","(UMI)AAI9314852"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The use of a private cache in each processor of large-scale shared-memory multiprocessor systems can reduce long global memory latency but also introduces the cache coherence problem. Cache design and performance in a large-scale multiprocessor are affected by the cache coherence problem and the cache coherence scheme implemented. The behavior of a parallel program usually differs from that of the same program executed sequentially. Consequently, the cache behaves differently and may not perform as well as the cache in a uniprocessor system. Some results of previous cache studies for a uniprocessor system are less applicable to multiprocessor caches. In this thesis, the cache design and performance using a directory and a software coherence scheme in multistage-interconnection-network-based multiprocessor systems are studied using trace-driven timing simulation of numerical benchmarks. Design complexity and performance trade-offs for both schemes are studied. Their performance problems are analyzed in detail, and several improvements are proposed and evaluated and are shown to be effective in improving the performance. Next, the performance of the directory and software schemes are compared; the simple software scheme is shown to have better performance for numerical programs. The performance advantages and disadvantages of the two schemes are analyzed, and a new coherence scheme combining the best of both schemes is proposed. This new scheme is shown to achieve higher hit ratios. Overall, the global memory remains one of the major performance bottlenecks for a multiprocessor system even though private caches are being used. The effectiveness of memory caches to reduce global memory access latency is demonstrated.","Made available in DSpace on 2014-12-16T22:23:06Z (GMT). No. of bitstreams: 1 9314852.pdf: 7502996 bytes, checksum: 07001ed60666bccf6516e6dd4026f59c (MD5) Previous issue date: 1993","Embargo set by: Seth Robbins for item 72157 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","200 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 1993."]},{"key":"dc:title","label":"Title","values":["Cache Design and Performance in a Large-Scale Shared-Memory Multiprocessor System"]}]}],"canonical_facts":{"dc:contributor":["Veidenbaum, Alexander V."],"dc:creator":["Chen, Yung-Chin"],"dc:date":["2014-12-16T22:23:06Z","10000-01-01","1993"],"dc:description":["The use of a private cache in each processor of large-scale shared-memory multiprocessor systems can reduce long global memory latency but also introduces the cache coherence problem. Cache design and performance in a large-scale multiprocessor are affected by the cache coherence problem and the cache coherence scheme implemented. The behavior of a parallel program usually differs from that of the same program executed sequentially. Consequently, the cache behaves differently and may not perform as well as the cache in a uniprocessor system. Some results of previous cache studies for a uniprocessor system are less applicable to multiprocessor caches. In this thesis, the cache design and performance using a directory and a software coherence scheme in multistage-interconnection-network-based multiprocessor systems are studied using trace-driven timing simulation of numerical benchmarks. Design complexity and performance trade-offs for both schemes are studied. Their performance problems are analyzed in detail, and several improvements are proposed and evaluated and are shown to be effective in improving the performance. Next, the performance of the directory and software schemes are compared; the simple software scheme is shown to have better performance for numerical programs. The performance advantages and disadvantages of the two schemes are analyzed, and a new coherence scheme combining the best of both schemes is proposed. This new scheme is shown to achieve higher hit ratios. Overall, the global memory remains one of the major performance bottlenecks for a multiprocessor system even though private caches are being used. The effectiveness of memory caches to reduce global memory access latency is demonstrated.","Made available in DSpace on 2014-12-16T22:23:06Z (GMT). No. of bitstreams: 1 9314852.pdf: 7502996 bytes, checksum: 07001ed60666bccf6516e6dd4026f59c (MD5) Previous issue date: 1993","Embargo set by: Seth Robbins for item 72157 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","200 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 1993."],"dc:identifier":["http://hdl.handle.net/2142/71991","(UMI)AAI9314852"],"dc:subject":["Engineering, Electronics and Electrical"],"dc:title":["Cache Design and Performance in a Large-Scale Shared-Memory Multiprocessor System"],"dc:type":["text"],"thesis:degree_discipline":["Electrical Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:05Z"}