{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/44133"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/44133","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Eliminating on-chip traffic waste: are we there yet?","abstract":"As technology continues to scale, the memory hierarchy in processors is predicted to be a major component of the overall system energy budget. This has led many researchers into focusing on techniques that minimize the amount of data moved, and the distance that it is moved. While many techniques have been shown to be successful at reducing the amount of on-chip network traffic, no studies have shown how close a combined approach would come to eliminating all unnecessary data traffic, nor have any studies provided insight into where the remaining challenges are. To answer these questions, this thesis systematically analyzes the on-chip traffic for six applications. For this study, we use an alternative hardware-software coherence protocol, DeNovo, and apply several optimizations to DeNovo to quantitatively show the traffic inefficiencies of a directory-based MESI protocol. With a fully optimized DeNovo protocol, we can remove most of the traffic inefficiencies caused by poor spatial locality, fetch-on-write write policy, poor L2 reuse, and MESI protocol overheads. In our discussion of these improvements, we highlight the software causes of these traffic overheads in order to generalize the results. The final DeNovo protocol with all optimizations applied is able to reduce network traffic by an average of 39.5% and execution time by an average of 10.5%, relative to a baseline MESI implementation. On average, 8.8% of the remaining traffic is spent fetching non-useful data. Using only our optimizations, reducing non-useful data movement further would not be possible without losing performance because of irregular access patterns in the applications.","abstract_html":"As technology continues to scale, the memory hierarchy in processors is predicted to be a major component of the overall system energy budget. This has led many researchers into focusing on techniques that minimize the amount of data moved, and the distance that it is moved. While many techniques have been shown to be successful at reducing the amount of on-chip network traffic, no studies have shown how close a combined approach would come to eliminating all unnecessary data traffic, nor have any studies provided insight into where the remaining challenges are. To answer these questions, this thesis systematically analyzes the on-chip traffic for six applications. For this study, we use an alternative hardware-software coherence protocol, DeNovo, and apply several optimizations to DeNovo to quantitatively show the traffic inefficiencies of a directory-based MESI protocol. With a fully optimized DeNovo protocol, we can remove most of the traffic inefficiencies caused by poor spatial locality, fetch-on-write write policy, poor L2 reuse, and MESI protocol overheads. In our discussion of these improvements, we highlight the software causes of these traffic overheads in order to generalize the results. The final DeNovo protocol with all optimizations applied is able to reduce network traffic by an average of 39.5% and execution time by an average of 10.5%, relative to a baseline MESI implementation. On average, 8.8% of the remaining traffic is spent fetching non-useful data. Using only our optimizations, reducing non-useful data movement further would not be possible without losing performance because of irregular access patterns in the applications.","abstract_has_math":false,"creators":["Smolinski, Robert"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Adve, Sarita V."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013-05-24T21:51:45Z","date_published":"2013-05-24T21:51:45Z","updated_at":"2026-07-22T22:25:33Z","subjects":["Multicore cache architecture","Coherence protocols","Network traffic","Energy efficiency"],"languages":["en"],"rights":["Copyright 2013 Robert Smolinski"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/44133","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Adve, Sarita V."]},{"key":"dc:creator","label":"Author","values":["Smolinski, Robert"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2013-05-24T21:51:45Z","2013-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Multicore cache architecture","Coherence protocols","Network traffic","Energy efficiency"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2013 Robert Smolinski"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/44133"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["As technology continues to scale, the memory hierarchy in processors is predicted to be a major component of the overall system energy budget. This has led many researchers into focusing on techniques that minimize the amount of data moved, and the distance that it is moved. While many techniques have been shown to be successful at reducing the amount of on-chip network traffic, no studies have shown how close a combined approach would come to eliminating all unnecessary data traffic, nor have any studies provided insight into where the remaining challenges are. To answer these questions, this thesis systematically analyzes the on-chip traffic for six applications. For this study, we use an alternative hardware-software coherence protocol, DeNovo, and apply several optimizations to DeNovo to quantitatively show the traffic inefficiencies of a directory-based MESI protocol. With a fully optimized DeNovo protocol, we can remove most of the traffic inefficiencies caused by poor spatial locality, fetch-on-write write policy, poor L2 reuse, and MESI protocol overheads. In our discussion of these improvements, we highlight the software causes of these traffic overheads in order to generalize the results. The final DeNovo protocol with all optimizations applied is able to reduce network traffic by an average of 39.5% and execution time by an average of 10.5%, relative to a baseline MESI implementation. On average, 8.8% of the remaining traffic is spent fetching non-useful data. Using only our optimizations, reducing non-useful data movement further would not be possible without losing performance because of irregular access patterns in the applications.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-04-22T19:37:10Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Smolinski_Robert.pdf: 698628 bytes, checksum: 5759d8d744021bc02149e5ce58f2c9a3 (MD5)","Made available in DSpace on 2013-05-24T21:51:45Z (GMT). No. of bitstreams: 2 Robert_Smolinski.pdf: 698628 bytes, checksum: 5759d8d744021bc02149e5ce58f2c9a3 (MD5) license.txt: 4066 bytes, checksum: d66209f59c6fec158b0ae54b90081502 (MD5)"]},{"key":"dc:title","label":"Title","values":["Eliminating on-chip traffic waste: are we there yet?"]}]}],"canonical_facts":{"dc:contributor":["Adve, Sarita V."],"dc:creator":["Smolinski, Robert"],"dc:date":["2013-05-24T21:51:45Z","2013-05"],"dc:description":["As technology continues to scale, the memory hierarchy in processors is predicted to be a major component of the overall system energy budget. This has led many researchers into focusing on techniques that minimize the amount of data moved, and the distance that it is moved. While many techniques have been shown to be successful at reducing the amount of on-chip network traffic, no studies have shown how close a combined approach would come to eliminating all unnecessary data traffic, nor have any studies provided insight into where the remaining challenges are. To answer these questions, this thesis systematically analyzes the on-chip traffic for six applications. For this study, we use an alternative hardware-software coherence protocol, DeNovo, and apply several optimizations to DeNovo to quantitatively show the traffic inefficiencies of a directory-based MESI protocol. With a fully optimized DeNovo protocol, we can remove most of the traffic inefficiencies caused by poor spatial locality, fetch-on-write write policy, poor L2 reuse, and MESI protocol overheads. In our discussion of these improvements, we highlight the software causes of these traffic overheads in order to generalize the results. The final DeNovo protocol with all optimizations applied is able to reduce network traffic by an average of 39.5% and execution time by an average of 10.5%, relative to a baseline MESI implementation. On average, 8.8% of the remaining traffic is spent fetching non-useful data. Using only our optimizations, reducing non-useful data movement further would not be possible without losing performance because of irregular access patterns in the applications.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-04-22T19:37:10Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Smolinski_Robert.pdf: 698628 bytes, checksum: 5759d8d744021bc02149e5ce58f2c9a3 (MD5)","Made available in DSpace on 2013-05-24T21:51:45Z (GMT). No. of bitstreams: 2 Robert_Smolinski.pdf: 698628 bytes, checksum: 5759d8d744021bc02149e5ce58f2c9a3 (MD5) license.txt: 4066 bytes, checksum: d66209f59c6fec158b0ae54b90081502 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/44133"],"dc:language":["en"],"dc:rights":["Copyright 2013 Robert Smolinski"],"dc:subject":["Multicore cache architecture","Coherence protocols","Network traffic","Energy efficiency"],"dc:title":["Eliminating on-chip traffic waste: are we there yet?"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:33Z"}