{"id":{"repo_id":"texas","oai_identifier":"oai:repositories.lib.utexas.edu:2152/1862"},"canonical_url":"https://search.dev.ndltd.org/etd/texas/oai:repositories.lib.utexas.edu:2152/1862","repository":{"repo_id":"texas","name":"University of Texas","base_url":"https://repositories.lib.utexas.edu/server/oai/request"},"display":{"title":"Scalable primary cache memory architectures","abstract":"For the past decade, microprocessors have been improving in overall performance at a rate of approximately 50–60% per year by exploiting a rapid increase in clock rate and improving instruction throughput. A part of this trend has included the growth of on-chip caches which in modern processor can be as large as 2MB. However, as smaller technologies become prevalent, achieving low average memory access time by simply scaling existing designs becomes more difficult because of process limitations. This research shows that scaling an existing design by either keeping the latency of various structures constant or allowing the latency to vary while keeping the capacity constant leads to degradation in the instructions per cycle (IPC). The goal of this research is to improve IPC at small feature sizes, using a combination of circuit and architectural techniques. This research develops technology-based models to estimate cache access times and uses the models for architectural performance estimation. The performance of a microarchitecture with clustered functional units coupled with a partitioned primary data cache is estimated using the cache access time models. This research evaluates both static and dynamic data mapping on the partitioned primary data cache and shows that using dynamic mapping in combination with the partitioned cache outperforms both the unified cache as well as the statically mapped design. In conjunction with the dynamic data mapping, this research proposes and evaluates predictive instruction steering strategies that help improve the performance of clustered processor designs. This research shows that a hybrid predictive instruction steering policy coupled with an aggressive dynamic mapping of data in a partitioned primary data cache can significantly improve the instructions per cycle (IPC) of a clustered processor, relative to dependence based steering with a unified data cache.","abstract_html":"For the past decade, microprocessors have been improving in overall performance at a rate of approximately 50–60% per year by exploiting a rapid increase in clock rate and improving instruction throughput. A part of this trend has included the growth of on-chip caches which in modern processor can be as large as 2MB. However, as smaller technologies become prevalent, achieving low average memory access time by simply scaling existing designs becomes more difficult because of process limitations. This research shows that scaling an existing design by either keeping the latency of various structures constant or allowing the latency to vary while keeping the capacity constant leads to degradation in the instructions per cycle (IPC). The goal of this research is to improve IPC at small feature sizes, using a combination of circuit and architectural techniques. This research develops technology-based models to estimate cache access times and uses the models for architectural performance estimation. The performance of a microarchitecture with clustered functional units coupled with a partitioned primary data cache is estimated using the cache access time models. This research evaluates both static and dynamic data mapping on the partitioned primary data cache and shows that using dynamic mapping in combination with the partitioned cache outperforms both the unified cache as well as the statically mapped design. In conjunction with the dynamic data mapping, this research proposes and evaluates predictive instruction steering strategies that help improve the performance of clustered processor designs. This research shows that a hybrid predictive instruction steering policy coupled with an aggressive dynamic mapping of data in a partitioned primary data cache can significantly improve the instructions per cycle (IPC) of a clustered processor, relative to dependence based steering with a unified data cache.","abstract_has_math":false,"creators":["Agarwal, Vikas"],"institution":"The University of Texas at Austin","degree_name":"Doctor of Philosophy","degree_level":"Doctoral","degree_discipline":"Electrical and Computer Engineering","degree_department":null,"school":null,"contributors":[],"advisors":["John, Lizy Kurian"],"committee_chairs":[],"committee_members":[],"year":2004,"date_issued":"2004","date_published":"2004","updated_at":"2026-07-24T05:01:04Z","subjects":[],"languages":["eng"],"rights":["Copyright is held by the author. Presentation of this material on the Libraries&apos; web site by University Libraries, The University of Texas at Austin was made possible under a limited license grant from the author who has retained all copyrights in the works."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["b60716824"],"render_values":[{"text":"b60716824","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2152/1862","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["John, Lizy Kurian"]},{"key":"dc:creator","label":"Author","values":["Agarwal, Vikas"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2008-08-28T22:20:55Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2008-08-28T22:20:55Z"]},{"key":"dc:date.issued","label":"Date","values":["2004"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical and Computer Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["The University of Texas at Austin"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright is held by the author. Presentation of this material on the Libraries&apos; web site by University Libraries, The University of Texas at Austin was made possible under a limited license grant from the author who has retained all copyrights in the works."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["b60716824"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/2152/1862"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["text"]},{"key":"dc:description.abstract","label":"Abstract","values":["For the past decade, microprocessors have been improving in overall performance at a rate of approximately 50–60% per year by exploiting a rapid increase in clock rate and improving instruction throughput. A part of this trend has included the growth of on-chip caches which in modern processor can be as large as 2MB. However, as smaller technologies become prevalent, achieving low average memory access time by simply scaling existing designs becomes more difficult because of process limitations. This research shows that scaling an existing design by either keeping the latency of various structures constant or allowing the latency to vary while keeping the capacity constant leads to degradation in the instructions per cycle (IPC). The goal of this research is to improve IPC at small feature sizes, using a combination of circuit and architectural techniques. This research develops technology-based models to estimate cache access times and uses the models for architectural performance estimation. The performance of a microarchitecture with clustered functional units coupled with a partitioned primary data cache is estimated using the cache access time models. This research evaluates both static and dynamic data mapping on the partitioned primary data cache and shows that using dynamic mapping in combination with the partitioned cache outperforms both the unified cache as well as the statically mapped design. In conjunction with the dynamic data mapping, this research proposes and evaluates predictive instruction steering strategies that help improve the performance of clustered processor designs. This research shows that a hybrid predictive instruction steering policy coupled with an aggressive dynamic mapping of data in a partitioned primary data cache can significantly improve the instructions per cycle (IPC) of a clustered processor, relative to dependence based steering with a unified data cache."]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["electronic"]},{"key":"dc:title","label":"Title","values":["Scalable primary cache memory architectures"]}]}],"canonical_facts":{"dc:contributor.advisor":["John, Lizy Kurian"],"dc:creator":["Agarwal, Vikas"],"dc:date.accessioned":["2008-08-28T22:20:55Z"],"dc:date.available":["2008-08-28T22:20:55Z"],"dc:date.issued":["2004"],"dc:description":["text"],"dc:description.abstract":["For the past decade, microprocessors have been improving in overall performance at a rate of approximately 50–60% per year by exploiting a rapid increase in clock rate and improving instruction throughput. A part of this trend has included the growth of on-chip caches which in modern processor can be as large as 2MB. However, as smaller technologies become prevalent, achieving low average memory access time by simply scaling existing designs becomes more difficult because of process limitations. This research shows that scaling an existing design by either keeping the latency of various structures constant or allowing the latency to vary while keeping the capacity constant leads to degradation in the instructions per cycle (IPC). The goal of this research is to improve IPC at small feature sizes, using a combination of circuit and architectural techniques. This research develops technology-based models to estimate cache access times and uses the models for architectural performance estimation. The performance of a microarchitecture with clustered functional units coupled with a partitioned primary data cache is estimated using the cache access time models. This research evaluates both static and dynamic data mapping on the partitioned primary data cache and shows that using dynamic mapping in combination with the partitioned cache outperforms both the unified cache as well as the statically mapped design. In conjunction with the dynamic data mapping, this research proposes and evaluates predictive instruction steering strategies that help improve the performance of clustered processor designs. This research shows that a hybrid predictive instruction steering policy coupled with an aggressive dynamic mapping of data in a partitioned primary data cache can significantly improve the instructions per cycle (IPC) of a clustered processor, relative to dependence based steering with a unified data cache."],"dc:format.medium":["electronic"],"dc:identifier":["b60716824"],"dc:identifier.uri":["http://hdl.handle.net/2152/1862"],"dc:language.iso":["eng"],"dc:rights":["Copyright is held by the author. Presentation of this material on the Libraries&apos; web site by University Libraries, The University of Texas at Austin was made possible under a limited license grant from the author who has retained all copyrights in the works."],"dc:title":["Scalable primary cache memory architectures"],"thesis:degree_discipline":["Electrical and Computer Engineering"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["The University of Texas at Austin"]},"updated_at":"2026-07-24T05:01:04Z"}