{"id":{"repo_id":"gatech","oai_identifier":"oai:repository.gatech.edu:1853/81660"},"canonical_url":"https://search.dev.ndltd.org/etd/gatech/oai:repository.gatech.edu:1853/81660","repository":{"repo_id":"gatech","name":"Georgia Tech","base_url":"https://repository.gatech.edu/server/oai/request"},"display":{"title":"Algorithm–Hardware Co-Design Of Digital Compute-In-Memory Architecture Supporting Flexible And Temporal N:M Sparsity","abstract":"Foundational models continue to grow in size and capability, increasing demands on memory, energy, and computation. This has motivated algorithm–hardware co-design approaches for efficient inference, with structured N:M sparsity emerging as a promising method to reduce overhead. However, fixed sparsity configurations across layers and decoding steps can limit model expressivity and degrade accuracy, while supporting multiple configurations introduces challenges in both pattern selection and hardware efficiency. This thesis addresses these challenges through a co-design approach that enables flexible and dynamic N:M sparsity. At the algorithm level, it introduces FLOW, a layer-wise framework that selects sparsity configurations based on the magnitude and distribution of outliers, and FLOW++, which extends this approach to the temporal domain by adapting sparsity during different decoding steps. At the hardware level, it presents FlexCiM, a low-overhead digital compute-in-memory architecture designed to efficiently support models with varying sparsity patterns. Together, these contributions demonstrate that flexible N:M sparsity, when co-designed with hardware, can effectively balance model expressivity and computational efficiency.","abstract_html":"Foundational models continue to grow in size and capability, increasing demands on memory, energy, and computation. This has motivated algorithm–hardware co-design approaches for efficient inference, with structured N:M sparsity emerging as a promising method to reduce overhead. However, fixed sparsity configurations across layers and decoding steps can limit model expressivity and degrade accuracy, while supporting multiple configurations introduces challenges in both pattern selection and hardware efficiency. This thesis addresses these challenges through a co-design approach that enables flexible and dynamic N:M sparsity. At the algorithm level, it introduces FLOW, a layer-wise framework that selects sparsity configurations based on the magnitude and distribution of outliers, and FLOW++, which extends this approach to the temporal domain by adapting sparsity during different decoding steps. At the hardware level, it presents FlexCiM, a low-overhead digital compute-in-memory architecture designed to efficiently support models with varying sparsity patterns. Together, these contributions demonstrate that flexible N:M sparsity, when co-designed with hardware, can effectively balance model expressivity and computational efficiency.","abstract_has_math":false,"creators":["Ramachandran, Akshat"],"institution":"Georgia Institute of Technology","degree_name":"Electrical and Computer Engineering, PhD","degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Lin, Yingyan (Celine)"],"committee_chairs":[],"committee_members":["Krishna, Tushar","Hao, Callie"],"year":2026,"date_issued":"2026-05","date_published":"2026-05","updated_at":"2026-07-27T19:49:34Z","subjects":[],"languages":["English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1853/81660","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Lin, Yingyan (Celine)"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Krishna, Tushar","Hao, Callie"]},{"key":"dc:creator","label":"Author","values":["Ramachandran, Akshat"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-05-27T15:11:37Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-05-27T15:11:37Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-05"]},{"key":"dc:type","label":"Dc Type","values":["Text"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Electrical and Computer Engineering, PhD"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Georgia Institute of Technology"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["English"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1853/81660"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Foundational models continue to grow in size and capability, increasing demands on memory, energy, and computation. This has motivated algorithm–hardware co-design approaches for efficient inference, with structured N:M sparsity emerging as a promising method to reduce overhead. However, fixed sparsity configurations across layers and decoding steps can limit model expressivity and degrade accuracy, while supporting multiple configurations introduces challenges in both pattern selection and hardware efficiency. This thesis addresses these challenges through a co-design approach that enables flexible and dynamic N:M sparsity. At the algorithm level, it introduces FLOW, a layer-wise framework that selects sparsity configurations based on the magnitude and distribution of outliers, and FLOW++, which extends this approach to the temporal domain by adapting sparsity during different decoding steps. At the hardware level, it presents FlexCiM, a low-overhead digital compute-in-memory architecture designed to efficiently support models with varying sparsity patterns. Together, these contributions demonstrate that flexible N:M sparsity, when co-designed with hardware, can effectively balance model expressivity and computational efficiency."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Algorithm–Hardware Co-Design Of Digital Compute-In-Memory Architecture Supporting Flexible And Temporal N:M Sparsity"]}]}],"canonical_facts":{"dc:contributor.advisor":["Lin, Yingyan (Celine)"],"dc:contributor.committeemember":["Krishna, Tushar","Hao, Callie"],"dc:creator":["Ramachandran, Akshat"],"dc:date.accessioned":["2026-05-27T15:11:37Z"],"dc:date.available":["2026-05-27T15:11:37Z"],"dc:date.issued":["2026-05"],"dc:description.abstract":["Foundational models continue to grow in size and capability, increasing demands on memory, energy, and computation. This has motivated algorithm–hardware co-design approaches for efficient inference, with structured N:M sparsity emerging as a promising method to reduce overhead. However, fixed sparsity configurations across layers and decoding steps can limit model expressivity and degrade accuracy, while supporting multiple configurations introduces challenges in both pattern selection and hardware efficiency. This thesis addresses these challenges through a co-design approach that enables flexible and dynamic N:M sparsity. At the algorithm level, it introduces FLOW, a layer-wise framework that selects sparsity configurations based on the magnitude and distribution of outliers, and FLOW++, which extends this approach to the temporal domain by adapting sparsity during different decoding steps. At the hardware level, it presents FlexCiM, a low-overhead digital compute-in-memory architecture designed to efficiently support models with varying sparsity patterns. Together, these contributions demonstrate that flexible N:M sparsity, when co-designed with hardware, can effectively balance model expressivity and computational efficiency."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/1853/81660"],"dc:language.iso":["English"],"dc:title":["Algorithm–Hardware Co-Design Of Digital Compute-In-Memory Architecture Supporting Flexible And Temporal N:M Sparsity"],"dc:type":["Text"],"thesis:degree_name":["Electrical and Computer Engineering, PhD"],"thesis:institution_name":["Georgia Institute of Technology"]},"updated_at":"2026-07-27T19:49:34Z"}