{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/130094"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/130094","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Architectures for machine learning and machine learning for architecture","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2027-08-01","abstract_has_math":false,"creators":["Nam, Hyoungwook"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Torrellas, Josep","Li, Bo","Mendis, Charith","Bose, Pradip","Pothukuchi, Raghavendra","Kim, Nam Sung"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-07-10","date_published":"2025-07-10","updated_at":"2026-07-22T22:25:06Z","subjects":["Machine Learning","Computer Architecture"],"languages":["en","eng"],"rights":["Copyright 2025 Hyoungwook Nam"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/130094","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Torrellas, Josep","Li, Bo","Mendis, Charith","Bose, Pradip","Pothukuchi, Raghavendra","Kim, Nam Sung"]},{"key":"dc:creator","label":"Author","values":["Nam, Hyoungwook"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-07-10","2025-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","Computer Architecture"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Hyoungwook Nam"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/130094"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-08-01","The student, Hyoungwook Nam, accepted the attached license on 2025-07-10 at 10:22.","The student, Hyoungwook Nam, submitted this Dissertation for approval on 2025-07-10 at 10:30.","This Dissertation was approved for publication on 2025-07-10 at 12:22.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22484 on 2025-10-25 at 15:30:58","The unprecedented success of machine learning (ML) has brought a problem and an opportunity to computer architecture research. The problem is how to build efficient and scalable computer systems for ML computations. The opportunity is how can ML methods help solving research problems in computer architecture. This dissertation explores both directions of research. In the direction of computer architecture for ML, this work focuses on two challenges for large-scale ML: scalability and power efficiency. This dissertation proposes two proposals for these: MeshSlice and PowerGrad. For the other direction, this dissertation proposes FriendlyFoe, which uses adversarial ML for hardware security. The first proposal is MeshSlice, a framework for efficient 2D tensor parallelism (TP) in distributed DNN training. MeshSlice consists of a novel 2D GeMM algorithm and an autotuner. The MeshSlice GeMM algorithm slices the collective communications into multiple partial collectives that allow overlapping communications with computations. As a result, MeshSlice hides most of the communication latency. The MeshSlice LLM autotuner automates finding the optimal configuration of 2D GeMM dataflow, the mesh shape, and the communication granularity using an analytical cost model. MeshSlice shows significant speedup in LLM training workloads compared to the state-of-the-art 2D TP method. The second proposal is PowerGrad, a gradient-based hierarchical power management framework for power-limited ML inference environments. The main idea of PowerGrad is simple: identify, at runtime, how much the performance of each workload benefits from extra power, and hierarchically shift power in the datacenter from workloads that benefit the least to those that benefit the most. In practical terms, PowerGrad dynamically computes the derivative of each compute unit’s performance over power (i.e., the performance gradient), and shifts power from lower-gradient units to higher-gradient ones. PowerGrad shows a promising result in local CPU power control, automatically achieving high power efficiency using only hardware performance counters. The final proposal is FriendlyFoe, which dynamically applies Adversarial Machine Learning (AML) to obfuscate side channels. FriendlyFoe defines a workflow to design obfuscation DNNs called Defenders with low overhead and information leakage, and to customize them for different environments. Defenders are transferable, i.e., they thwart attacker classifiers that are different from those used to train the Defenders. They also resist adaptive attacks, where attackers train using the obfuscated signals collected while the Defender is active. Finally, the approach is general enough to be applicable to different environments. FriendlyFoe is demonstrated against two side channel attacks: one based on memory contention and one on system power. FriendlyFoe is an efficient obfuscation method to defend against hardware side channels. Compared to current defenses, FriendlyFoe either 1) reduces the performance overhead with a similar level of security or 2) improves the security with a similar level of performance overhead."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Architectures for machine learning and machine learning for architecture"]}]}],"canonical_facts":{"dc:contributor":["Torrellas, Josep","Li, Bo","Mendis, Charith","Bose, Pradip","Pothukuchi, Raghavendra","Kim, Nam Sung"],"dc:creator":["Nam, Hyoungwook"],"dc:date":["2025-07-10","2025-08"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-08-01","The student, Hyoungwook Nam, accepted the attached license on 2025-07-10 at 10:22.","The student, Hyoungwook Nam, submitted this Dissertation for approval on 2025-07-10 at 10:30.","This Dissertation was approved for publication on 2025-07-10 at 12:22.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22484 on 2025-10-25 at 15:30:58","The unprecedented success of machine learning (ML) has brought a problem and an opportunity to computer architecture research. The problem is how to build efficient and scalable computer systems for ML computations. The opportunity is how can ML methods help solving research problems in computer architecture. This dissertation explores both directions of research. In the direction of computer architecture for ML, this work focuses on two challenges for large-scale ML: scalability and power efficiency. This dissertation proposes two proposals for these: MeshSlice and PowerGrad. For the other direction, this dissertation proposes FriendlyFoe, which uses adversarial ML for hardware security. The first proposal is MeshSlice, a framework for efficient 2D tensor parallelism (TP) in distributed DNN training. MeshSlice consists of a novel 2D GeMM algorithm and an autotuner. The MeshSlice GeMM algorithm slices the collective communications into multiple partial collectives that allow overlapping communications with computations. As a result, MeshSlice hides most of the communication latency. The MeshSlice LLM autotuner automates finding the optimal configuration of 2D GeMM dataflow, the mesh shape, and the communication granularity using an analytical cost model. MeshSlice shows significant speedup in LLM training workloads compared to the state-of-the-art 2D TP method. The second proposal is PowerGrad, a gradient-based hierarchical power management framework for power-limited ML inference environments. The main idea of PowerGrad is simple: identify, at runtime, how much the performance of each workload benefits from extra power, and hierarchically shift power in the datacenter from workloads that benefit the least to those that benefit the most. In practical terms, PowerGrad dynamically computes the derivative of each compute unit’s performance over power (i.e., the performance gradient), and shifts power from lower-gradient units to higher-gradient ones. PowerGrad shows a promising result in local CPU power control, automatically achieving high power efficiency using only hardware performance counters. The final proposal is FriendlyFoe, which dynamically applies Adversarial Machine Learning (AML) to obfuscate side channels. FriendlyFoe defines a workflow to design obfuscation DNNs called Defenders with low overhead and information leakage, and to customize them for different environments. Defenders are transferable, i.e., they thwart attacker classifiers that are different from those used to train the Defenders. They also resist adaptive attacks, where attackers train using the obfuscated signals collected while the Defender is active. Finally, the approach is general enough to be applicable to different environments. FriendlyFoe is demonstrated against two side channel attacks: one based on memory contention and one on system power. FriendlyFoe is an efficient obfuscation method to defend against hardware side channels. Compared to current defenses, FriendlyFoe either 1) reduces the performance overhead with a similar level of security or 2) improves the security with a similar level of performance overhead."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/130094"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Hyoungwook Nam"],"dc:subject":["Machine Learning","Computer Architecture"],"dc:title":["Architectures for machine learning and machine learning for architecture"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:06Z"}