{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/116194"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/116194","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Infrastructure to enable and exploit GPU orchestrated high-throughput storage access","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_has_math":false,"creators":["Qureshi, Zaid"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Hwu, Wen-mei","Torrellas, Josep","Padua, David","Chung, I-Hsin"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08","date_published":"2022-08","updated_at":"2026-07-22T22:24:55Z","subjects":["GPU","memory","storage","HPC","graph analytics","data analytics","memory wall","SSD","Flash","software cache","caching","accelerators","memory system","storage system","memory capacity","memory capacity wall"],"languages":["en","eng"],"rights":["Copyright 2022 Zaid Qureshi"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/116194","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hwu, Wen-mei","Torrellas, Josep","Padua, David","Chung, I-Hsin"]},{"key":"dc:creator","label":"Author","values":["Qureshi, Zaid"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-08","2022-07-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["GPU","memory","storage","HPC","graph analytics","data analytics","memory wall","SSD","Flash","software cache","caching","accelerators","memory system","storage system","memory capacity","memory capacity wall"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Zaid Qureshi"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/116194"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Zaid Qureshi, accepted the attached license on 2022-07-08 at 15:49.","The student, Zaid Qureshi, submitted this Dissertation for approval on 2022-07-08 at 16:05.","This Dissertation was approved for publication on 2022-07-08 at 17:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18174 on 2022-11-15 at 17:38:18","Graphics Processing Units (GPUs) have traditionally relied on the CPU to orchestrate access to the data storage. This approach is well-suited for GPU applications with known data access patterns that enable partitioning of their data-set to be processed in a pipelined fashion in the GPU. However, many emerging applications, such as graph and data analytics, recommender systems, or graph neural networks, require fine-grained, data-dependent access to storage. CPU orchestration of storage access is unsuitable for these applications due to high CPU-GPU synchronization overheads, I/O traffic amplification, and long CPU processing latencies. GPU self-orchestrated storage access avoids these overheads by removing the CPU from the storage control path and, thus, seems better-suited for these applications. However, existing system architectures and software infrastructure lack support for such GPU-orchestrated storage access. In this work, we present a novel system architecture, BaM, that offers mechanisms for GPU code to efficiently access storage and enables GPU self-orchestrated storage access. BaM features a fined-grained, scalable software cache to coalesce data storage requests while minimizing I/O amplification effects. This software cache communicates with the storage system through high-throughput queues that enable the massive number of concurrent threads in modern GPUs to generate I/O requests at a sufficiently high rate to fully utilize the available bandwidth of the interconnect and the storage system. Furthermore, we provide array-based abstractions that not only make integrating BaM into GPU kernels trivial for programmers but also transparently optimize the number of BaM cache accesses by exploiting common GPU thread access patterns. We evaluate the end-to-end performance and efficiency impact of each optimization for each layer of BaM’s software stack with a variety of workloads with multiple data-sets. Experimental results show that GPU self-orchestrated storage access running on BaM delivers 1.04× and 1.05× end-to-end speed up for BFS and CC graph analytics. Our experiments also show GPU self-orchestrated storage access speeds up data-analytics workloads by 4.9× when running on the same hardware. In this work, we show that with carefully optimized systems software on the GPU, it is possible to use solid-state storage as a means to extend the GPU’s effective memory capacity since it provides performance (up to 4.62×), cost (up to 21.8×), and I/O efficiency (up to 3.72×) benefits, even over much more expensive state-of-the-art solutions using fast DRAM, for important GPU accelerated applications."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Infrastructure to enable and exploit GPU orchestrated high-throughput storage access"]}]}],"canonical_facts":{"dc:contributor":["Hwu, Wen-mei","Torrellas, Josep","Padua, David","Chung, I-Hsin"],"dc:creator":["Qureshi, Zaid"],"dc:date":["2022-08","2022-07-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Zaid Qureshi, accepted the attached license on 2022-07-08 at 15:49.","The student, Zaid Qureshi, submitted this Dissertation for approval on 2022-07-08 at 16:05.","This Dissertation was approved for publication on 2022-07-08 at 17:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18174 on 2022-11-15 at 17:38:18","Graphics Processing Units (GPUs) have traditionally relied on the CPU to orchestrate access to the data storage. This approach is well-suited for GPU applications with known data access patterns that enable partitioning of their data-set to be processed in a pipelined fashion in the GPU. However, many emerging applications, such as graph and data analytics, recommender systems, or graph neural networks, require fine-grained, data-dependent access to storage. CPU orchestration of storage access is unsuitable for these applications due to high CPU-GPU synchronization overheads, I/O traffic amplification, and long CPU processing latencies. GPU self-orchestrated storage access avoids these overheads by removing the CPU from the storage control path and, thus, seems better-suited for these applications. However, existing system architectures and software infrastructure lack support for such GPU-orchestrated storage access. In this work, we present a novel system architecture, BaM, that offers mechanisms for GPU code to efficiently access storage and enables GPU self-orchestrated storage access. BaM features a fined-grained, scalable software cache to coalesce data storage requests while minimizing I/O amplification effects. This software cache communicates with the storage system through high-throughput queues that enable the massive number of concurrent threads in modern GPUs to generate I/O requests at a sufficiently high rate to fully utilize the available bandwidth of the interconnect and the storage system. Furthermore, we provide array-based abstractions that not only make integrating BaM into GPU kernels trivial for programmers but also transparently optimize the number of BaM cache accesses by exploiting common GPU thread access patterns. We evaluate the end-to-end performance and efficiency impact of each optimization for each layer of BaM’s software stack with a variety of workloads with multiple data-sets. Experimental results show that GPU self-orchestrated storage access running on BaM delivers 1.04× and 1.05× end-to-end speed up for BFS and CC graph analytics. Our experiments also show GPU self-orchestrated storage access speeds up data-analytics workloads by 4.9× when running on the same hardware. In this work, we show that with carefully optimized systems software on the GPU, it is possible to use solid-state storage as a means to extend the GPU’s effective memory capacity since it provides performance (up to 4.62×), cost (up to 21.8×), and I/O efficiency (up to 3.72×) benefits, even over much more expensive state-of-the-art solutions using fast DRAM, for important GPU accelerated applications."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/116194"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Zaid Qureshi"],"dc:subject":["GPU","memory","storage","HPC","graph analytics","data analytics","memory wall","SSD","Flash","software cache","caching","accelerators","memory system","storage system","memory capacity","memory capacity wall"],"dc:title":["Infrastructure to enable and exploit GPU orchestrated high-throughput storage access"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:55Z"}