{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/138562"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/138562","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Programming Abstractions for Scalable and High-Performance Memory Disaggregation","abstract":"Memory-intensive applications, such as in-memory analytics, databases, and caching, are driving the memory demand in datacenters. Unfortunately, DRAM scaling is challenging as the technology reaches its physical limit. Although its demand continues to rise, memory is an underutilized resource within datacenters while simultaneously accounting for a significant portion of the total infrastructure cost. Consequently, memory capacity and performance have become significant bottlenecks, impacting both the overall system cost and efficiency. To address the aforementioned inefficiency, memory disaggregation has emerged as a promising solution. Memory disaggregation physically separates compute from memory, allowing each resource to be separately provisioned to enable independent resource scaling. Applications run at the compute nodes that have a small amount of memory and access the disaggregated memory over a fast network/fabric technology, such as Remote Direct Memory Access (RDMA), and the recent Compute Express Link (CXL). The disaggregated architecture is promising as it can efficiently utilize the compute and memory resources. However, despite its promise, memory disaggregation must overcome several challenges to enable its widespread adoption. First, the scalability of the disaggregated architecture depends on the scalability of the components driving it. RDMA network interconnects are a key enabler to make memory disaggregation practical; hence, a scalable RDMA communication framework is imperative to achieving system scalability. However, RDMA performance does not scale with the cluster size, pointing towards a fundamental trade-off between performance and scalability in RDMA-capable networks. Second, memory disaggregation is a paradigm shift from the conventional model of programming with machine-local memory as the application state could be distributed across disaggregated memory. A lack of programming abstractions that offer easy-to-use interfaces would limit the applicability of memory disaggregation. Finally, apart from the efficiency gains, performance is crucial to the adoption of memory disaggregation; hence, the system must deliver good performance in this architecture. This dissertation aims to overcome the challenges listed above. I will first describe my prior work, FLOCK, which achieves scalability and high performance in RDMA-capable networks. FLOCK exposes a connection handle abstraction and uses it to transparently coalesce network messages at the application to efficiently utilize the network bandwidth and reduce CPU cycles spent on network I/O. FLOCK introduces a symbiotic load-control scheduling policy that enables a server to dynamically control the maximum network load it wants to handle, which is key to achieving scalability. This dissertation also proposes a programmable, scalable, and high-performance disaggregated memory system, SONIC. In this work, we offer programming abstractions that are user-friendly and general to make programming for disaggregated memory similar to working with machine-local memory. These abstractions can be applied to a wide variety of applications, demonstrating their generality and benefiting many applications from disaggregated memory. The key idea in SONIC is using transactions that enable applications to access disaggregated memory with location transparency. We present multiple commit protocols for efficient transaction processing in our system. Accessing disaggregated memory over the network is slower than local memory; hence, our design incorporates several optimizations to reduce the network round trips and deliver good performance, demonstrating that general-purpose programming abstractions can deliver good performance in the disaggregated architecture. Finally, this dissertation presents directions for future research on memory disaggregation. I will discuss my ongoing work on extending SONIC to leverage application knowledge for better remote memory management that can improve application performance while efficiently utilizing system resources (compute, memory, and network). Furthermore, I will discuss the challenges and interesting problems that memory disaggregation brings, such as fault tolerance, and how to address them.","abstract_html":"Memory-intensive applications, such as in-memory analytics, databases, and caching, are driving the memory demand in datacenters. Unfortunately, DRAM scaling is challenging as the technology reaches its physical limit. Although its demand continues to rise, memory is an underutilized resource within datacenters while simultaneously accounting for a significant portion of the total infrastructure cost. Consequently, memory capacity and performance have become significant bottlenecks, impacting both the overall system cost and efficiency. To address the aforementioned inefficiency, memory disaggregation has emerged as a promising solution. Memory disaggregation physically separates compute from memory, allowing each resource to be separately provisioned to enable independent resource scaling. Applications run at the compute nodes that have a small amount of memory and access the disaggregated memory over a fast network/fabric technology, such as Remote Direct Memory Access (RDMA), and the recent Compute Express Link (CXL). The disaggregated architecture is promising as it can efficiently utilize the compute and memory resources. However, despite its promise, memory disaggregation must overcome several challenges to enable its widespread adoption. First, the scalability of the disaggregated architecture depends on the scalability of the components driving it. RDMA network interconnects are a key enabler to make memory disaggregation practical; hence, a scalable RDMA communication framework is imperative to achieving system scalability. However, RDMA performance does not scale with the cluster size, pointing towards a fundamental trade-off between performance and scalability in RDMA-capable networks. Second, memory disaggregation is a paradigm shift from the conventional model of programming with machine-local memory as the application state could be distributed across disaggregated memory. A lack of programming abstractions that offer easy-to-use interfaces would limit the applicability of memory disaggregation. Finally, apart from the efficiency gains, performance is crucial to the adoption of memory disaggregation; hence, the system must deliver good performance in this architecture. This dissertation aims to overcome the challenges listed above. I will first describe my prior work, FLOCK, which achieves scalability and high performance in RDMA-capable networks. FLOCK exposes a connection handle abstraction and uses it to transparently coalesce network messages at the application to efficiently utilize the network bandwidth and reduce CPU cycles spent on network I/O. FLOCK introduces a symbiotic load-control scheduling policy that enables a server to dynamically control the maximum network load it wants to handle, which is key to achieving scalability. This dissertation also proposes a programmable, scalable, and high-performance disaggregated memory system, SONIC. In this work, we offer programming abstractions that are user-friendly and general to make programming for disaggregated memory similar to working with machine-local memory. These abstractions can be applied to a wide variety of applications, demonstrating their generality and benefiting many applications from disaggregated memory. The key idea in SONIC is using transactions that enable applications to access disaggregated memory with location transparency. We present multiple commit protocols for efficient transaction processing in our system. Accessing disaggregated memory over the network is slower than local memory; hence, our design incorporates several optimizations to reduce the network round trips and deliver good performance, demonstrating that general-purpose programming abstractions can deliver good performance in the disaggregated architecture. Finally, this dissertation presents directions for future research on memory disaggregation. I will discuss my ongoing work on extending SONIC to leverage application knowledge for better remote memory management that can improve application performance while efficiently utilizing system resources (compute, memory, and network). Furthermore, I will discuss the challenges and interesting problems that memory disaggregation brings, such as fault tolerance, and how to address them.","abstract_has_math":false,"creators":["Monga, Sumit Kumar"],"institution":"Virginia Tech","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"Computer Engineering","degree_department":"Electrical and Computer Engineering","school":null,"contributors":[],"advisors":[],"committee_chairs":["Nikolopoulos, Dimitrios S.","Li, Huaicheng"],"committee_members":["Stavrou, Angelos","Butt, Ali","Min, Chang Woo"],"year":2025,"date_issued":"2025-10-22","date_published":"2025-10-22","updated_at":"2026-07-22T22:19:18Z","subjects":["Datacenter networking","Remote Direct Memory Access (RDMA)","Distributed Systems","Memory Disaggregation","Scalability"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44847"],"render_values":[{"text":"vt_gsexam:44847","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/138562","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Nikolopoulos, Dimitrios S.","Li, Huaicheng"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Stavrou, Angelos","Butt, Ali","Min, Chang Woo"]},{"key":"dc:contributor.department","label":"Department","values":["Electrical and Computer Engineering"]},{"key":"dc:creator","label":"Author","values":["Monga, Sumit Kumar"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-10-23T08:00:13Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-10-23T08:00:13Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-10-22"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Datacenter networking","Remote Direct Memory Access (RDMA)","Distributed Systems","Memory Disaggregation","Scalability"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44847"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/138562"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Memory-intensive applications, such as in-memory analytics, databases, and caching, are driving the memory demand in datacenters. Unfortunately, DRAM scaling is challenging as the technology reaches its physical limit. Although its demand continues to rise, memory is an underutilized resource within datacenters while simultaneously accounting for a significant portion of the total infrastructure cost. Consequently, memory capacity and performance have become significant bottlenecks, impacting both the overall system cost and efficiency. To address the aforementioned inefficiency, memory disaggregation has emerged as a promising solution. Memory disaggregation physically separates compute from memory, allowing each resource to be separately provisioned to enable independent resource scaling. Applications run at the compute nodes that have a small amount of memory and access the disaggregated memory over a fast network/fabric technology, such as Remote Direct Memory Access (RDMA), and the recent Compute Express Link (CXL). The disaggregated architecture is promising as it can efficiently utilize the compute and memory resources. However, despite its promise, memory disaggregation must overcome several challenges to enable its widespread adoption. First, the scalability of the disaggregated architecture depends on the scalability of the components driving it. RDMA network interconnects are a key enabler to make memory disaggregation practical; hence, a scalable RDMA communication framework is imperative to achieving system scalability. However, RDMA performance does not scale with the cluster size, pointing towards a fundamental trade-off between performance and scalability in RDMA-capable networks. Second, memory disaggregation is a paradigm shift from the conventional model of programming with machine-local memory as the application state could be distributed across disaggregated memory. A lack of programming abstractions that offer easy-to-use interfaces would limit the applicability of memory disaggregation. Finally, apart from the efficiency gains, performance is crucial to the adoption of memory disaggregation; hence, the system must deliver good performance in this architecture. This dissertation aims to overcome the challenges listed above. I will first describe my prior work, FLOCK, which achieves scalability and high performance in RDMA-capable networks. FLOCK exposes a connection handle abstraction and uses it to transparently coalesce network messages at the application to efficiently utilize the network bandwidth and reduce CPU cycles spent on network I/O. FLOCK introduces a symbiotic load-control scheduling policy that enables a server to dynamically control the maximum network load it wants to handle, which is key to achieving scalability. This dissertation also proposes a programmable, scalable, and high-performance disaggregated memory system, SONIC. In this work, we offer programming abstractions that are user-friendly and general to make programming for disaggregated memory similar to working with machine-local memory. These abstractions can be applied to a wide variety of applications, demonstrating their generality and benefiting many applications from disaggregated memory. The key idea in SONIC is using transactions that enable applications to access disaggregated memory with location transparency. We present multiple commit protocols for efficient transaction processing in our system. Accessing disaggregated memory over the network is slower than local memory; hence, our design incorporates several optimizations to reduce the network round trips and deliver good performance, demonstrating that general-purpose programming abstractions can deliver good performance in the disaggregated architecture. Finally, this dissertation presents directions for future research on memory disaggregation. I will discuss my ongoing work on extending SONIC to leverage application knowledge for better remote memory management that can improve application performance while efficiently utilizing system resources (compute, memory, and network). Furthermore, I will discuss the challenges and interesting problems that memory disaggregation brings, such as fault tolerance, and how to address them."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["End-user applications, such as search, social media, and recommendation systems, rely on memory-intensive datacenter services. The proliferation of these applications has led to rising memory demand in datacenters; consequently, memory accounts for up to one-third of their infrastructure cost. At the same time, memory is an underutilized resource as the traditional model of provisioning compute and memory within a single server doesn't allow datacenter operators to independently scale each resource. For instance, compute capacity can be increased by adding more servers, but this simultaneously increases memory capacity, thus unnecessarily increasing the system cost as well as memory underutilization. Memory disaggregation offers a solution to this problem by decoupling the provisioning of compute and memory such that each resource can be dynamically provisioned based on its demand. The disaggregated architecture physically separates compute from memory, wherein datacenter services running at the compute nodes access memory over a fast network interconnect, such as Remote Direct Memory Access (RDMA). Despite the advantages of memory disaggregation, its adoption is contingent on overcoming several challenges. First, as memory is accessed over RDMA network, the scalability of the disaggregated architecture depends on the scalability of the RDMA communication framework. Unfortunately, RDMA performance does not scale with the cluster size; hence, simultaneously achieving high performance and scalability remains a challenge with RDMA. Second, easy-to-use interfaces to program disaggregated memory would allow many datacenter services to benefit from memory disaggregation. However, existing systems either offer programmability through transparent interfaces that are not performant or they are specialized for particular services (e.g., key-value stores) to deliver high performance; hence, programmability and performance are not achieved together. This dissertation aims to overcome the aforementioned challenges. I will first describe FLOCK, a communication library that balances the performance-scalability trade-off in RDMA networks. Through symbiotic send-recv scheduling, a server in FLOCK dynamically controls the maximum network load it handles, preventing overload, which is key to achieving high performance and scalability in RDMA networks. This dissertation then discusses SONIC, a platform for programmable, scalable, and high-performance memory disaggregation. SONIC uses transactions as the core programming abstraction, which offers a local-memory-like interface to access disaggregated memory. SONIC delivers high performance by reducing the network accesses to disaggregated memory through caching, prefetching, and efficient transaction processing through our proposed commit protocols."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Programming Abstractions for Scalable and High-Performance Memory Disaggregation"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Nikolopoulos, Dimitrios S.","Li, Huaicheng"],"dc:contributor.committeemember":["Stavrou, Angelos","Butt, Ali","Min, Chang Woo"],"dc:contributor.department":["Electrical and Computer Engineering"],"dc:creator":["Monga, Sumit Kumar"],"dc:date.accessioned":["2025-10-23T08:00:13Z"],"dc:date.available":["2025-10-23T08:00:13Z"],"dc:date.issued":["2025-10-22"],"dc:description.abstract":["Memory-intensive applications, such as in-memory analytics, databases, and caching, are driving the memory demand in datacenters. Unfortunately, DRAM scaling is challenging as the technology reaches its physical limit. Although its demand continues to rise, memory is an underutilized resource within datacenters while simultaneously accounting for a significant portion of the total infrastructure cost. Consequently, memory capacity and performance have become significant bottlenecks, impacting both the overall system cost and efficiency. To address the aforementioned inefficiency, memory disaggregation has emerged as a promising solution. Memory disaggregation physically separates compute from memory, allowing each resource to be separately provisioned to enable independent resource scaling. Applications run at the compute nodes that have a small amount of memory and access the disaggregated memory over a fast network/fabric technology, such as Remote Direct Memory Access (RDMA), and the recent Compute Express Link (CXL). The disaggregated architecture is promising as it can efficiently utilize the compute and memory resources. However, despite its promise, memory disaggregation must overcome several challenges to enable its widespread adoption. First, the scalability of the disaggregated architecture depends on the scalability of the components driving it. RDMA network interconnects are a key enabler to make memory disaggregation practical; hence, a scalable RDMA communication framework is imperative to achieving system scalability. However, RDMA performance does not scale with the cluster size, pointing towards a fundamental trade-off between performance and scalability in RDMA-capable networks. Second, memory disaggregation is a paradigm shift from the conventional model of programming with machine-local memory as the application state could be distributed across disaggregated memory. A lack of programming abstractions that offer easy-to-use interfaces would limit the applicability of memory disaggregation. Finally, apart from the efficiency gains, performance is crucial to the adoption of memory disaggregation; hence, the system must deliver good performance in this architecture. This dissertation aims to overcome the challenges listed above. I will first describe my prior work, FLOCK, which achieves scalability and high performance in RDMA-capable networks. FLOCK exposes a connection handle abstraction and uses it to transparently coalesce network messages at the application to efficiently utilize the network bandwidth and reduce CPU cycles spent on network I/O. FLOCK introduces a symbiotic load-control scheduling policy that enables a server to dynamically control the maximum network load it wants to handle, which is key to achieving scalability. This dissertation also proposes a programmable, scalable, and high-performance disaggregated memory system, SONIC. In this work, we offer programming abstractions that are user-friendly and general to make programming for disaggregated memory similar to working with machine-local memory. These abstractions can be applied to a wide variety of applications, demonstrating their generality and benefiting many applications from disaggregated memory. The key idea in SONIC is using transactions that enable applications to access disaggregated memory with location transparency. We present multiple commit protocols for efficient transaction processing in our system. Accessing disaggregated memory over the network is slower than local memory; hence, our design incorporates several optimizations to reduce the network round trips and deliver good performance, demonstrating that general-purpose programming abstractions can deliver good performance in the disaggregated architecture. Finally, this dissertation presents directions for future research on memory disaggregation. I will discuss my ongoing work on extending SONIC to leverage application knowledge for better remote memory management that can improve application performance while efficiently utilizing system resources (compute, memory, and network). Furthermore, I will discuss the challenges and interesting problems that memory disaggregation brings, such as fault tolerance, and how to address them."],"dc:description.abstractgeneral":["End-user applications, such as search, social media, and recommendation systems, rely on memory-intensive datacenter services. The proliferation of these applications has led to rising memory demand in datacenters; consequently, memory accounts for up to one-third of their infrastructure cost. At the same time, memory is an underutilized resource as the traditional model of provisioning compute and memory within a single server doesn't allow datacenter operators to independently scale each resource. For instance, compute capacity can be increased by adding more servers, but this simultaneously increases memory capacity, thus unnecessarily increasing the system cost as well as memory underutilization. Memory disaggregation offers a solution to this problem by decoupling the provisioning of compute and memory such that each resource can be dynamically provisioned based on its demand. The disaggregated architecture physically separates compute from memory, wherein datacenter services running at the compute nodes access memory over a fast network interconnect, such as Remote Direct Memory Access (RDMA). Despite the advantages of memory disaggregation, its adoption is contingent on overcoming several challenges. First, as memory is accessed over RDMA network, the scalability of the disaggregated architecture depends on the scalability of the RDMA communication framework. Unfortunately, RDMA performance does not scale with the cluster size; hence, simultaneously achieving high performance and scalability remains a challenge with RDMA. Second, easy-to-use interfaces to program disaggregated memory would allow many datacenter services to benefit from memory disaggregation. However, existing systems either offer programmability through transparent interfaces that are not performant or they are specialized for particular services (e.g., key-value stores) to deliver high performance; hence, programmability and performance are not achieved together. This dissertation aims to overcome the aforementioned challenges. I will first describe FLOCK, a communication library that balances the performance-scalability trade-off in RDMA networks. Through symbiotic send-recv scheduling, a server in FLOCK dynamically controls the maximum network load it handles, preventing overload, which is key to achieving high performance and scalability in RDMA networks. This dissertation then discusses SONIC, a platform for programmable, scalable, and high-performance memory disaggregation. SONIC uses transactions as the core programming abstraction, which offers a local-memory-like interface to access disaggregated memory. SONIC delivers high performance by reducing the network accesses to disaggregated memory through caching, prefetching, and efficient transaction processing through our proposed commit protocols."],"dc:description.degree":["Doctor of Philosophy"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:44847"],"dc:identifier.uri":["https://hdl.handle.net/10919/138562"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Datacenter networking","Remote Direct Memory Access (RDMA)","Distributed Systems","Memory Disaggregation","Scalability"],"dc:title":["Programming Abstractions for Scalable and High-Performance Memory Disaggregation"],"dc:type":["Dissertation"],"thesis:degree_discipline":["Computer Engineering"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:18Z"}