{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/90503"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/90503","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Dynamic resource provisioning for data center workloads with data constraints","abstract":"Dynamic resource provisioning, as an important data center software building block, helps to achieve high resource usage efficiency, leading to enormous monetary benefits. Most existing work for data center dynamic provisioning target on stateless servers, where any request can be routed to any server. However, the assumption of stateless behaviors no longer holds for subsystems that subject to data constraints, as a request may depend on a certain dataset stored on a small subset of servers. Routing a request to a server without the required dataset violates data locality or data availability properties, which may negatively impact on the response times. To solve this problem, this thesis provides an unified framework consisting of two main steps: 1) determining the proper amount of resources to serve the workload by analyzing the schedulability utilization bound; 2) avoiding transition penalties during cluster resizing operations by deliberately design data distribution policies. We apply this framework to both storage and computing subsystems, where the former includes distributed file systems, databases, memory caches, and the latter refers to systems such as Hadoop, Spark, and Storm. Proposed solutions are implemented into MemCached, HBase/HDFS, and Spark, and evaluated using various datasets, including Wikipedia, NYC taxi trace, Twitter traces, etc.","abstract_html":"Dynamic resource provisioning, as an important data center software building block, helps to achieve high resource usage efficiency, leading to enormous monetary benefits. Most existing work for data center dynamic provisioning target on stateless servers, where any request can be routed to any server. However, the assumption of stateless behaviors no longer holds for subsystems that subject to data constraints, as a request may depend on a certain dataset stored on a small subset of servers. Routing a request to a server without the required dataset violates data locality or data availability properties, which may negatively impact on the response times. To solve this problem, this thesis provides an unified framework consisting of two main steps: 1) determining the proper amount of resources to serve the workload by analyzing the schedulability utilization bound; 2) avoiding transition penalties during cluster resizing operations by deliberately design data distribution policies. We apply this framework to both storage and computing subsystems, where the former includes distributed file systems, databases, memory caches, and the latter refers to systems such as Hadoop, Spark, and Storm. Proposed solutions are implemented into MemCached, HBase/HDFS, and Spark, and evaluated using various datasets, including Wikipedia, NYC taxi trace, Twitter traces, etc.","abstract_has_math":false,"creators":["Li, Shen"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Abdelzaher, Tarek F.","Gupta, Indranil","Liu, Jie","Sha, Lui"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-07-07T19:53:19Z","date_published":"2016-07-07T19:53:19Z","updated_at":"2026-07-22T22:26:32Z","subjects":["Data Center","Dynamic Provisioning","Big Data","Scheduling","Distributed Systems"],"languages":["en"],"rights":["Copyright 2016 Shen Li"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/90503","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Abdelzaher, Tarek F.","Gupta, Indranil","Liu, Jie","Sha, Lui"]},{"key":"dc:creator","label":"Author","values":["Li, Shen"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-07-07T19:53:19Z","2016-04-07","2016-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Data Center","Dynamic Provisioning","Big Data","Scheduling","Distributed Systems"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2016 Shen Li"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/90503"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Dynamic resource provisioning, as an important data center software building block, helps to achieve high resource usage efficiency, leading to enormous monetary benefits. Most existing work for data center dynamic provisioning target on stateless servers, where any request can be routed to any server. However, the assumption of stateless behaviors no longer holds for subsystems that subject to data constraints, as a request may depend on a certain dataset stored on a small subset of servers. Routing a request to a server without the required dataset violates data locality or data availability properties, which may negatively impact on the response times. To solve this problem, this thesis provides an unified framework consisting of two main steps: 1) determining the proper amount of resources to serve the workload by analyzing the schedulability utilization bound; 2) avoiding transition penalties during cluster resizing operations by deliberately design data distribution policies. We apply this framework to both storage and computing subsystems, where the former includes distributed file systems, databases, memory caches, and the latter refers to systems such as Hadoop, Spark, and Storm. Proposed solutions are implemented into MemCached, HBase/HDFS, and Spark, and evaluated using various datasets, including Wikipedia, NYC taxi trace, Twitter traces, etc.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-07-07 without embargo terms","The student, Shen Li, accepted the attached license on 2016-04-04 at 11:18.","The student, Shen Li, submitted this Dissertation for approval on 2016-04-04 at 11:24.","This Dissertation was approved for publication on 2016-04-07 at 11:39.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9143 on 2016-07-07 at 13:28:33","Made available in DSpace on 2016-07-07T19:53:19Z (GMT). No. of bitstreams: 3 LI-DISSERTATION-2016.pdf: 5340852 bytes, checksum: 58196e76bad299845438afdd51e589a1 (MD5) LICENSE.txt: 4204 bytes, checksum: bc35961e0a82fa0fed3b184cf7e9e0f2 (MD5) PROQUEST_LICENSE.txt: 4550 bytes, checksum: 459b8882dd891e9016aac275cda58be0 (MD5) Previous issue date: 2016-04-07"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Dynamic resource provisioning for data center workloads with data constraints"]}]}],"canonical_facts":{"dc:contributor":["Abdelzaher, Tarek F.","Gupta, Indranil","Liu, Jie","Sha, Lui"],"dc:creator":["Li, Shen"],"dc:date":["2016-07-07T19:53:19Z","2016-04-07","2016-05"],"dc:description":["Dynamic resource provisioning, as an important data center software building block, helps to achieve high resource usage efficiency, leading to enormous monetary benefits. Most existing work for data center dynamic provisioning target on stateless servers, where any request can be routed to any server. However, the assumption of stateless behaviors no longer holds for subsystems that subject to data constraints, as a request may depend on a certain dataset stored on a small subset of servers. Routing a request to a server without the required dataset violates data locality or data availability properties, which may negatively impact on the response times. To solve this problem, this thesis provides an unified framework consisting of two main steps: 1) determining the proper amount of resources to serve the workload by analyzing the schedulability utilization bound; 2) avoiding transition penalties during cluster resizing operations by deliberately design data distribution policies. We apply this framework to both storage and computing subsystems, where the former includes distributed file systems, databases, memory caches, and the latter refers to systems such as Hadoop, Spark, and Storm. Proposed solutions are implemented into MemCached, HBase/HDFS, and Spark, and evaluated using various datasets, including Wikipedia, NYC taxi trace, Twitter traces, etc.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-07-07 without embargo terms","The student, Shen Li, accepted the attached license on 2016-04-04 at 11:18.","The student, Shen Li, submitted this Dissertation for approval on 2016-04-04 at 11:24.","This Dissertation was approved for publication on 2016-04-07 at 11:39.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9143 on 2016-07-07 at 13:28:33","Made available in DSpace on 2016-07-07T19:53:19Z (GMT). No. of bitstreams: 3 LI-DISSERTATION-2016.pdf: 5340852 bytes, checksum: 58196e76bad299845438afdd51e589a1 (MD5) LICENSE.txt: 4204 bytes, checksum: bc35961e0a82fa0fed3b184cf7e9e0f2 (MD5) PROQUEST_LICENSE.txt: 4550 bytes, checksum: 459b8882dd891e9016aac275cda58be0 (MD5) Previous issue date: 2016-04-07"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/90503"],"dc:language":["en"],"dc:rights":["Copyright 2016 Shen Li"],"dc:subject":["Data Center","Dynamic Provisioning","Big Data","Scheduling","Distributed Systems"],"dc:title":["Dynamic resource provisioning for data center workloads with data constraints"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:32Z"}