{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/102420"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/102420","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Efficient data reconfiguration for today's cloud systems","abstract":"DSpace SAF Submission Ingestion Package generated from Vireo submission #13052 on 2019-02-05 at 11:08:44","abstract_html":"DSpace SAF Submission Ingestion Package generated from Vireo submission #13052 on 2019-02-05 at 11:08:44","abstract_has_math":false,"creators":["Ghosh, Mainak"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Gupta, Indranil","Vaidya, Nitin","Olson, Luke","Elmore, Aaron"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-02-06T19:32:48Z","date_published":"2019-02-06T19:32:48Z","updated_at":"2026-07-22T22:24:40Z","subjects":["reconfiguration","partitioning","replication","prefetching","compaction","nosql databases","interactive analytics engines","tiered storage systems"],"languages":["en"],"rights":["2018 Mainak Ghosh"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/102420","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Gupta, Indranil","Vaidya, Nitin","Olson, Luke","Elmore, Aaron"]},{"key":"dc:creator","label":"Author","values":["Ghosh, Mainak"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-02-06T19:32:48Z","2018-11-12","2018-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["reconfiguration","partitioning","replication","prefetching","compaction","nosql databases","interactive analytics engines","tiered storage systems"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["2018 Mainak Ghosh"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/102420"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["DSpace SAF Submission Ingestion Package generated from Vireo submission #13052 on 2019-02-05 at 11:08:44","Made available in DSpace on 2019-02-06T19:32:48Z (GMT). No. of bitstreams: 2 GHOSH-DISSERTATION-2018.pdf: 5484335 bytes, checksum: 8c90f90b5f130b0c44209847bcaf0f2b (MD5) LICENSE.txt: 4209 bytes, checksum: 9d9dadc04fb0b5269b4c11d92e3ed858 (MD5) Previous issue date: 2018-11-12","Performance of big data systems largely relies on efficient data reconfiguration techniques. Data reconfiguration operations deal with changing configuration parameters that affect data layout in a system. They could be user-initiated like changing shard key, block size in NoSQL databases, or system-initiated like changing replication in distributed interactive analytics engine. Current data reconfiguration schemes are heuristics at best and often do not scale well as data volume grows. As a result, system performance suffers. In this thesis, we show that {\\it data reconfiguration mechanisms can be done in the background by using new optimal or near-optimal algorithms coupling them with performant system designs}. We explore four different data reconfiguration operations affecting three popular types of systems -- storage, real-time analytics and batch analytics. In NoSQL databases (storage), we explore new strategies for changing table-level configuration and for compaction as they improve read/write latencies. In distributed interactive analytics engines, a good replication algorithm can save costs by judiciously using memory that is sufficient to provide the highest throughput and low latency for queries. Finally, in batch processing systems, we explore prefetching and caching strategies that can improve the number of production jobs meeting their SLOs. All these operations happen in the background without affecting the fast path. Our contributions in each of the problems are two-fold -- 1) we model the problem and design algorithms inspired from well-known theoretical abstractions, 2) we design and build a system on top of popular open source systems used in companies today. Finally, using real-life workloads, we evaluate the efficacy of our solutions. Morphus and Parqua provide several 9s of availability while changing table level configuration parameters in databases. By halving memory usage in distributed interactive analytics engine, Getafix reduces cost of deploying the system by 10 million dollars annually and improves query throughput. We are the first to model the problem of compaction and provide formal bounds on their runtime. Finally, NetCachier helps 30\\% more production jobs to meet their SLOs compared to existing state-of-the-art.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-02-05 without embargo terms","The student, Mainak Ghosh, accepted the attached license on 2018-10-30 at 01:27.","The student, Mainak Ghosh, submitted this Dissertation for approval on 2018-10-30 at 01:54.","This Dissertation was approved for publication on 2018-11-12 at 11:52."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Efficient data reconfiguration for today's cloud systems"]}]}],"canonical_facts":{"dc:contributor":["Gupta, Indranil","Vaidya, Nitin","Olson, Luke","Elmore, Aaron"],"dc:creator":["Ghosh, Mainak"],"dc:date":["2019-02-06T19:32:48Z","2018-11-12","2018-12"],"dc:description":["DSpace SAF Submission Ingestion Package generated from Vireo submission #13052 on 2019-02-05 at 11:08:44","Made available in DSpace on 2019-02-06T19:32:48Z (GMT). No. of bitstreams: 2 GHOSH-DISSERTATION-2018.pdf: 5484335 bytes, checksum: 8c90f90b5f130b0c44209847bcaf0f2b (MD5) LICENSE.txt: 4209 bytes, checksum: 9d9dadc04fb0b5269b4c11d92e3ed858 (MD5) Previous issue date: 2018-11-12","Performance of big data systems largely relies on efficient data reconfiguration techniques. Data reconfiguration operations deal with changing configuration parameters that affect data layout in a system. They could be user-initiated like changing shard key, block size in NoSQL databases, or system-initiated like changing replication in distributed interactive analytics engine. Current data reconfiguration schemes are heuristics at best and often do not scale well as data volume grows. As a result, system performance suffers. In this thesis, we show that {\\it data reconfiguration mechanisms can be done in the background by using new optimal or near-optimal algorithms coupling them with performant system designs}. We explore four different data reconfiguration operations affecting three popular types of systems -- storage, real-time analytics and batch analytics. In NoSQL databases (storage), we explore new strategies for changing table-level configuration and for compaction as they improve read/write latencies. In distributed interactive analytics engines, a good replication algorithm can save costs by judiciously using memory that is sufficient to provide the highest throughput and low latency for queries. Finally, in batch processing systems, we explore prefetching and caching strategies that can improve the number of production jobs meeting their SLOs. All these operations happen in the background without affecting the fast path. Our contributions in each of the problems are two-fold -- 1) we model the problem and design algorithms inspired from well-known theoretical abstractions, 2) we design and build a system on top of popular open source systems used in companies today. Finally, using real-life workloads, we evaluate the efficacy of our solutions. Morphus and Parqua provide several 9s of availability while changing table level configuration parameters in databases. By halving memory usage in distributed interactive analytics engine, Getafix reduces cost of deploying the system by 10 million dollars annually and improves query throughput. We are the first to model the problem of compaction and provide formal bounds on their runtime. Finally, NetCachier helps 30\\% more production jobs to meet their SLOs compared to existing state-of-the-art.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-02-05 without embargo terms","The student, Mainak Ghosh, accepted the attached license on 2018-10-30 at 01:27.","The student, Mainak Ghosh, submitted this Dissertation for approval on 2018-10-30 at 01:54.","This Dissertation was approved for publication on 2018-11-12 at 11:52."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/102420"],"dc:language":["en"],"dc:rights":["2018 Mainak Ghosh"],"dc:subject":["reconfiguration","partitioning","replication","prefetching","compaction","nosql databases","interactive analytics engines","tiered storage systems"],"dc:title":["Efficient data reconfiguration for today's cloud systems"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:40Z"}