{"id":{"repo_id":"lethbridge","oai_identifier":"oai:opus.uleth.ca:10133/6939"},"canonical_url":"https://search.dev.ndltd.org/etd/lethbridge/oai:opus.uleth.ca:10133/6939","repository":{"repo_id":"lethbridge","name":"University of Lethbridge","base_url":"https://opus.uleth.ca/server/oai/request"},"display":{"title":"Cost-effective batch-based migration strategies for NewSQL-based big data systems","abstract":"Modern, high-performance applications demand scalable and efficient databases, leading to the evolution of NewSQL systems. The challenge lies in migrating data from Shardingsphere with PostgreSQL to AWS (AmazonWeb Services) cloud object storage. Implementing batch migration algorithms in Apache Spark, specifically targeting Delta Lake format, introduces complexities to ensure seamless data integration and storage within AWS environments. This thesis explores tailored batch-based migration algorithms for transferring data from Shardingsphere with PostgreSQL to AWS cloud object storage, emphasizing performance optimization by transferring the data faster. The study evaluates various batch loading techniques in Apache Spark, including sequential and concurrent strategies for shard-by-shard and aggregated-shards based algorithms. These techniques aim to maximize efficiency in storing data in Delta Lake format within AWS cloud storage, facilitating effective data management, visualization, and utilization for modern applications, business intelligence, AI and ML. Leveraging the Lakehouse architecture for integrated data processing and analytics.","abstract_html":"Modern, high-performance applications demand scalable and efficient databases, leading to the evolution of NewSQL systems. The challenge lies in migrating data from Shardingsphere with PostgreSQL to AWS (AmazonWeb Services) cloud object storage. Implementing batch migration algorithms in Apache Spark, specifically targeting Delta Lake format, introduces complexities to ensure seamless data integration and storage within AWS environments. This thesis explores tailored batch-based migration algorithms for transferring data from Shardingsphere with PostgreSQL to AWS cloud object storage, emphasizing performance optimization by transferring the data faster. The study evaluates various batch loading techniques in Apache Spark, including sequential and concurrent strategies for shard-by-shard and aggregated-shards based algorithms. These techniques aim to maximize efficiency in storing data in Delta Lake format within AWS cloud storage, facilitating effective data management, visualization, and utilization for modern applications, business intelligence, AI and ML. Leveraging the Lakehouse architecture for integrated data processing and analytics.","abstract_has_math":false,"creators":["Vadlamudi, Naveen Kumar","University of Lethbridge. Faculty of Arts and Science"],"institution":"Lethbridge, Alta. : University of Lethbridge, Dept. of Mathematics and Computer Science","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Osborn, Wendy"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024","date_published":"2024","updated_at":"2026-08-21T16:45:58Z","subjects":["batch-based migration algorithms","cloud computing","NewSQL systems","data migration","data pipelines","documentation"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["hdl:10133/6939"],"render_values":[{"text":"hdl:10133/6939","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10133/6939","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"source_record":{"url":"https://opus.uleth.ca/server/oai/request?verb=GetRecord&metadataPrefix=dim&identifier=oai%3Aopus.uleth.ca%3A10133%2F6939","prefix":"dim"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.supervisor","label":"Supervisor","values":["Osborn, Wendy"]},{"key":"dc:creator","label":"Author","values":["Vadlamudi, Naveen Kumar","University of Lethbridge. Faculty of Arts and Science"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2024-10-10T20:42:31Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2024-10-10T20:42:31Z"]},{"key":"dc:date.issued","label":"Date","values":["2024"]},{"key":"dc:publisher","label":"Institution","values":["Lethbridge, Alta. : University of Lethbridge, Dept. of Mathematics and Computer Science"]},{"key":"dc:publisher.department","label":"Dc Publisher Department","values":["Department of Mathematics and Computer Science"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["batch-based migration algorithms","cloud computing","NewSQL systems","data migration","data pipelines","documentation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["hdl:10133/6939"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10133/6939"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Modern, high-performance applications demand scalable and efficient databases, leading to the evolution of NewSQL systems. The challenge lies in migrating data from Shardingsphere with PostgreSQL to AWS (AmazonWeb Services) cloud object storage. Implementing batch migration algorithms in Apache Spark, specifically targeting Delta Lake format, introduces complexities to ensure seamless data integration and storage within AWS environments. This thesis explores tailored batch-based migration algorithms for transferring data from Shardingsphere with PostgreSQL to AWS cloud object storage, emphasizing performance optimization by transferring the data faster. The study evaluates various batch loading techniques in Apache Spark, including sequential and concurrent strategies for shard-by-shard and aggregated-shards based algorithms. These techniques aim to maximize efficiency in storing data in Delta Lake format within AWS cloud storage, facilitating effective data management, visualization, and utilization for modern applications, business intelligence, AI and ML. Leveraging the Lakehouse architecture for integrated data processing and analytics."]},{"key":"dc:description.other","label":"Dc Description Other","values":["Modern, high-performance applications demand scalable and efficient databases, leading to the evolution of NewSQL systems. The challenge lies in migrating data from Shardingsphere with PostgreSQL to AWS (AmazonWeb Services) cloud object storage. Implementing batch migration algorithms in Apache Spark, specifically targeting Delta Lake format, introduces complexities to ensure seamless data integration and storage within AWS environments. This thesis explores tailored batch-based migration algorithms for transferring data from Shardingsphere with PostgreSQL to AWS cloud object storage, emphasizing performance optimization by transferring the data faster. The study evaluates various batch loading techniques in Apache Spark, including sequential and concurrent strategies for shard-by-shard and aggregated-shards based algorithms. These techniques aim to maximize efficiency in storing data in Delta Lake format within AWS cloud storage, facilitating effective data management, visualization, and utilization for modern applications, business intelligence, AI and ML. Leveraging the Lakehouse architecture for integrated data processing and analytics."]},{"key":"dc:title","label":"Title","values":["Cost-effective batch-based migration strategies for NewSQL-based big data systems"]}]}],"canonical_facts":{"dc:contributor.supervisor":["Osborn, Wendy"],"dc:creator":["Vadlamudi, Naveen Kumar","University of Lethbridge. Faculty of Arts and Science"],"dc:date.accessioned":["2024-10-10T20:42:31Z"],"dc:date.available":["2024-10-10T20:42:31Z"],"dc:date.issued":["2024"],"dc:description.abstract":["Modern, high-performance applications demand scalable and efficient databases, leading to the evolution of NewSQL systems. The challenge lies in migrating data from Shardingsphere with PostgreSQL to AWS (AmazonWeb Services) cloud object storage. Implementing batch migration algorithms in Apache Spark, specifically targeting Delta Lake format, introduces complexities to ensure seamless data integration and storage within AWS environments. This thesis explores tailored batch-based migration algorithms for transferring data from Shardingsphere with PostgreSQL to AWS cloud object storage, emphasizing performance optimization by transferring the data faster. The study evaluates various batch loading techniques in Apache Spark, including sequential and concurrent strategies for shard-by-shard and aggregated-shards based algorithms. These techniques aim to maximize efficiency in storing data in Delta Lake format within AWS cloud storage, facilitating effective data management, visualization, and utilization for modern applications, business intelligence, AI and ML. Leveraging the Lakehouse architecture for integrated data processing and analytics."],"dc:description.other":["Modern, high-performance applications demand scalable and efficient databases, leading to the evolution of NewSQL systems. The challenge lies in migrating data from Shardingsphere with PostgreSQL to AWS (AmazonWeb Services) cloud object storage. Implementing batch migration algorithms in Apache Spark, specifically targeting Delta Lake format, introduces complexities to ensure seamless data integration and storage within AWS environments. This thesis explores tailored batch-based migration algorithms for transferring data from Shardingsphere with PostgreSQL to AWS cloud object storage, emphasizing performance optimization by transferring the data faster. The study evaluates various batch loading techniques in Apache Spark, including sequential and concurrent strategies for shard-by-shard and aggregated-shards based algorithms. These techniques aim to maximize efficiency in storing data in Delta Lake format within AWS cloud storage, facilitating effective data management, visualization, and utilization for modern applications, business intelligence, AI and ML. Leveraging the Lakehouse architecture for integrated data processing and analytics."],"dc:identifier":["hdl:10133/6939"],"dc:identifier.uri":["https://hdl.handle.net/10133/6939"],"dc:language.iso":["en"],"dc:publisher":["Lethbridge, Alta. : University of Lethbridge, Dept. of Mathematics and Computer Science"],"dc:publisher.department":["Department of Mathematics and Computer Science"],"dc:subject":["batch-based migration algorithms","cloud computing","NewSQL systems","data migration","data pipelines","documentation"],"dc:title":["Cost-effective batch-based migration strategies for NewSQL-based big data systems"],"dc:type":["Thesis"]},"updated_at":"2026-08-21T16:45:58Z"}