{"id":{"repo_id":"columbus-state","oai_identifier":"oai:csuepress.columbusstate.edu:theses_dissertations-1395"},"canonical_url":"https://search.dev.ndltd.org/etd/columbus-state/oai:csuepress.columbusstate.edu:theses_dissertations-1395","repository":{"repo_id":"columbus-state","name":"Columbus State University","base_url":"https://csuepress.columbusstate.edu/do/oai/"},"display":{"title":"Improving Energy-Efficiency through Smart Data Placement in Hadoop Clusters","abstract":"<p>Hadoop, a pioneering open source framework, has revolutionized the big data world because of its ability to process vast amounts of unstructured and semi-structured data. This ability makes Hadoop the ‘go-to’ technology for many industries that generate big data, thus it also aids in being cost effective, unlike other legacy systems. Hadoop MapReduce is used in large scale data parallel applications to process massive amounts of data across a cluster and is used for scheduling, processing, and executing jobs. Basically, MapReduce is the right hand of Hadoop, as its library is needed to process these large data sets. In this research thesis, this study proposes a smart framework model that profiles MapReduce tasks with the use of Machine Learning (ML) algorithms to effectively place the data in Hadoop clusters; activate only sufficient number of nodes to accomplish the data processing within the planned deadline time for the task. The model will ensure achieving energy efficiency by utilizing the minimum number of necessary nodes, with maximum utilization and least energy consumption to reduce the overall cost of operations in data centers that deploy the Hadoop clusters.</p>","abstract_html":"&lt;p&gt;Hadoop, a pioneering open source framework, has revolutionized the big data world because of its ability to process vast amounts of unstructured and semi-structured data. This ability makes Hadoop the ‘go-to’ technology for many industries that generate big data, thus it also aids in being cost effective, unlike other legacy systems. Hadoop MapReduce is used in large scale data parallel applications to process massive amounts of data across a cluster and is used for scheduling, processing, and executing jobs. Basically, MapReduce is the right hand of Hadoop, as its library is needed to process these large data sets. In this research thesis, this study proposes a smart framework model that profiles MapReduce tasks with the use of Machine Learning (ML) algorithms to effectively place the data in Hadoop clusters; activate only sufficient number of nodes to accomplish the data processing within the planned deadline time for the task. The model will ensure achieving energy efficiency by utilizing the minimum number of necessary nodes, with maximum utilization and least energy consumption to reduce the overall cost of operations in data centers that deploy the Hadoop clusters.&lt;/p&gt;","abstract_has_math":false,"creators":["Mostafa, Ahmed"],"institution":null,"degree_name":"Computer Science - Applied Computing Track","degree_level":"Thesis","degree_discipline":"TSYS School of Computer Science","degree_department":null,"school":null,"contributors":["Dr. Yi Zhou","Dr. Rania Hodhod","Dr. Shamim Khan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-05-01T07:00:00Z","date_published":"2020-05-01T07:00:00Z","updated_at":"2026-07-24T01:45:01Z","subjects":["Hadoop","ML","Energy-Aware","Big Data","MapReduce","HDFS","Computer Sciences","Databases and Information Systems"],"languages":["English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://csuepress.columbusstate.edu/theses_dissertations/392","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dr. Yi Zhou","Dr. Rania Hodhod","Dr. Shamim Khan"]},{"key":"dc:creator","label":"Author","values":["Mostafa, Ahmed"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2020-06-25T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["TSYS School of Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Computer Science - Applied Computing Track"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Hadoop","ML","Energy-Aware","Big Data","MapReduce","HDFS","Computer Sciences","Databases and Information Systems"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://csuepress.columbusstate.edu/theses_dissertations/392"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Hadoop, a pioneering open source framework, has revolutionized the big data world because of its ability to process vast amounts of unstructured and semi-structured data. This ability makes Hadoop the ‘go-to’ technology for many industries that generate big data, thus it also aids in being cost effective, unlike other legacy systems. Hadoop MapReduce is used in large scale data parallel applications to process massive amounts of data across a cluster and is used for scheduling, processing, and executing jobs. Basically, MapReduce is the right hand of Hadoop, as its library is needed to process these large data sets. In this research thesis, this study proposes a smart framework model that profiles MapReduce tasks with the use of Machine Learning (ML) algorithms to effectively place the data in Hadoop clusters; activate only sufficient number of nodes to accomplish the data processing within the planned deadline time for the task. The model will ensure achieving energy efficiency by utilizing the minimum number of necessary nodes, with maximum utilization and least energy consumption to reduce the overall cost of operations in data centers that deploy the Hadoop clusters.</p>"]},{"key":"dc:title","label":"Title","values":["Improving Energy-Efficiency through Smart Data Placement in Hadoop Clusters"]}]}],"canonical_facts":{"dc:contributor":["Dr. Yi Zhou","Dr. Rania Hodhod","Dr. Shamim Khan"],"dc:creator":["Mostafa, Ahmed"],"dc:date.available":["2020-06-25T07:00:00Z"],"dc:description.abstract":["<p>Hadoop, a pioneering open source framework, has revolutionized the big data world because of its ability to process vast amounts of unstructured and semi-structured data. This ability makes Hadoop the ‘go-to’ technology for many industries that generate big data, thus it also aids in being cost effective, unlike other legacy systems. Hadoop MapReduce is used in large scale data parallel applications to process massive amounts of data across a cluster and is used for scheduling, processing, and executing jobs. Basically, MapReduce is the right hand of Hadoop, as its library is needed to process these large data sets. In this research thesis, this study proposes a smart framework model that profiles MapReduce tasks with the use of Machine Learning (ML) algorithms to effectively place the data in Hadoop clusters; activate only sufficient number of nodes to accomplish the data processing within the planned deadline time for the task. The model will ensure achieving energy efficiency by utilizing the minimum number of necessary nodes, with maximum utilization and least energy consumption to reduce the overall cost of operations in data centers that deploy the Hadoop clusters.</p>"],"dc:identifier":["https://csuepress.columbusstate.edu/theses_dissertations/392"],"dc:language":["English"],"dc:subject":["Hadoop","ML","Energy-Aware","Big Data","MapReduce","HDFS","Computer Sciences","Databases and Information Systems"],"dc:title":["Improving Energy-Efficiency through Smart Data Placement in Hadoop Clusters"],"thesis:degree_discipline":["TSYS School of Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Computer Science - Applied Computing Track"]},"updated_at":"2026-07-24T01:45:01Z"}