{"id":{"repo_id":"york","oai_identifier":"oai:yorkspace.library.yorku.ca:10315/35023"},"canonical_url":"https://search.dev.ndltd.org/etd/york/oai:yorkspace.library.yorku.ca:10315/35023","repository":{"repo_id":"york","name":"York University","base_url":"https://yorkspace.library.yorku.ca/oai/request"},"display":{"title":"Exploring and Evaluating the Scalability and Efficiency of Apache Spark using Educational Datasets","abstract":"Research into the combination of data mining and machine learning technology with web-based education systems (known as education data mining, or EDM) is becoming imperative in order to enhance the quality of education by moving beyond traditional methods. With the worldwide growth of the Information Communication Technology (ICT), data are becoming available at a significantly large volume, with high velocity and extensive variety. In this thesis, four popular data mining methods are applied to Apache Spark, using large volumes of datasets from Online Cognitive Learning Systems to explore the scalability and efficiency of Spark. Various volumes of datasets are tested on Spark MLlib with different running configurations and parameter tunings. The thesis convincingly presents useful strategies for allocating computing resources and tuning to take full advantage of the in-memory system of Apache Spark to conduct the tasks of data mining and machine learning. Moreover, it offers insights that education experts and data scientists can use to manage and improve the quality of education, as well as to analyze and discover hidden knowledge in the era of big data.","abstract_html":"Research into the combination of data mining and machine learning technology with web-based education systems (known as education data mining, or EDM) is becoming imperative in order to enhance the quality of education by moving beyond traditional methods. With the worldwide growth of the Information Communication Technology (ICT), data are becoming available at a significantly large volume, with high velocity and extensive variety. In this thesis, four popular data mining methods are applied to Apache Spark, using large volumes of datasets from Online Cognitive Learning Systems to explore the scalability and efficiency of Spark. Various volumes of datasets are tested on Spark MLlib with different running configurations and parameter tunings. The thesis convincingly presents useful strategies for allocating computing resources and tuning to take full advantage of the in-memory system of Apache Spark to conduct the tasks of data mining and machine learning. Moreover, it offers insights that education experts and data scientists can use to manage and improve the quality of education, as well as to analyze and discover hidden knowledge in the era of big data.","abstract_has_math":false,"creators":["Zhang, Jian"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Yang, Zijiang Cynthia"],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-08-27","date_published":"2018-08-27","updated_at":"2026-07-24T06:34:05Z","subjects":["Educational technology"],"languages":["en"],"rights":["Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/10315/35023","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Yang, Zijiang Cynthia"]},{"key":"dc:creator","label":"Author","values":["Zhang, Jian"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2018-08-27T16:42:44Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2018-08-27T16:42:44Z"]},{"key":"dc:date.issued","label":"Date","values":["2018-08-27"]},{"key":"dc:type","label":"Dc Type","values":["Electronic Thesis or Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Educational technology"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/10315/35023"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Research into the combination of data mining and machine learning technology with web-based education systems (known as education data mining, or EDM) is becoming imperative in order to enhance the quality of education by moving beyond traditional methods. With the worldwide growth of the Information Communication Technology (ICT), data are becoming available at a significantly large volume, with high velocity and extensive variety. In this thesis, four popular data mining methods are applied to Apache Spark, using large volumes of datasets from Online Cognitive Learning Systems to explore the scalability and efficiency of Spark. Various volumes of datasets are tested on Spark MLlib with different running configurations and parameter tunings. The thesis convincingly presents useful strategies for allocating computing resources and tuning to take full advantage of the in-memory system of Apache Spark to conduct the tasks of data mining and machine learning. Moreover, it offers insights that education experts and data scientists can use to manage and improve the quality of education, as well as to analyze and discover hidden knowledge in the era of big data."]},{"key":"dc:title","label":"Title","values":["Exploring and Evaluating the Scalability and Efficiency of Apache Spark using Educational Datasets"]}]}],"canonical_facts":{"dc:contributor.advisor":["Yang, Zijiang Cynthia"],"dc:creator":["Zhang, Jian"],"dc:date.accessioned":["2018-08-27T16:42:44Z"],"dc:date.available":["2018-08-27T16:42:44Z"],"dc:date.issued":["2018-08-27"],"dc:description.abstract":["Research into the combination of data mining and machine learning technology with web-based education systems (known as education data mining, or EDM) is becoming imperative in order to enhance the quality of education by moving beyond traditional methods. With the worldwide growth of the Information Communication Technology (ICT), data are becoming available at a significantly large volume, with high velocity and extensive variety. In this thesis, four popular data mining methods are applied to Apache Spark, using large volumes of datasets from Online Cognitive Learning Systems to explore the scalability and efficiency of Spark. Various volumes of datasets are tested on Spark MLlib with different running configurations and parameter tunings. The thesis convincingly presents useful strategies for allocating computing resources and tuning to take full advantage of the in-memory system of Apache Spark to conduct the tasks of data mining and machine learning. Moreover, it offers insights that education experts and data scientists can use to manage and improve the quality of education, as well as to analyze and discover hidden knowledge in the era of big data."],"dc:identifier.uri":["http://hdl.handle.net/10315/35023"],"dc:language.iso":["en"],"dc:rights":["Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests."],"dc:subject":["Educational technology"],"dc:title":["Exploring and Evaluating the Scalability and Efficiency of Apache Spark using Educational Datasets"],"dc:type":["Electronic Thesis or Dissertation"]},"updated_at":"2026-07-24T06:34:05Z"}