Back to results

York University

Exploring and Evaluating the Scalability and Efficiency of Apache Spark using Educational Datasets

Abstract

dc:description.abstract

Research into the combination of data mining and machine learning technology with web-based education systems (known as education data mining, or EDM) is becoming imperative in order to enhance the quality of education by moving beyond traditional methods. With the worldwide growth of the Information Communication Technology (ICT), data are becoming available at a significantly large volume, with high velocity and extensive variety. In this thesis, four popular data mining methods are applied to Apache Spark, using large volumes of datasets from Online Cognitive Learning Systems to explore the scalability and efficiency of Spark. Various volumes of datasets are tested on Spark MLlib with different running configurations and parameter tunings. The thesis convincingly presents useful strategies for allocating computing resources and tuning to take full advantage of the in-memory system of Apache Spark to conduct the tasks of data mining and machine learning. Moreover, it offers insights that education experts and data scientists can use to manage and improve the quality of education, as well as to analyze and discover hidden knowledge in the era of big data.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zhang, Jian
Advisor dc:contributor.advisor
  • Yang, Zijiang Cynthia

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests.
Language dc:language.iso
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/10315/35023
OAI identifier oai:identifier
oai:yorkspace.library.yorku.ca:10315/35023

Chain of custody

source
Harvested from
York University
Base URL
yorkspace.library.yorku.ca/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Zhang, Jian. Exploring and Evaluating the Scalability and Efficiency of Apache Spark using Educational Datasets. 2018. http://hdl.handle.net/10315/35023