University of Illinois at Urbana-Champaign
A faster reinforcement learning approach to efficient job scheduling in Apache Spark
Abstract
dc:descriptionJob scheduling problems have been widely studied in theoretical computer science and operations research, and are commonly encountered in applied settings such as computer systems, manufacturing, and construction. There are many variants of job scheduling, but they share a common goal: designat- ing jobs to run on a set of parallel machines at different times, such that the machines are efficiently utilized. This thesis focuses on job scheduling in the context of Apache SparkTM, a popular data analytics engine that harnesses the power of distributed computing. Job scheduling is central to Spark, as each Spark application needs a scheduler to orchestrate its job submissions. The basic scheduling rules provided by Spark work well on lighter workloads, but sophisticated scheduling algorithms can greatly increase cluster efficiency when workloads are heavier. Previous work has introduced such algorithms, some hand-tuned and others learned. This thesis thoroughly documents Dec- ima, the state-of-the-art, reinforcement-learned Spark job scheduler, includ- ing a close look into their simulator and model architectures, and a new SMDP formulation of the problem. This thesis also proposes Decima++, an update to Decima which improves scheduling performance and reduces training time by over 11×.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Industrial Engineering
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Gertsman, Arkadiy
- Contributors dc:contributor
-
- Nagi, Rakesh
Subjects
dc:subject × 4Rights
dc:rights- Statement dc:rights
-
- Copyright 2023 Arkadiy Gertsman
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/121563