Back to search

NJIT

Data analytics with mapreduce in apache spark and hadoop systems

Abstract

dc:description.abstract

MapReduce comes from a traditional problem solving method: separating a big problem and solving each small parts. With the target of computing larger dataset in more efficient and cheaper way, this is implement into a programming mode to deal with massive quantity of data. The users get a map function and use it to abstract dataset into key / value logical pair and then use a reduce function to group all value with the same key. With this mode, task can be automatic spread the job into clusters grouped by lots of normal computers. MapReduce program can be easily implemented and gain much more efficiency than tradition computing programs. In this paper there are some sample programs and one GRN detection algorithm program to study about it. Detecting gene regulatory networks (GRN), the regulatory molecules connection among various genes, is one of the main subjects in understanding gene biology. Although there are algorithms developed for this target, the increase of gene size and their complexity make the processing time more and more hard and slow. MapReduce mode with parallelize computing can be one way to overcome these problems. In this paper, a well-defined framework to parallelize mutual information algorithm is presented. The experiments and result performances shows the improvement of using parallelizing MapReduce model.

Degree

thesis:*
Name thesis:degree_name
Master of Science in Computer Science - (M.S.)
Discipline thesis:degree_discipline
Computer Science
Year
2016

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Du, Zongxuan
Contributors dc:contributor
  • Jason T. L. Wang
  • Xiaoning Ding
  • Chase Qishi Wu

Subjects

dc:subject × 3

Identifiers

dc:identifier.*
Repository record dc:identifier
https://digitalcommons.njit.edu/theses/269
OAI identifier oai:identifier
oai:digitalcommons.njit.edu:theses-1268

Chain of custody

source
Harvested from
NJIT
Base URL
digitalcommons.njit.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Du, Zongxuan. Data analytics with mapreduce in apache spark and hadoop systems. 2016. https://digitalcommons.njit.edu/theses/269