Back to results

University of Victoria (Canada)

A distributed approach to Frequent Itemset Mining at low support levels

Abstract

dc:description.abstract

Frequent Itemset Mining, the process of finding frequently co-occurring sets of items in a dataset, has been at the core of the field of data mining for the past 25 years. During this time the datasets have grown much faster than the algorithms capacity to process them. Great progress was made at optimizing this task on a single computer however, despite years of research, very little progress has been made on parallelizing this task. FPGrowth based algorithms have proven notoriously difficult to parallelize and Apriori has largely fallen out of favor with the research community. In this thesis we introduce a parallel, Apriori based, Frequent Itemset Mining algo- rithm capable of distributing computation across large commodity clusters. Our case study demonstrates that our algorithm can efficiently scale to hundreds of cores, on a standard Hadoop MapReduce cluster, and can improve executions times by at least an order of magnitude at the lowest support levels.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Clark, Neal
Advisor dc:contributor.supervisor
  • Coady, Yvonne

Subjects

dc:subject × 7

Rights

Language dc:language.iso
en, English

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/1828/5803
OAI identifier oai:identifier
oai:dspace.library.uvic.ca:1828/5803

Chain of custody

source
Harvested from
University of Victoria (Canada)
Base URL
dspace.library.uvic.ca/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Clark, Neal. A distributed approach to Frequent Itemset Mining at low support levels. 2014. http://hdl.handle.net/1828/5803