Abstract
dc:descriptionThis thesis addresses the issue of enhancing the scalability of data mining techniques, with specific emphasis on association rule and frequent itemset mining. In particular, it proposes a scalable itemset mining approach relying on (i) a persistent (disk-based) representation of the transactional data, (ii) ad-hoc data retrieval techniques, and (iii)~strategies for the integration of existing itemset mining algorithms. A parallel design based on the same approach, to perform itemset extraction in a parallel and/or distributed environment, is also described. To address the manageability of frequent itemsets, a concise disk-based representation, with a set of querying techniques, is proposed. This work has been preliminarly validated in the Semantic Web domain, to identify semantic relationships from textual collections with a semi-automatic approach. As a minor topic, the extracion of frequent itemsets from streams of data, modelled as a set of transactional data windows, has also been tackled by proposing an online/offline analysis approach. Finally, a software platform, developed in a joint effort with the Institute for Cancer Research and Treatment (Candiolo), is presented, allowing the collection, the integration, and the analysis of heterogeneous data from the molecular oncology field.
Degree
thesis:*- Grantor dc:publisher
- Politecnico di Torino
- Year dc:date
- 2013
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- GRAND, ALBERTO
Subjects
dc:subject × 5Rights
- Language dc:language
- eng
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/11583/2507907
- OAI identifier oai:identifier
- oai:iris.polito.it:11583/2507907