{"id":{"repo_id":"poli-torino","oai_identifier":"oai:iris.polito.it:11583/2507907"},"canonical_url":"https://search.dev.ndltd.org/etd/poli-torino/oai:iris.polito.it:11583/2507907","repository":{"repo_id":"poli-torino","name":"Politecnico di Torino","base_url":"https://iris.polito.it/oai/request"},"display":{"title":"Scaling data mining activities on very large datasets","abstract":"This thesis addresses the issue of enhancing the scalability of data mining techniques, with specific emphasis on association rule and frequent itemset mining. In particular, it proposes a scalable itemset mining approach relying on (i) a persistent (disk-based) representation of the transactional data, (ii) ad-hoc data retrieval techniques, and (iii)~strategies for the integration of existing itemset mining algorithms. A parallel design based on the same approach, to perform itemset extraction in a parallel and/or distributed environment, is also described. To address the manageability of frequent itemsets, a concise disk-based representation, with a set of querying techniques, is proposed. This work has been preliminarly validated in the Semantic Web domain, to identify semantic relationships from textual collections with a semi-automatic approach. As a minor topic, the extracion of frequent itemsets from streams of data, modelled as a set of transactional data windows, has also been tackled by proposing an online/offline analysis approach. Finally, a software platform, developed in a joint effort with the Institute for Cancer Research and Treatment (Candiolo), is presented, allowing the collection, the integration, and the analysis of heterogeneous data from the molecular oncology field.","abstract_html":"This thesis addresses the issue of enhancing the scalability of data mining techniques, with specific emphasis on association rule and frequent itemset mining. In particular, it proposes a scalable itemset mining approach relying on (i) a persistent (disk-based) representation of the transactional data, (ii) ad-hoc data retrieval techniques, and (iii)~strategies for the integration of existing itemset mining algorithms. A parallel design based on the same approach, to perform itemset extraction in a parallel and/or distributed environment, is also described. To address the manageability of frequent itemsets, a concise disk-based representation, with a set of querying techniques, is proposed. This work has been preliminarly validated in the Semantic Web domain, to identify semantic relationships from textual collections with a semi-automatic approach. As a minor topic, the extracion of frequent itemsets from streams of data, modelled as a set of transactional data windows, has also been tackled by proposing an online/offline analysis approach. Finally, a software platform, developed in a joint effort with the Institute for Cancer Research and Treatment (Candiolo), is presented, allowing the collection, the integration, and the analysis of heterogeneous data from the molecular oncology field.","abstract_has_math":false,"creators":["GRAND, ALBERTO"],"institution":"Politecnico di Torino","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013","date_published":"2013","updated_at":"2026-07-24T03:51:05Z","subjects":["data mining","association rule","itemset","scalability","Settore ING-INF/05 - Sistemi di Elaborazione delle Informazioni"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/11583/2507907","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Grand, Alberto"]},{"key":"dc:creator","label":"Author","values":["GRAND, ALBERTO"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2013"]},{"key":"dc:publisher","label":"Institution","values":["Politecnico di Torino","country:Italy"]},{"key":"dc:relation","label":"Dc Relation","values":["numberofpages:146"]},{"key":"dc:type","label":"Dc Type","values":["info:eu-repo/semantics/doctoralThesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["data mining","association rule","itemset","scalability","Settore ING-INF/05 - Sistemi di Elaborazione delle Informazioni"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/11583/2507907"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This thesis addresses the issue of enhancing the scalability of data mining techniques, with specific emphasis on association rule and frequent itemset mining. In particular, it proposes a scalable itemset mining approach relying on (i) a persistent (disk-based) representation of the transactional data, (ii) ad-hoc data retrieval techniques, and (iii)~strategies for the integration of existing itemset mining algorithms. A parallel design based on the same approach, to perform itemset extraction in a parallel and/or distributed environment, is also described. To address the manageability of frequent itemsets, a concise disk-based representation, with a set of querying techniques, is proposed. This work has been preliminarly validated in the Semantic Web domain, to identify semantic relationships from textual collections with a semi-automatic approach. As a minor topic, the extracion of frequent itemsets from streams of data, modelled as a set of transactional data windows, has also been tackled by proposing an online/offline analysis approach. Finally, a software platform, developed in a joint effort with the Institute for Cancer Research and Treatment (Candiolo), is presented, allowing the collection, the integration, and the analysis of heterogeneous data from the molecular oncology field."]},{"key":"dc:format","label":"Dc Format","values":["STAMPA"]},{"key":"dc:title","label":"Title","values":["Scaling data mining activities on very large datasets"]}]}],"canonical_facts":{"dc:contributor":["Grand, Alberto"],"dc:creator":["GRAND, ALBERTO"],"dc:date":["2013"],"dc:description":["This thesis addresses the issue of enhancing the scalability of data mining techniques, with specific emphasis on association rule and frequent itemset mining. In particular, it proposes a scalable itemset mining approach relying on (i) a persistent (disk-based) representation of the transactional data, (ii) ad-hoc data retrieval techniques, and (iii)~strategies for the integration of existing itemset mining algorithms. A parallel design based on the same approach, to perform itemset extraction in a parallel and/or distributed environment, is also described. To address the manageability of frequent itemsets, a concise disk-based representation, with a set of querying techniques, is proposed. This work has been preliminarly validated in the Semantic Web domain, to identify semantic relationships from textual collections with a semi-automatic approach. As a minor topic, the extracion of frequent itemsets from streams of data, modelled as a set of transactional data windows, has also been tackled by proposing an online/offline analysis approach. Finally, a software platform, developed in a joint effort with the Institute for Cancer Research and Treatment (Candiolo), is presented, allowing the collection, the integration, and the analysis of heterogeneous data from the molecular oncology field."],"dc:format":["STAMPA"],"dc:identifier":["http://hdl.handle.net/11583/2507907"],"dc:language":["eng"],"dc:publisher":["Politecnico di Torino","country:Italy"],"dc:relation":["numberofpages:146"],"dc:subject":["data mining","association rule","itemset","scalability","Settore ING-INF/05 - Sistemi di Elaborazione delle Informazioni"],"dc:title":["Scaling data mining activities on very large datasets"],"dc:type":["info:eu-repo/semantics/doctoralThesis"]},"updated_at":"2026-07-24T03:51:05Z"}