{"id":{"repo_id":"milano","oai_identifier":"oai:air.unimi.it:2434/1202718"},"canonical_url":"https://search.dev.ndltd.org/etd/milano/oai:air.unimi.it:2434/1202718","repository":{"repo_id":"milano","name":"Università degli Studi di Milano","base_url":"https://air.unimi.it/oai/request"},"display":{"title":"ADAPTIVE FRAMEWORKS FOR KNOWLEDGE EXTRACTION IN HETEROGENEOUS DATA ENVIRONMENTS","abstract":"The proliferation of unstructured, multimodal data presents a significant challenge for effective knowledge extraction, due to the heterogeneous nature and the complexity of extracting meaningful patterns in environments presenting diverse data types. This thesis proposes SHIFT, the first seed-guided hierarchical topic modelling framework specifically designed for heterogeneous data environments. It combines unsupervised information extraction with advanced representation learning techniques, incorporating external knowledge bases to enhance semantic understanding. SHIFT modular architecture provides seamless adaptation to different data modalities and domain requirements. The framework’s adaptability is demonstrated through comprehensive applications across distinct domains, from legal text analysis, through scientific literature understanding, to digital humanities research. Beyond the core SHIFT framework, this thesis presents the development and application of complementary approaches tailored to domain-specific challenges. Through extensive evaluation, the research validates the technical effectiveness and practical utility of these frameworks for real-world knowledge extraction challenges. This work contributes to advances in multimodal topic modelling while demonstrating the critical importance of adaptive, modular approaches for handling the complexity and diversity of contemporary unstructured data across multiple academic and professional contexts.","abstract_html":"The proliferation of unstructured, multimodal data presents a significant challenge for effective knowledge extraction, due to the heterogeneous nature and the complexity of extracting meaningful patterns in environments presenting diverse data types. This thesis proposes SHIFT, the first seed-guided hierarchical topic modelling framework specifically designed for heterogeneous data environments. It combines unsupervised information extraction with advanced representation learning techniques, incorporating external knowledge bases to enhance semantic understanding. SHIFT modular architecture provides seamless adaptation to different data modalities and domain requirements. The framework’s adaptability is demonstrated through comprehensive applications across distinct domains, from legal text analysis, through scientific literature understanding, to digital humanities research. Beyond the core SHIFT framework, this thesis presents the development and application of complementary approaches tailored to domain-specific challenges. Through extensive evaluation, the research validates the technical effectiveness and practical utility of these frameworks for real-world knowledge extraction challenges. This work contributes to advances in multimodal topic modelling while demonstrating the critical importance of adaptive, modular approaches for handling the complexity and diversity of contemporary unstructured data across multiple academic and professional contexts.","abstract_has_math":false,"creators":["PICASCIA, SERGIO"],"institution":"Università degli Studi di Milano","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["tutor: A. Ferrara ; co-tutor: S. Montanelli ; coordinatore: R. Sassi","S. Picascia","FERRARA, ALFIO","SASSI, ROBERTO"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12-19","date_published":"2025-12-19","updated_at":"2026-07-27T20:19:15Z","subjects":["Settore INFO-01/A - Informatica"],"languages":["eng"],"rights":["info:eu-repo/semantics/openAccess","license:Creative commons","license uri:http://creativecommons.org/licenses/by-sa/4.0/"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["http://dx.doi.org/10.13130/picascia-sergio_phd2025-12-19","10.13130/picascia-sergio_phd2025-12-19"],"render_values":[{"text":"http://dx.doi.org/10.13130/picascia-sergio_phd2025-12-19","href":"http://dx.doi.org/10.13130/picascia-sergio_phd2025-12-19","code":true},{"text":"10.13130/picascia-sergio_phd2025-12-19","href":"https://doi.org/10.13130/picascia-sergio_phd2025-12-19","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/2434/1202718","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["tutor: A. Ferrara ; co-tutor: S. Montanelli ; coordinatore: R. Sassi","S. Picascia","FERRARA, ALFIO","SASSI, ROBERTO"]},{"key":"dc:creator","label":"Author","values":["PICASCIA, SERGIO"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12-19"]},{"key":"dc:publisher","label":"Institution","values":["Università degli Studi di Milano"]},{"key":"dc:relation","label":"Dc Relation","values":["numberofpages:215"]},{"key":"dc:type","label":"Dc Type","values":["info:eu-repo/semantics/doctoralThesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Settore INFO-01/A - Informatica"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["info:eu-repo/semantics/openAccess","license:Creative commons","license uri:http://creativecommons.org/licenses/by-sa/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://dx.doi.org/10.13130/picascia-sergio_phd2025-12-19","https://hdl.handle.net/2434/1202718","10.13130/picascia-sergio_phd2025-12-19"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The proliferation of unstructured, multimodal data presents a significant challenge for effective knowledge extraction, due to the heterogeneous nature and the complexity of extracting meaningful patterns in environments presenting diverse data types. This thesis proposes SHIFT, the first seed-guided hierarchical topic modelling framework specifically designed for heterogeneous data environments. It combines unsupervised information extraction with advanced representation learning techniques, incorporating external knowledge bases to enhance semantic understanding. SHIFT modular architecture provides seamless adaptation to different data modalities and domain requirements. The framework’s adaptability is demonstrated through comprehensive applications across distinct domains, from legal text analysis, through scientific literature understanding, to digital humanities research. Beyond the core SHIFT framework, this thesis presents the development and application of complementary approaches tailored to domain-specific challenges. Through extensive evaluation, the research validates the technical effectiveness and practical utility of these frameworks for real-world knowledge extraction challenges. This work contributes to advances in multimodal topic modelling while demonstrating the critical importance of adaptive, modular approaches for handling the complexity and diversity of contemporary unstructured data across multiple academic and professional contexts."]},{"key":"dc:title","label":"Title","values":["ADAPTIVE FRAMEWORKS FOR KNOWLEDGE EXTRACTION IN HETEROGENEOUS DATA ENVIRONMENTS"]}]}],"canonical_facts":{"dc:contributor":["tutor: A. Ferrara ; co-tutor: S. Montanelli ; coordinatore: R. Sassi","S. Picascia","FERRARA, ALFIO","SASSI, ROBERTO"],"dc:creator":["PICASCIA, SERGIO"],"dc:date":["2025-12-19"],"dc:description":["The proliferation of unstructured, multimodal data presents a significant challenge for effective knowledge extraction, due to the heterogeneous nature and the complexity of extracting meaningful patterns in environments presenting diverse data types. This thesis proposes SHIFT, the first seed-guided hierarchical topic modelling framework specifically designed for heterogeneous data environments. It combines unsupervised information extraction with advanced representation learning techniques, incorporating external knowledge bases to enhance semantic understanding. SHIFT modular architecture provides seamless adaptation to different data modalities and domain requirements. The framework’s adaptability is demonstrated through comprehensive applications across distinct domains, from legal text analysis, through scientific literature understanding, to digital humanities research. Beyond the core SHIFT framework, this thesis presents the development and application of complementary approaches tailored to domain-specific challenges. Through extensive evaluation, the research validates the technical effectiveness and practical utility of these frameworks for real-world knowledge extraction challenges. This work contributes to advances in multimodal topic modelling while demonstrating the critical importance of adaptive, modular approaches for handling the complexity and diversity of contemporary unstructured data across multiple academic and professional contexts."],"dc:identifier":["http://dx.doi.org/10.13130/picascia-sergio_phd2025-12-19","https://hdl.handle.net/2434/1202718","10.13130/picascia-sergio_phd2025-12-19"],"dc:language":["eng"],"dc:publisher":["Università degli Studi di Milano"],"dc:relation":["numberofpages:215"],"dc:rights":["info:eu-repo/semantics/openAccess","license:Creative commons","license uri:http://creativecommons.org/licenses/by-sa/4.0/"],"dc:subject":["Settore INFO-01/A - Informatica"],"dc:title":["ADAPTIVE FRAMEWORKS FOR KNOWLEDGE EXTRACTION IN HETEROGENEOUS DATA ENVIRONMENTS"],"dc:type":["info:eu-repo/semantics/doctoralThesis"]},"updated_at":"2026-07-27T20:19:15Z"}