{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132501"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132501","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Advancing semantic modeling: addressing coordination, interpretability, and data scarcity in domain representation","abstract":"The rapid growth of online text—from scientific articles and product catalogs to news and social media—offers an unprecedented opportunity to uncover structure and meaning from unorganized information. Traditional topic models provide a powerful framework for discovering latent themes but often fall short in practice: their outputs are generic and difficult to interpret, alignment across corpora is not guaranteed, and they struggle with both short documents and domains with limited data. These challenges limit their usefulness for applications such as search, recommendation, trend analysis, and domain comparison, where interpretability and adaptability are essential. This thesis develops four complementary solutions that address these limitations. Coordinated Topic Modeling (CTM) introduces a framework for aligning corpus-specific topics with interpretable reference topics, enabling both interpretability and cross-corpus comparability. Domain Representative Keyword Selection (DRKS) proposes a probabilistic method for extracting distinctive, context-aware keywords that capture the semantics of a target domain. Short-Text Topic Modeling (STTM) leverages large language models and prefix-tuned autoencoders to enrich sparse inputs and produce coherent topics under extreme document-level scarcity. Low-Resource Topic Modeling (LRTM) presents DALTA, a domain-adaptation framework that transfers knowledge from data-rich corpora while preserving target-specific nuance. Across multiple datasets and domains, these methods demonstrate substantial improvements over state-of-the-art baselines in coherence, stability, and adaptability. Together, they establish a principled foundation for building structured, interpretable, and adaptable semantic models in diverse and resource-constrained settings.","abstract_html":"The rapid growth of online text—from scientific articles and product catalogs to news and social media—offers an unprecedented opportunity to uncover structure and meaning from unorganized information. Traditional topic models provide a powerful framework for discovering latent themes but often fall short in practice: their outputs are generic and difficult to interpret, alignment across corpora is not guaranteed, and they struggle with both short documents and domains with limited data. These challenges limit their usefulness for applications such as search, recommendation, trend analysis, and domain comparison, where interpretability and adaptability are essential. This thesis develops four complementary solutions that address these limitations. Coordinated Topic Modeling (CTM) introduces a framework for aligning corpus-specific topics with interpretable reference topics, enabling both interpretability and cross-corpus comparability. Domain Representative Keyword Selection (DRKS) proposes a probabilistic method for extracting distinctive, context-aware keywords that capture the semantics of a target domain. Short-Text Topic Modeling (STTM) leverages large language models and prefix-tuned autoencoders to enrich sparse inputs and produce coherent topics under extreme document-level scarcity. Low-Resource Topic Modeling (LRTM) presents DALTA, a domain-adaptation framework that transfers knowledge from data-rich corpora while preserving target-specific nuance. Across multiple datasets and domains, these methods demonstrate substantial improvements over state-of-the-art baselines in coherence, stability, and adaptability. Together, they establish a principled foundation for building structured, interpretable, and adaptable semantic models in diverse and resource-constrained settings.","abstract_has_math":false,"creators":["Akash, Pritom Saha"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Chang, Kevin Chen-Chuan","Zhai, ChengXiang","He, Jingrui","Popa, Lucian"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Semantic Modeling","Topic Modeling","Domain Representation","Keyword Selection","Domain Adaptation","Variational Auto Encoder"],"languages":["en"],"rights":["Copyright 2025 Pritom Saha Akash"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132501","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chang, Kevin Chen-Chuan","Zhai, ChengXiang","He, Jingrui","Popa, Lucian"]},{"key":"dc:creator","label":"Author","values":["Akash, Pritom Saha"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-11-19"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Semantic Modeling","Topic Modeling","Domain Representation","Keyword Selection","Domain Adaptation","Variational Auto Encoder"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Pritom Saha Akash"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132501"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The rapid growth of online text—from scientific articles and product catalogs to news and social media—offers an unprecedented opportunity to uncover structure and meaning from unorganized information. Traditional topic models provide a powerful framework for discovering latent themes but often fall short in practice: their outputs are generic and difficult to interpret, alignment across corpora is not guaranteed, and they struggle with both short documents and domains with limited data. These challenges limit their usefulness for applications such as search, recommendation, trend analysis, and domain comparison, where interpretability and adaptability are essential. This thesis develops four complementary solutions that address these limitations. Coordinated Topic Modeling (CTM) introduces a framework for aligning corpus-specific topics with interpretable reference topics, enabling both interpretability and cross-corpus comparability. Domain Representative Keyword Selection (DRKS) proposes a probabilistic method for extracting distinctive, context-aware keywords that capture the semantics of a target domain. Short-Text Topic Modeling (STTM) leverages large language models and prefix-tuned autoencoders to enrich sparse inputs and produce coherent topics under extreme document-level scarcity. Low-Resource Topic Modeling (LRTM) presents DALTA, a domain-adaptation framework that transfers knowledge from data-rich corpora while preserving target-specific nuance. Across multiple datasets and domains, these methods demonstrate substantial improvements over state-of-the-art baselines in coherence, stability, and adaptability. Together, they establish a principled foundation for building structured, interpretable, and adaptable semantic models in diverse and resource-constrained settings.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Pritom Saha Akash, accepted the attached license on 2025-11-18 at 14:27.","The student, Pritom Saha Akash, submitted this Dissertation for approval on 2025-11-18 at 17:23.","This Dissertation was approved for publication on 2025-11-19 at 08:41.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22883 on 2026-02-19 at 18:24:55"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Advancing semantic modeling: addressing coordination, interpretability, and data scarcity in domain representation"]}]}],"canonical_facts":{"dc:contributor":["Chang, Kevin Chen-Chuan","Zhai, ChengXiang","He, Jingrui","Popa, Lucian"],"dc:creator":["Akash, Pritom Saha"],"dc:date":["2025-12","2025-11-19"],"dc:description":["The rapid growth of online text—from scientific articles and product catalogs to news and social media—offers an unprecedented opportunity to uncover structure and meaning from unorganized information. Traditional topic models provide a powerful framework for discovering latent themes but often fall short in practice: their outputs are generic and difficult to interpret, alignment across corpora is not guaranteed, and they struggle with both short documents and domains with limited data. These challenges limit their usefulness for applications such as search, recommendation, trend analysis, and domain comparison, where interpretability and adaptability are essential. This thesis develops four complementary solutions that address these limitations. Coordinated Topic Modeling (CTM) introduces a framework for aligning corpus-specific topics with interpretable reference topics, enabling both interpretability and cross-corpus comparability. Domain Representative Keyword Selection (DRKS) proposes a probabilistic method for extracting distinctive, context-aware keywords that capture the semantics of a target domain. Short-Text Topic Modeling (STTM) leverages large language models and prefix-tuned autoencoders to enrich sparse inputs and produce coherent topics under extreme document-level scarcity. Low-Resource Topic Modeling (LRTM) presents DALTA, a domain-adaptation framework that transfers knowledge from data-rich corpora while preserving target-specific nuance. Across multiple datasets and domains, these methods demonstrate substantial improvements over state-of-the-art baselines in coherence, stability, and adaptability. Together, they establish a principled foundation for building structured, interpretable, and adaptable semantic models in diverse and resource-constrained settings.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Pritom Saha Akash, accepted the attached license on 2025-11-18 at 14:27.","The student, Pritom Saha Akash, submitted this Dissertation for approval on 2025-11-18 at 17:23.","This Dissertation was approved for publication on 2025-11-19 at 08:41.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22883 on 2026-02-19 at 18:24:55"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132501"],"dc:language":["en"],"dc:rights":["Copyright 2025 Pritom Saha Akash"],"dc:subject":["Semantic Modeling","Topic Modeling","Domain Representation","Keyword Selection","Domain Adaptation","Variational Auto Encoder"],"dc:title":["Advancing semantic modeling: addressing coordination, interpretability, and data scarcity in domain representation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}