{"id":{"repo_id":"iastate","oai_identifier":"oai:dr.lib.iastate.edu:20.500.12876/Qr9mg7Jr"},"canonical_url":"https://search.dev.ndltd.org/etd/iastate/oai:dr.lib.iastate.edu:20.500.12876/Qr9mg7Jr","repository":{"repo_id":"iastate","name":"Iowa State University","base_url":"https://dr.lib.iastate.edu/server/oai/request"},"display":{"title":"Information extraction with weak supervision","abstract":"This dissertation explores the development and application of weak supervision techniques to address key challenges in three fundamental information extraction (IE) tasks: Named Entity Recognition (NER), Relation Extraction (RE), and Entity Linking (EL). Traditional supervised learning methods in these domains often require extensive human annotations, which are costly and time-consuming, limiting their scalability and applicability in real-world scenarios. To overcome these limitations, this research introduces innovative weakly supervised methodologies for each of these tasks, aiming to reduce reliance on manual labeling while maintaining high performance. The first part of the dissertation presents a novel framework, Confidence-Based Multi-Class Positive and Unlabeled (Conf-MPU) learning, designed to enhance the performance of distantly supervised NER. By incorporating confidence scores into a multi-class PU learning approach, Conf-MPU effectively handles incomplete labeling and varying false negative rates inherent in distantly supervised data. Experimental results on benchmark datasets demonstrate that Conf-MPU significantly outperforms existing state-of-the-art methods, advancing the field of distantly supervised NER. The second part focuses on improving Relation Extraction through the integration of indirect supervision. A novel approach, DSRE-NLI, is introduced, which leverages a Natural Language Inference (NLI) engine and a Semi-Automatic Relation Verbalization (SARV) mechanism to diagnose and mitigate label noise in distantly supervised RE tasks. This method enhances the semantic diversity of relation templates with minimal human input, resulting in a significant performance boost over traditional distantly supervised methods on real and simulated datasets. The third part of the dissertation addresses challenges in Zero-Shot Entity Linking (ZSEL) with a new re-ranking approach, GenDecider, which incorporates “None of the Candidates” (NoC) judgments into the re-ranking process. By formulating the task as a generative process using the Llama model, GenDecider effectively detects scenarios where the correct entity is not among the retrieved candidates. This approach significantly improves the accuracy and reliability of ZSEL systems, as evidenced by its performance on the benchmark ZESHEL dataset. Collectively, the contributions of this dissertation lie in advancing weak supervision techniques across three critical IE tasks, reducing the dependency on extensive manual annotations, and improving the robustness and scalability of information extraction systems. The findings have broad implications for the development of practical, scalable IE solutions in data-rich environments. Future research directions include refining noise-handling mechanisms, optimizing computational efficiency, and expanding the proposed methods to multilingual and low-resource settings.","abstract_html":"This dissertation explores the development and application of weak supervision techniques to address key challenges in three fundamental information extraction (IE) tasks: Named Entity Recognition (NER), Relation Extraction (RE), and Entity Linking (EL). Traditional supervised learning methods in these domains often require extensive human annotations, which are costly and time-consuming, limiting their scalability and applicability in real-world scenarios. To overcome these limitations, this research introduces innovative weakly supervised methodologies for each of these tasks, aiming to reduce reliance on manual labeling while maintaining high performance. The first part of the dissertation presents a novel framework, Confidence-Based Multi-Class Positive and Unlabeled (Conf-MPU) learning, designed to enhance the performance of distantly supervised NER. By incorporating confidence scores into a multi-class PU learning approach, Conf-MPU effectively handles incomplete labeling and varying false negative rates inherent in distantly supervised data. Experimental results on benchmark datasets demonstrate that Conf-MPU significantly outperforms existing state-of-the-art methods, advancing the field of distantly supervised NER. The second part focuses on improving Relation Extraction through the integration of indirect supervision. A novel approach, DSRE-NLI, is introduced, which leverages a Natural Language Inference (NLI) engine and a Semi-Automatic Relation Verbalization (SARV) mechanism to diagnose and mitigate label noise in distantly supervised RE tasks. This method enhances the semantic diversity of relation templates with minimal human input, resulting in a significant performance boost over traditional distantly supervised methods on real and simulated datasets. The third part of the dissertation addresses challenges in Zero-Shot Entity Linking (ZSEL) with a new re-ranking approach, GenDecider, which incorporates “None of the Candidates” (NoC) judgments into the re-ranking process. By formulating the task as a generative process using the Llama model, GenDecider effectively detects scenarios where the correct entity is not among the retrieved candidates. This approach significantly improves the accuracy and reliability of ZSEL systems, as evidenced by its performance on the benchmark ZESHEL dataset. Collectively, the contributions of this dissertation lie in advancing weak supervision techniques across three critical IE tasks, reducing the dependency on extensive manual annotations, and improving the robustness and scalability of information extraction systems. The findings have broad implications for the development of practical, scalable IE solutions in data-rich environments. Future research directions include refining noise-handling mechanisms, optimizing computational efficiency, and expanding the proposed methods to multilingual and low-resource settings.","abstract_has_math":false,"creators":["Zhou, Kang"],"institution":"Iowa State University","degree_name":"Doctor of Philosophy","degree_level":"dissertation","degree_discipline":"Computer science","degree_department":"Department of Computer Science","school":null,"contributors":[],"advisors":["Li, Qi","Cai, Ying","Liu, Kevin","Gao, Hongyang","Huai, Mengdi"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-12","date_published":"2024-12","updated_at":"2026-07-24T02:39:53Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.31274/td-20250502-2"],"render_values":[{"text":"https://doi.org/10.31274/td-20250502-2","href":"https://doi.org/10.31274/td-20250502-2","code":true}]}]},"links":{"outbound_url":"https://dr.lib.iastate.edu/handle/20.500.12876/Qr9mg7Jr","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Li, Qi","Cai, Ying","Liu, Kevin","Gao, Hongyang","Huai, Mengdi"]},{"key":"dc:contributor.department","label":"Department","values":["Department of Computer Science","Computer Science"]},{"key":"dc:creator","label":"Author","values":["Zhou, Kang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-02-11T17:23:01Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-02-11T17:23:01Z"]},{"key":"dc:date.issued","label":"Date","values":["2024-12"]},{"key":"dc:type","label":"Dc Type","values":["dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Iowa State University"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.31274/td-20250502-2"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://dr.lib.iastate.edu/handle/20.500.12876/Qr9mg7Jr"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This dissertation explores the development and application of weak supervision techniques to address key challenges in three fundamental information extraction (IE) tasks: Named Entity Recognition (NER), Relation Extraction (RE), and Entity Linking (EL). Traditional supervised learning methods in these domains often require extensive human annotations, which are costly and time-consuming, limiting their scalability and applicability in real-world scenarios. To overcome these limitations, this research introduces innovative weakly supervised methodologies for each of these tasks, aiming to reduce reliance on manual labeling while maintaining high performance. The first part of the dissertation presents a novel framework, Confidence-Based Multi-Class Positive and Unlabeled (Conf-MPU) learning, designed to enhance the performance of distantly supervised NER. By incorporating confidence scores into a multi-class PU learning approach, Conf-MPU effectively handles incomplete labeling and varying false negative rates inherent in distantly supervised data. Experimental results on benchmark datasets demonstrate that Conf-MPU significantly outperforms existing state-of-the-art methods, advancing the field of distantly supervised NER. The second part focuses on improving Relation Extraction through the integration of indirect supervision. A novel approach, DSRE-NLI, is introduced, which leverages a Natural Language Inference (NLI) engine and a Semi-Automatic Relation Verbalization (SARV) mechanism to diagnose and mitigate label noise in distantly supervised RE tasks. This method enhances the semantic diversity of relation templates with minimal human input, resulting in a significant performance boost over traditional distantly supervised methods on real and simulated datasets. The third part of the dissertation addresses challenges in Zero-Shot Entity Linking (ZSEL) with a new re-ranking approach, GenDecider, which incorporates “None of the Candidates” (NoC) judgments into the re-ranking process. By formulating the task as a generative process using the Llama model, GenDecider effectively detects scenarios where the correct entity is not among the retrieved candidates. This approach significantly improves the accuracy and reliability of ZSEL systems, as evidenced by its performance on the benchmark ZESHEL dataset. Collectively, the contributions of this dissertation lie in advancing weak supervision techniques across three critical IE tasks, reducing the dependency on extensive manual annotations, and improving the robustness and scalability of information extraction systems. The findings have broad implications for the development of practical, scalable IE solutions in data-rich environments. Future research directions include refining noise-handling mechanisms, optimizing computational efficiency, and expanding the proposed methods to multilingual and low-resource settings."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["PDF"]},{"key":"dc:title","label":"Title","values":["Information extraction with weak supervision"]}]}],"canonical_facts":{"dc:contributor.advisor":["Li, Qi","Cai, Ying","Liu, Kevin","Gao, Hongyang","Huai, Mengdi"],"dc:contributor.department":["Department of Computer Science","Computer Science"],"dc:creator":["Zhou, Kang"],"dc:date.accessioned":["2025-02-11T17:23:01Z"],"dc:date.available":["2025-02-11T17:23:01Z"],"dc:date.issued":["2024-12"],"dc:description.abstract":["This dissertation explores the development and application of weak supervision techniques to address key challenges in three fundamental information extraction (IE) tasks: Named Entity Recognition (NER), Relation Extraction (RE), and Entity Linking (EL). Traditional supervised learning methods in these domains often require extensive human annotations, which are costly and time-consuming, limiting their scalability and applicability in real-world scenarios. To overcome these limitations, this research introduces innovative weakly supervised methodologies for each of these tasks, aiming to reduce reliance on manual labeling while maintaining high performance. The first part of the dissertation presents a novel framework, Confidence-Based Multi-Class Positive and Unlabeled (Conf-MPU) learning, designed to enhance the performance of distantly supervised NER. By incorporating confidence scores into a multi-class PU learning approach, Conf-MPU effectively handles incomplete labeling and varying false negative rates inherent in distantly supervised data. Experimental results on benchmark datasets demonstrate that Conf-MPU significantly outperforms existing state-of-the-art methods, advancing the field of distantly supervised NER. The second part focuses on improving Relation Extraction through the integration of indirect supervision. A novel approach, DSRE-NLI, is introduced, which leverages a Natural Language Inference (NLI) engine and a Semi-Automatic Relation Verbalization (SARV) mechanism to diagnose and mitigate label noise in distantly supervised RE tasks. This method enhances the semantic diversity of relation templates with minimal human input, resulting in a significant performance boost over traditional distantly supervised methods on real and simulated datasets. The third part of the dissertation addresses challenges in Zero-Shot Entity Linking (ZSEL) with a new re-ranking approach, GenDecider, which incorporates “None of the Candidates” (NoC) judgments into the re-ranking process. By formulating the task as a generative process using the Llama model, GenDecider effectively detects scenarios where the correct entity is not among the retrieved candidates. This approach significantly improves the accuracy and reliability of ZSEL systems, as evidenced by its performance on the benchmark ZESHEL dataset. Collectively, the contributions of this dissertation lie in advancing weak supervision techniques across three critical IE tasks, reducing the dependency on extensive manual annotations, and improving the robustness and scalability of information extraction systems. The findings have broad implications for the development of practical, scalable IE solutions in data-rich environments. Future research directions include refining noise-handling mechanisms, optimizing computational efficiency, and expanding the proposed methods to multilingual and low-resource settings."],"dc:format.mimetype":["PDF"],"dc:identifier.doi":["https://doi.org/10.31274/td-20250502-2"],"dc:identifier.uri":["https://dr.lib.iastate.edu/handle/20.500.12876/Qr9mg7Jr"],"dc:language.iso":["en"],"dc:title":["Information extraction with weak supervision"],"dc:type":["dissertation"],"thesis:degree_discipline":["Computer science"],"thesis:degree_level":["dissertation"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Iowa State University"]},"updated_at":"2026-07-24T02:39:53Z"}