{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/98336"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/98336","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Concept and entity grounding using indirect supervision","abstract":"Extracting and disambiguating entities and concepts is a crucial step toward understanding natural language text. In this thesis, we consider the problem of grounding concepts and entities mentioned in text to one or more knowledge bases (KBs). A well-studied scenario of this problem is the one in which documents are given in English and the goal is to identify concept and entity mentions, and find the corresponding entries the mentions refer to in Wikipedia. We extend this problem in two directions: First, we study identifying and grounding entities written in any language to the English Wikipedia. Second, we investigate using multiple KBs which do not contain rich textual and structural information Wikipedia does. These more involved settings pose a few additional challenges beyond those addressed in the standard English Wikification problem. Key among them is that no supervision is available to facilitate training machine learning models. The first extension, cross-lingual Wikification, introduces problems such as recognizing multilingual named entities mentioned in text, translating non-English names into English, and computing word similarity across languages. Since it is impossible to acquire manually annotated examples for all languages, building models for all languages in Wikipedia requires exploring indirect or incidental supervision signals which already exist in Wikipedia. For the second setting, we need to deal with the fact that most KBs do not contain the rich information Wikipedia has; consequently, the main supervision signal used to train Wikification rankers does not exist anymore. In this thesis, we show that supervision signals can be obtained by carefully examining the redundancy and relations between multiple KBs. By developing algorithms and models which harvest these incidental signals, we can achieve better performance on these tasks.","abstract_html":"Extracting and disambiguating entities and concepts is a crucial step toward understanding natural language text. In this thesis, we consider the problem of grounding concepts and entities mentioned in text to one or more knowledge bases (KBs). A well-studied scenario of this problem is the one in which documents are given in English and the goal is to identify concept and entity mentions, and find the corresponding entries the mentions refer to in Wikipedia. We extend this problem in two directions: First, we study identifying and grounding entities written in any language to the English Wikipedia. Second, we investigate using multiple KBs which do not contain rich textual and structural information Wikipedia does. These more involved settings pose a few additional challenges beyond those addressed in the standard English Wikification problem. Key among them is that no supervision is available to facilitate training machine learning models. The first extension, cross-lingual Wikification, introduces problems such as recognizing multilingual named entities mentioned in text, translating non-English names into English, and computing word similarity across languages. Since it is impossible to acquire manually annotated examples for all languages, building models for all languages in Wikipedia requires exploring indirect or incidental supervision signals which already exist in Wikipedia. For the second setting, we need to deal with the fact that most KBs do not contain the rich information Wikipedia has; consequently, the main supervision signal used to train Wikification rankers does not exist anymore. In this thesis, we show that supervision signals can be obtained by carefully examining the redundancy and relations between multiple KBs. By developing algorithms and models which harvest these incidental signals, we can achieve better performance on these tasks.","abstract_has_math":false,"creators":["Tsai, Chen-Tse"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Roth, Dan","Chang, Kevin","Zhai, ChengXiang","Mihalcea, Rada"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-09-29T17:56:26Z","date_published":"2017-09-29T17:56:26Z","updated_at":"2026-07-22T22:24:35Z","subjects":["Wikification","Entity linking","Cross-lingual wikification","Named entity recognition","Indirect supervision","Incidental supervision","Entity disambiguation","Concept disambiguation"],"languages":["en"],"rights":["Copyright 2017 Chen-Tse Tsai"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/98336","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Roth, Dan","Chang, Kevin","Zhai, ChengXiang","Mihalcea, Rada"]},{"key":"dc:creator","label":"Author","values":["Tsai, Chen-Tse"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-09-29T17:56:26Z","2017-07-06","2017-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Wikification","Entity linking","Cross-lingual wikification","Named entity recognition","Indirect supervision","Incidental supervision","Entity disambiguation","Concept disambiguation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Chen-Tse Tsai"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/98336"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Extracting and disambiguating entities and concepts is a crucial step toward understanding natural language text. In this thesis, we consider the problem of grounding concepts and entities mentioned in text to one or more knowledge bases (KBs). A well-studied scenario of this problem is the one in which documents are given in English and the goal is to identify concept and entity mentions, and find the corresponding entries the mentions refer to in Wikipedia. We extend this problem in two directions: First, we study identifying and grounding entities written in any language to the English Wikipedia. Second, we investigate using multiple KBs which do not contain rich textual and structural information Wikipedia does. These more involved settings pose a few additional challenges beyond those addressed in the standard English Wikification problem. Key among them is that no supervision is available to facilitate training machine learning models. The first extension, cross-lingual Wikification, introduces problems such as recognizing multilingual named entities mentioned in text, translating non-English names into English, and computing word similarity across languages. Since it is impossible to acquire manually annotated examples for all languages, building models for all languages in Wikipedia requires exploring indirect or incidental supervision signals which already exist in Wikipedia. For the second setting, we need to deal with the fact that most KBs do not contain the rich information Wikipedia has; consequently, the main supervision signal used to train Wikification rankers does not exist anymore. In this thesis, we show that supervision signals can be obtained by carefully examining the redundancy and relations between multiple KBs. By developing algorithms and models which harvest these incidental signals, we can achieve better performance on these tasks.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-09-29 without embargo terms","The student, Chen-Tse Tsai, accepted the attached license on 2017-07-06 at 00:19.","The student, Chen-Tse Tsai, submitted this Dissertation for approval on 2017-07-06 at 00:33.","This Dissertation was approved for publication on 2017-07-06 at 16:13.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11320 on 2017-09-29 at 11:28:01","Made available in DSpace on 2017-09-29T17:56:26Z (GMT). No. of bitstreams: 3 TSAI-DISSERTATION-2017.pdf: 3356413 bytes, checksum: 5abccfe79304a8ac502b2e95c7d0f3c8 (MD5) LICENSE.txt: 4210 bytes, checksum: 76014038fa4685d0998dbb26a3551db7 (MD5) PROQUEST_LICENSE.txt: 4556 bytes, checksum: 689b6a2414256a55b00a32a7ee274d4a (MD5) Previous issue date: 2017-07-06"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Concept and entity grounding using indirect supervision"]}]}],"canonical_facts":{"dc:contributor":["Roth, Dan","Chang, Kevin","Zhai, ChengXiang","Mihalcea, Rada"],"dc:creator":["Tsai, Chen-Tse"],"dc:date":["2017-09-29T17:56:26Z","2017-07-06","2017-08"],"dc:description":["Extracting and disambiguating entities and concepts is a crucial step toward understanding natural language text. In this thesis, we consider the problem of grounding concepts and entities mentioned in text to one or more knowledge bases (KBs). A well-studied scenario of this problem is the one in which documents are given in English and the goal is to identify concept and entity mentions, and find the corresponding entries the mentions refer to in Wikipedia. We extend this problem in two directions: First, we study identifying and grounding entities written in any language to the English Wikipedia. Second, we investigate using multiple KBs which do not contain rich textual and structural information Wikipedia does. These more involved settings pose a few additional challenges beyond those addressed in the standard English Wikification problem. Key among them is that no supervision is available to facilitate training machine learning models. The first extension, cross-lingual Wikification, introduces problems such as recognizing multilingual named entities mentioned in text, translating non-English names into English, and computing word similarity across languages. Since it is impossible to acquire manually annotated examples for all languages, building models for all languages in Wikipedia requires exploring indirect or incidental supervision signals which already exist in Wikipedia. For the second setting, we need to deal with the fact that most KBs do not contain the rich information Wikipedia has; consequently, the main supervision signal used to train Wikification rankers does not exist anymore. In this thesis, we show that supervision signals can be obtained by carefully examining the redundancy and relations between multiple KBs. By developing algorithms and models which harvest these incidental signals, we can achieve better performance on these tasks.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-09-29 without embargo terms","The student, Chen-Tse Tsai, accepted the attached license on 2017-07-06 at 00:19.","The student, Chen-Tse Tsai, submitted this Dissertation for approval on 2017-07-06 at 00:33.","This Dissertation was approved for publication on 2017-07-06 at 16:13.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11320 on 2017-09-29 at 11:28:01","Made available in DSpace on 2017-09-29T17:56:26Z (GMT). No. of bitstreams: 3 TSAI-DISSERTATION-2017.pdf: 3356413 bytes, checksum: 5abccfe79304a8ac502b2e95c7d0f3c8 (MD5) LICENSE.txt: 4210 bytes, checksum: 76014038fa4685d0998dbb26a3551db7 (MD5) PROQUEST_LICENSE.txt: 4556 bytes, checksum: 689b6a2414256a55b00a32a7ee274d4a (MD5) Previous issue date: 2017-07-06"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/98336"],"dc:language":["en"],"dc:rights":["Copyright 2017 Chen-Tse Tsai"],"dc:subject":["Wikification","Entity linking","Cross-lingual wikification","Named entity recognition","Indirect supervision","Incidental supervision","Entity disambiguation","Concept disambiguation"],"dc:title":["Concept and entity grounding using indirect supervision"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:35Z"}