{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/109431"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/109431","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Cross-lingual entity extraction and linking for 300 languages","abstract":"The student, Xiaoman Pan, accepted the attached license on 2020-12-02 at 17:38.","abstract_html":"The student, Xiaoman Pan, accepted the attached license on 2020-12-02 at 17:38.","abstract_has_math":false,"creators":["Pan, Xiaoman"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng","Han, Jiawei","Tong, Hanghang","Knight, Kevin"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-03-05T21:38:20Z","date_published":"2021-03-05T21:38:20Z","updated_at":"2026-07-22T22:24:50Z","subjects":["cross-lingual","entity extraction","entity linking"],"languages":["en"],"rights":["Copyright 2020 Xiaoman Pan"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/109431","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng","Han, Jiawei","Tong, Hanghang","Knight, Kevin"]},{"key":"dc:creator","label":"Author","values":["Pan, Xiaoman"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-03-05T21:38:20Z","2020-12-03","2020-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["cross-lingual","entity extraction","entity linking"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Xiaoman Pan"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/109431"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The student, Xiaoman Pan, accepted the attached license on 2020-12-02 at 17:38.","The student, Xiaoman Pan, submitted this Dissertation for approval on 2020-12-02 at 17:43.","This Dissertation was approved for publication on 2020-12-03 at 14:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16045 on 2021-03-04 at 15:36:02","Made available in DSpace on 2021-03-05T21:38:20Z (GMT). No. of bitstreams: 3 PAN-DISSERTATION-2020.pdf: 3042895 bytes, checksum: aa6402cdf19336e3cf816b57f7db4a3e (MD5) LICENSE.txt: 4208 bytes, checksum: b77f4a45cd5af6da81ae5c04ebf2a41f (MD5) PROQUEST_LICENSE.txt: 4554 bytes, checksum: 9cc63616e8f43b9269af5b4c211c8d36 (MD5) Previous issue date: 2020-12-03","Information provided in languages that people can understand saves lives in crises. For example, the language barrier was one of the main difficulties faced by humanitarian workers responding to the Ebola crisis in 2014. We propose to break language barriers by extracting information (e.g., entities) from a massive variety of languages and ground the information into an existing Knowledge Base (KB) which is accessible to a user in their own language (e.g., a reporter from the World Health Organization who speaks English only). The ambitious goal of this thesis is to develop a Cross-lingual Entity Extraction and Linking framework for 1,000 fine-grained entity types and 300 languages that exist in Wikipedia. Given a document in any of these languages, our framework is able to identify entity name mentions, assign a fine-grained type to each mention, and link it to an English KB if it is linkable. Traditional entity linking methods rely on costly human-annotated data to train supervised learning-to-rank models to select the best candidate entity for each mention. In contrast, we propose a novel unsupervised represent-and-compare approach that can accurately capture the semantic meaning representation of each mention, and directly compare its representation with the representation of each candidate entity in the target KB. First, we leverage a deep symbolic semantic representation of the Abstract Meaning Representation to represent contextual properties of mentions. Then we enrich the representation of each contextual word and entity mention with a novel distributed semantic representation based on cross-lingual joint entity and word embedding. We develop a novel method to generate cross-lingual data that is a mix of entities and contextual words based on Wikipedia. This distributed semantics enables effective entity extraction and linking. Because the joint entity and word embedding space is constructed across languages, we further extend it to all 300 Wikipedia languages and fine-grained entity extraction and linking for 1,000 entity types defined in YAGO. Finally, using knowledge-driven question answering as a case study, we demonstrate the effectiveness of acquiring external knowledge using entity extraction and linking to improve downstream applications.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-03-04 without embargo terms"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Cross-lingual entity extraction and linking for 300 languages"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng","Han, Jiawei","Tong, Hanghang","Knight, Kevin"],"dc:creator":["Pan, Xiaoman"],"dc:date":["2021-03-05T21:38:20Z","2020-12-03","2020-12"],"dc:description":["The student, Xiaoman Pan, accepted the attached license on 2020-12-02 at 17:38.","The student, Xiaoman Pan, submitted this Dissertation for approval on 2020-12-02 at 17:43.","This Dissertation was approved for publication on 2020-12-03 at 14:21.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16045 on 2021-03-04 at 15:36:02","Made available in DSpace on 2021-03-05T21:38:20Z (GMT). No. of bitstreams: 3 PAN-DISSERTATION-2020.pdf: 3042895 bytes, checksum: aa6402cdf19336e3cf816b57f7db4a3e (MD5) LICENSE.txt: 4208 bytes, checksum: b77f4a45cd5af6da81ae5c04ebf2a41f (MD5) PROQUEST_LICENSE.txt: 4554 bytes, checksum: 9cc63616e8f43b9269af5b4c211c8d36 (MD5) Previous issue date: 2020-12-03","Information provided in languages that people can understand saves lives in crises. For example, the language barrier was one of the main difficulties faced by humanitarian workers responding to the Ebola crisis in 2014. We propose to break language barriers by extracting information (e.g., entities) from a massive variety of languages and ground the information into an existing Knowledge Base (KB) which is accessible to a user in their own language (e.g., a reporter from the World Health Organization who speaks English only). The ambitious goal of this thesis is to develop a Cross-lingual Entity Extraction and Linking framework for 1,000 fine-grained entity types and 300 languages that exist in Wikipedia. Given a document in any of these languages, our framework is able to identify entity name mentions, assign a fine-grained type to each mention, and link it to an English KB if it is linkable. Traditional entity linking methods rely on costly human-annotated data to train supervised learning-to-rank models to select the best candidate entity for each mention. In contrast, we propose a novel unsupervised represent-and-compare approach that can accurately capture the semantic meaning representation of each mention, and directly compare its representation with the representation of each candidate entity in the target KB. First, we leverage a deep symbolic semantic representation of the Abstract Meaning Representation to represent contextual properties of mentions. Then we enrich the representation of each contextual word and entity mention with a novel distributed semantic representation based on cross-lingual joint entity and word embedding. We develop a novel method to generate cross-lingual data that is a mix of entities and contextual words based on Wikipedia. This distributed semantics enables effective entity extraction and linking. Because the joint entity and word embedding space is constructed across languages, we further extend it to all 300 Wikipedia languages and fine-grained entity extraction and linking for 1,000 entity types defined in YAGO. Finally, using knowledge-driven question answering as a case study, we demonstrate the effectiveness of acquiring external knowledge using entity extraction and linking to improve downstream applications.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-03-04 without embargo terms"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/109431"],"dc:language":["en"],"dc:rights":["Copyright 2020 Xiaoman Pan"],"dc:subject":["cross-lingual","entity extraction","entity linking"],"dc:title":["Cross-lingual entity extraction and linking for 300 languages"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:50Z"}