{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/104934"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/104934","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Character language models for generalization of multilingual named entity recognition","abstract":"\"State-of-the-art Named Entity Recognition (NER) models usually achieve high performance on entities that they have seen in training data, but a significantly lower performance on unseen entities. This is one of the key reasons in performance degradation observed when NER models are evaluated on new domains. Motivated by this observation, quantified for the first time in this thesis, we study an improved, multi-domain and multi-lingual, capability for identifying \\what is a name\"\". Character-level patterns have been widely used as features in English Named Entity Recognition (NER) systems. However, to date there has been no direct investigation of the inherent differences between name and non-name tokens in text, nor whether this property holds across multiple languages. The key contribution of this thesis is to develop a Character-level Language Model (CLM) that, as we show, allow us to better learn \\what is a name\"\". We analyze the capabilities of corpus-agnostic Character-level Language Models (CLMs) in the binary task of distinguishing name tokens from non-name tokens and demonstrate that CLMs provide a simple yet powerful model for capturing these differences. Specifically, we show that it can identify named entity tokens in a diverse set of languages at close to the performance of full NER systems. Moreover, by adding very simple CLM-based features we can significantly improve the performance of an o -the-shelf NER system for multiple languages.\"","abstract_html":"&quot;State-of-the-art Named Entity Recognition (NER) models usually achieve high performance on entities that they have seen in training data, but a significantly lower performance on unseen entities. This is one of the key reasons in performance degradation observed when NER models are evaluated on new domains. Motivated by this observation, quantified for the first time in this thesis, we study an improved, multi-domain and multi-lingual, capability for identifying \\what is a name&quot;&quot;. Character-level patterns have been widely used as features in English Named Entity Recognition (NER) systems. However, to date there has been no direct investigation of the inherent differences between name and non-name tokens in text, nor whether this property holds across multiple languages. The key contribution of this thesis is to develop a Character-level Language Model (CLM) that, as we show, allow us to better learn \\what is a name&quot;&quot;. We analyze the capabilities of corpus-agnostic Character-level Language Models (CLMs) in the binary task of distinguishing name tokens from non-name tokens and demonstrate that CLMs provide a simple yet powerful model for capturing these differences. Specifically, we show that it can identify named entity tokens in a diverse set of languages at close to the performance of full NER systems. Moreover, by adding very simple CLM-based features we can significantly improve the performance of an o -the-shelf NER system for multiple languages.&quot;","abstract_has_math":false,"creators":["Yu, Xiaodong"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Roth, Dan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T20:02:10Z","date_published":"2019-08-23T20:02:10Z","updated_at":"2026-07-22T22:24:42Z","subjects":["Character Language Models","Named Entity Recognition","Generalization","Multilingual","Multilingual Named Entity Recognition","NER"],"languages":["en"],"rights":["Copyright 2019 Xiaodong Yu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/104934","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Roth, Dan"]},{"key":"dc:creator","label":"Author","values":["Yu, Xiaodong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T20:02:10Z","2019-04-25","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Character Language Models","Named Entity Recognition","Generalization","Multilingual","Multilingual Named Entity Recognition","NER"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Xiaodong Yu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/104934"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"State-of-the-art Named Entity Recognition (NER) models usually achieve high performance on entities that they have seen in training data, but a significantly lower performance on unseen entities. This is one of the key reasons in performance degradation observed when NER models are evaluated on new domains. Motivated by this observation, quantified for the first time in this thesis, we study an improved, multi-domain and multi-lingual, capability for identifying \\what is a name\"\". Character-level patterns have been widely used as features in English Named Entity Recognition (NER) systems. However, to date there has been no direct investigation of the inherent differences between name and non-name tokens in text, nor whether this property holds across multiple languages. The key contribution of this thesis is to develop a Character-level Language Model (CLM) that, as we show, allow us to better learn \\what is a name\"\". We analyze the capabilities of corpus-agnostic Character-level Language Models (CLMs) in the binary task of distinguishing name tokens from non-name tokens and demonstrate that CLMs provide a simple yet powerful model for capturing these differences. Specifically, we show that it can identify named entity tokens in a diverse set of languages at close to the performance of full NER systems. Moreover, by adding very simple CLM-based features we can significantly improve the performance of an o -the-shelf NER system for multiple languages.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Xiaodong Yu, accepted the attached license on 2019-04-24 at 20:55.","The student, Xiaodong Yu, submitted this Thesis for approval on 2019-04-24 at 21:02.","This Thesis was approved for publication on 2019-04-25 at 12:01.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13899 on 2019-08-22 at 14:46:43","Made available in DSpace on 2019-08-23T20:02:10Z (GMT). No. of bitstreams: 2 YU-THESIS-2019.pdf: 535492 bytes, checksum: 99324f33e6ae79e26537d750db6b2839 (MD5) LICENSE.txt: 4208 bytes, checksum: 2c772a6052c5da92f39462c193abfa99 (MD5) Previous issue date: 2019-04-25"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Character language models for generalization of multilingual named entity recognition"]}]}],"canonical_facts":{"dc:contributor":["Roth, Dan"],"dc:creator":["Yu, Xiaodong"],"dc:date":["2019-08-23T20:02:10Z","2019-04-25","2019-05"],"dc:description":["\"State-of-the-art Named Entity Recognition (NER) models usually achieve high performance on entities that they have seen in training data, but a significantly lower performance on unseen entities. This is one of the key reasons in performance degradation observed when NER models are evaluated on new domains. Motivated by this observation, quantified for the first time in this thesis, we study an improved, multi-domain and multi-lingual, capability for identifying \\what is a name\"\". Character-level patterns have been widely used as features in English Named Entity Recognition (NER) systems. However, to date there has been no direct investigation of the inherent differences between name and non-name tokens in text, nor whether this property holds across multiple languages. The key contribution of this thesis is to develop a Character-level Language Model (CLM) that, as we show, allow us to better learn \\what is a name\"\". We analyze the capabilities of corpus-agnostic Character-level Language Models (CLMs) in the binary task of distinguishing name tokens from non-name tokens and demonstrate that CLMs provide a simple yet powerful model for capturing these differences. Specifically, we show that it can identify named entity tokens in a diverse set of languages at close to the performance of full NER systems. Moreover, by adding very simple CLM-based features we can significantly improve the performance of an o -the-shelf NER system for multiple languages.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Xiaodong Yu, accepted the attached license on 2019-04-24 at 20:55.","The student, Xiaodong Yu, submitted this Thesis for approval on 2019-04-24 at 21:02.","This Thesis was approved for publication on 2019-04-25 at 12:01.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13899 on 2019-08-22 at 14:46:43","Made available in DSpace on 2019-08-23T20:02:10Z (GMT). No. of bitstreams: 2 YU-THESIS-2019.pdf: 535492 bytes, checksum: 99324f33e6ae79e26537d750db6b2839 (MD5) LICENSE.txt: 4208 bytes, checksum: 2c772a6052c5da92f39462c193abfa99 (MD5) Previous issue date: 2019-04-25"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/104934"],"dc:language":["en"],"dc:rights":["Copyright 2019 Xiaodong Yu"],"dc:subject":["Character Language Models","Named Entity Recognition","Generalization","Multilingual","Multilingual Named Entity Recognition","NER"],"dc:title":["Character language models for generalization of multilingual named entity recognition"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:42Z"}