Back to results

University of Illinois at Urbana-Champaign

Character language models for generalization of multilingual named entity recognition

Abstract

dc:description

"State-of-the-art Named Entity Recognition (NER) models usually achieve high performance on entities that they have seen in training data, but a significantly lower performance on unseen entities. This is one of the key reasons in performance degradation observed when NER models are evaluated on new domains. Motivated by this observation, quantified for the first time in this thesis, we study an improved, multi-domain and multi-lingual, capability for identifying \what is a name"". Character-level patterns have been widely used as features in English Named Entity Recognition (NER) systems. However, to date there has been no direct investigation of the inherent differences between name and non-name tokens in text, nor whether this property holds across multiple languages. The key contribution of this thesis is to develop a Character-level Language Model (CLM) that, as we show, allow us to better learn \what is a name"". We analyze the capabilities of corpus-agnostic Character-level Language Models (CLMs) in the binary task of distinguishing name tokens from non-name tokens and demonstrate that CLMs provide a simple yet powerful model for capturing these differences. Specifically, we show that it can identify named entity tokens in a diverse set of languages at close to the performance of full NER systems. Moreover, by adding very simple CLM-based features we can significantly improve the performance of an o -the-shelf NER system for multiple languages."

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2019

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Yu, Xiaodong
Contributors dc:contributor
  • Roth, Dan

Subjects

dc:subject × 6

Rights

dc:rights
Statement dc:rights
  • Copyright 2019 Xiaodong Yu
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/104934
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/104934

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Yu, Xiaodong. Character language models for generalization of multilingual named entity recognition. Thesis thesis, University of Illinois at Urbana-Champaign, 2019. http://hdl.handle.net/2142/104934