{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/97395"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/97395","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Autoentity: automated entity detection from massive text corpora","abstract":"Entity detection is one of the fundamental tasks in Natural Language Processing and Information Retrieval. Most existing methods rely on human annotated data and hand-crafted linguistic features, which makes it hard to apply the model to an emerging domain. In this paper, we propose a novel automated entity detection framework, called AutoEntity, that performs automated phrase mining to create entity mention candidates and enforces lexico-syntactic rules to select entity mentions from candidates. Our experiments on real-world datasets in different domains and multiple languages have demonstrated the effectiveness and robustness of the proposed method.","abstract_html":"Entity detection is one of the fundamental tasks in Natural Language Processing and Information Retrieval. Most existing methods rely on human annotated data and hand-crafted linguistic features, which makes it hard to apply the model to an emerging domain. In this paper, we propose a novel automated entity detection framework, called AutoEntity, that performs automated phrase mining to create entity mention candidates and enforces lexico-syntactic rules to select entity mentions from candidates. Our experiments on real-world datasets in different domains and multiple languages have demonstrated the effectiveness and robustness of the proposed method.","abstract_has_math":false,"creators":["He, Wenqi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-08-10T19:15:23Z","date_published":"2017-08-10T19:15:23Z","updated_at":"2026-07-22T22:24:34Z","subjects":["Entity detection","Phrase mining"],"languages":["en"],"rights":["Copyright 2017 Wenqi He"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/97395","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei"]},{"key":"dc:creator","label":"Author","values":["He, Wenqi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-08-10T19:15:23Z","2017-04-24","2017-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Entity detection","Phrase mining"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Wenqi He"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/97395"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Entity detection is one of the fundamental tasks in Natural Language Processing and Information Retrieval. Most existing methods rely on human annotated data and hand-crafted linguistic features, which makes it hard to apply the model to an emerging domain. In this paper, we propose a novel automated entity detection framework, called AutoEntity, that performs automated phrase mining to create entity mention candidates and enforces lexico-syntactic rules to select entity mentions from candidates. Our experiments on real-world datasets in different domains and multiple languages have demonstrated the effectiveness and robustness of the proposed method.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-08-10 without embargo terms","The student, Wenqi He, accepted the attached license on 2017-04-18 at 15:57.","The student, Wenqi He, submitted this Thesis for approval on 2017-04-18 at 16:04.","This Thesis was approved for publication on 2017-04-24 at 09:55.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10863 on 2017-08-10 at 13:42:07","Made available in DSpace on 2017-08-10T19:15:23Z (GMT). No. of bitstreams: 2 HE-THESIS-2017.pdf: 709327 bytes, checksum: 21ce0bc3039a9b9fc1cbfa9373bada29 (MD5) LICENSE.txt: 4205 bytes, checksum: b6c5324a07f04d495e4e3b9305210e2f (MD5) Previous issue date: 2017-04-24"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Autoentity: automated entity detection from massive text corpora"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei"],"dc:creator":["He, Wenqi"],"dc:date":["2017-08-10T19:15:23Z","2017-04-24","2017-05"],"dc:description":["Entity detection is one of the fundamental tasks in Natural Language Processing and Information Retrieval. Most existing methods rely on human annotated data and hand-crafted linguistic features, which makes it hard to apply the model to an emerging domain. In this paper, we propose a novel automated entity detection framework, called AutoEntity, that performs automated phrase mining to create entity mention candidates and enforces lexico-syntactic rules to select entity mentions from candidates. Our experiments on real-world datasets in different domains and multiple languages have demonstrated the effectiveness and robustness of the proposed method.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-08-10 without embargo terms","The student, Wenqi He, accepted the attached license on 2017-04-18 at 15:57.","The student, Wenqi He, submitted this Thesis for approval on 2017-04-18 at 16:04.","This Thesis was approved for publication on 2017-04-24 at 09:55.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10863 on 2017-08-10 at 13:42:07","Made available in DSpace on 2017-08-10T19:15:23Z (GMT). No. of bitstreams: 2 HE-THESIS-2017.pdf: 709327 bytes, checksum: 21ce0bc3039a9b9fc1cbfa9373bada29 (MD5) LICENSE.txt: 4205 bytes, checksum: b6c5324a07f04d495e4e3b9305210e2f (MD5) Previous issue date: 2017-04-24"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/97395"],"dc:language":["en"],"dc:rights":["Copyright 2017 Wenqi He"],"dc:subject":["Entity detection","Phrase mining"],"dc:title":["Autoentity: automated entity detection from massive text corpora"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:34Z"}