{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/88088"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/88088","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Entity recognition for multi-modal socio-technical systems","abstract":"Entity Recognition (ER) can be used as a method for extracting information about socio-technical systems from unstructured, natural language text data. This process is limited by the set of entity classes considered in many current ER solutions. In this thesis, we report on the development of an ER classifier that supports a wide range of entity classes that are relevant for analyzing multi-modal, socio-technical systems. Another limitation with current entity extractors is that they mainly support the detection of named entities, typically in the form of proper nouns. The presented solution also detects entities not referred to by a name, such as general references to places (e.g. forest) or natural resources (e.g. timber). We use supervised machine learning for this project. To overcome data sparseness issues that results from considering a large number of entity classes, we built two separate classifiers for predicting labels for entity boundary and class. We herein investigate rules for merging both labels while minimizing the loss of accuracy due to this step. The accuracy of our classifier for the largest model with 94 classes achieves 75.9%. We compare the performance of our solution to other standard systems on several datasets, finding that with the same number of classes, the accuracy of our classifier is comparable to other state-of-the-art ER packages.","abstract_html":"Entity Recognition (ER) can be used as a method for extracting information about socio-technical systems from unstructured, natural language text data. This process is limited by the set of entity classes considered in many current ER solutions. In this thesis, we report on the development of an ER classifier that supports a wide range of entity classes that are relevant for analyzing multi-modal, socio-technical systems. Another limitation with current entity extractors is that they mainly support the detection of named entities, typically in the form of proper nouns. The presented solution also detects entities not referred to by a name, such as general references to places (e.g. forest) or natural resources (e.g. timber). We use supervised machine learning for this project. To overcome data sparseness issues that results from considering a large number of entity classes, we built two separate classifiers for predicting labels for entity boundary and class. We herein investigate rules for merging both labels while minimizing the loss of accuracy due to this step. The accuracy of our classifier for the largest model with 94 classes achieves 75.9%. We compare the performance of our solution to other standard systems on several datasets, finding that with the same number of classes, the accuracy of our classifier is comparable to other state-of-the-art ER packages.","abstract_has_math":false,"creators":["Aleyasen, Amirhossein"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Winslett, Marianne","Diesner, Jana"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-09-29T20:38:43Z","date_published":"2015-09-29T20:38:43Z","updated_at":"2026-07-22T22:26:31Z","subjects":["Entity Recognition","Supervised Learning","Information Extraction","Conditional Random Field"],"languages":["en"],"rights":["Copyright 2015 Amirhossein Aleyasen"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/88088","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Winslett, Marianne","Diesner, Jana"]},{"key":"dc:creator","label":"Author","values":["Aleyasen, Amirhossein"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-09-29T20:38:43Z","2015-08","2015-07-22","2015-8"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Entity Recognition","Supervised Learning","Information Extraction","Conditional Random Field"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2015 Amirhossein Aleyasen"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/88088"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Entity Recognition (ER) can be used as a method for extracting information about socio-technical systems from unstructured, natural language text data. This process is limited by the set of entity classes considered in many current ER solutions. In this thesis, we report on the development of an ER classifier that supports a wide range of entity classes that are relevant for analyzing multi-modal, socio-technical systems. Another limitation with current entity extractors is that they mainly support the detection of named entities, typically in the form of proper nouns. The presented solution also detects entities not referred to by a name, such as general references to places (e.g. forest) or natural resources (e.g. timber). We use supervised machine learning for this project. To overcome data sparseness issues that results from considering a large number of entity classes, we built two separate classifiers for predicting labels for entity boundary and class. We herein investigate rules for merging both labels while minimizing the loss of accuracy due to this step. The accuracy of our classifier for the largest model with 94 classes achieves 75.9%. We compare the performance of our solution to other standard systems on several datasets, finding that with the same number of classes, the accuracy of our classifier is comparable to other state-of-the-art ER packages.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2015-09-29 without embargo terms","The student, Amirhossein Aleyasen, accepted the attached license on 2015-07-20 at 11:18.","The student, Amirhossein Aleyasen, submitted this Thesis for approval on 2015-07-22 at 11:53.","This Thesis was approved for publication on 2015-07-22 at 12:04.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8559 on 2015-09-29 at 13:23:17","Made available in DSpace on 2015-09-29T20:38:43Z (GMT). No. of bitstreams: 3 ALEYASEN-THESIS-2015.pdf: 317989 bytes, checksum: f7fc9ac9e6eb39aec703d808e50d92a5 (MD5) Master-Thesis-Aleyasen.zip: 396461 bytes, checksum: da3ff3ae6ee7aa81314f10df13aa8669 (MD5) LICENSE.txt: 4217 bytes, checksum: 1e5aa8050a9e5d5365fa56737a141c11 (MD5) Previous issue date: 2015-07-22"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Entity recognition for multi-modal socio-technical systems"]}]}],"canonical_facts":{"dc:contributor":["Winslett, Marianne","Diesner, Jana"],"dc:creator":["Aleyasen, Amirhossein"],"dc:date":["2015-09-29T20:38:43Z","2015-08","2015-07-22","2015-8"],"dc:description":["Entity Recognition (ER) can be used as a method for extracting information about socio-technical systems from unstructured, natural language text data. This process is limited by the set of entity classes considered in many current ER solutions. In this thesis, we report on the development of an ER classifier that supports a wide range of entity classes that are relevant for analyzing multi-modal, socio-technical systems. Another limitation with current entity extractors is that they mainly support the detection of named entities, typically in the form of proper nouns. The presented solution also detects entities not referred to by a name, such as general references to places (e.g. forest) or natural resources (e.g. timber). We use supervised machine learning for this project. To overcome data sparseness issues that results from considering a large number of entity classes, we built two separate classifiers for predicting labels for entity boundary and class. We herein investigate rules for merging both labels while minimizing the loss of accuracy due to this step. The accuracy of our classifier for the largest model with 94 classes achieves 75.9%. We compare the performance of our solution to other standard systems on several datasets, finding that with the same number of classes, the accuracy of our classifier is comparable to other state-of-the-art ER packages.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2015-09-29 without embargo terms","The student, Amirhossein Aleyasen, accepted the attached license on 2015-07-20 at 11:18.","The student, Amirhossein Aleyasen, submitted this Thesis for approval on 2015-07-22 at 11:53.","This Thesis was approved for publication on 2015-07-22 at 12:04.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8559 on 2015-09-29 at 13:23:17","Made available in DSpace on 2015-09-29T20:38:43Z (GMT). No. of bitstreams: 3 ALEYASEN-THESIS-2015.pdf: 317989 bytes, checksum: f7fc9ac9e6eb39aec703d808e50d92a5 (MD5) Master-Thesis-Aleyasen.zip: 396461 bytes, checksum: da3ff3ae6ee7aa81314f10df13aa8669 (MD5) LICENSE.txt: 4217 bytes, checksum: 1e5aa8050a9e5d5365fa56737a141c11 (MD5) Previous issue date: 2015-07-22"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/88088"],"dc:language":["en"],"dc:rights":["Copyright 2015 Amirhossein Aleyasen"],"dc:subject":["Entity Recognition","Supervised Learning","Information Extraction","Conditional Random Field"],"dc:title":["Entity recognition for multi-modal socio-technical systems"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:31Z"}