{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105826"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105826","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"The surprising effectiveness of explicit semantic analysis in dataless classification","abstract":"Organizing textual content into broad labels is one of the most basic tasks that some people carry out on a regular basis. This simple task helps people navigate through large document collections by exposing the labels of the documents, which can then be used for selecting the documents of interest. Currently, the most popular techniques for providing this basic functionality are supervised in nature, wherein someone has to annotate a collection of documents with the labels of interest. However, it might not always be possible to create a sizeable labeled dataset for every scenario or domain of interest. Thus, techniques like “Dataless Classification” have been proposed in the past that are able to bootstrap the creation of a classifier by only requiring semantic descriptions of the labels. However, despite the encouraging performance of Dataless Classification on Text Classification tasks, there is still a room for large improvement. In this thesis, we identify the limitations of ESA-driven Dataless Classification and systematically design techniques for addressing each limitation. In the process, we end up developing 4 new embeddings – EntityESA, Entity2Vec, Topic2Vec and Word2Concept. However, despite our best efforts, we found it difficult to outperform the original Dataless Classification system. For some of the techniques we provide an explanation for this observed behavior, however we also attribute some of these observations to the datasets that are being used for evaluation purposes. We then propose a way to create a new dataset that can used for future Dataless evaluations. The new embedding methods proposed in this work are generic enough that they can be of independent interest as well.","abstract_html":"Organizing textual content into broad labels is one of the most basic tasks that some people carry out on a regular basis. This simple task helps people navigate through large document collections by exposing the labels of the documents, which can then be used for selecting the documents of interest. Currently, the most popular techniques for providing this basic functionality are supervised in nature, wherein someone has to annotate a collection of documents with the labels of interest. However, it might not always be possible to create a sizeable labeled dataset for every scenario or domain of interest. Thus, techniques like “Dataless Classification” have been proposed in the past that are able to bootstrap the creation of a classifier by only requiring semantic descriptions of the labels. However, despite the encouraging performance of Dataless Classification on Text Classification tasks, there is still a room for large improvement. In this thesis, we identify the limitations of ESA-driven Dataless Classification and systematically design techniques for addressing each limitation. In the process, we end up developing 4 new embeddings – EntityESA, Entity2Vec, Topic2Vec and Word2Concept. However, despite our best efforts, we found it difficult to outperform the original Dataless Classification system. For some of the techniques we provide an explanation for this observed behavior, however we also attribute some of these observations to the datasets that are being used for evaluation purposes. We then propose a way to create a new dataset that can used for future Dataless evaluations. The new embedding methods proposed in this work are generic enough that they can be of independent interest as well.","abstract_has_math":false,"creators":["Gupta, Shashank"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Roth, Dan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-11-26T20:49:31Z","date_published":"2019-11-26T20:49:31Z","updated_at":"2026-07-22T22:24:45Z","subjects":["ESA","Dataless Classification","Embeddings","Unsupervised Learning","EntityESA","Entity2Vec","Topic2Vec","Word2Concept"],"languages":["en"],"rights":["Copyright 2019 Shashank Gupta"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105826","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Roth, Dan"]},{"key":"dc:creator","label":"Author","values":["Gupta, Shashank"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-11-26T20:49:31Z","2021-11-27T10:15:20Z","2019-07-16","2019-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["ESA","Dataless Classification","Embeddings","Unsupervised Learning","EntityESA","Entity2Vec","Topic2Vec","Word2Concept"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Shashank Gupta"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105826"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Organizing textual content into broad labels is one of the most basic tasks that some people carry out on a regular basis. This simple task helps people navigate through large document collections by exposing the labels of the documents, which can then be used for selecting the documents of interest. Currently, the most popular techniques for providing this basic functionality are supervised in nature, wherein someone has to annotate a collection of documents with the labels of interest. However, it might not always be possible to create a sizeable labeled dataset for every scenario or domain of interest. Thus, techniques like “Dataless Classification” have been proposed in the past that are able to bootstrap the creation of a classifier by only requiring semantic descriptions of the labels. However, despite the encouraging performance of Dataless Classification on Text Classification tasks, there is still a room for large improvement. In this thesis, we identify the limitations of ESA-driven Dataless Classification and systematically design techniques for addressing each limitation. In the process, we end up developing 4 new embeddings – EntityESA, Entity2Vec, Topic2Vec and Word2Concept. However, despite our best efforts, we found it difficult to outperform the original Dataless Classification system. For some of the techniques we provide an explanation for this observed behavior, however we also attribute some of these observations to the datasets that are being used for evaluation purposes. We then propose a way to create a new dataset that can used for future Dataless evaluations. The new embedding methods proposed in this work are generic enough that they can be of independent interest as well.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-08-01","The student, Shashank Gupta, accepted the attached license on 2019-07-15 at 16:16.","The student, Shashank Gupta, submitted this Thesis for approval on 2019-07-15 at 16:30.","This Thesis was approved for publication on 2019-07-16 at 09:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14330 on 2019-11-26 at 13:05:59","Made available in DSpace on 2019-11-26T20:49:31Z (GMT). No. of bitstreams: 2 GUPTA-THESIS-2019.pdf: 2559723 bytes, checksum: 627d25e8a76ba33bd36b7f7048a8e3c7 (MD5) LICENSE.txt: 4211 bytes, checksum: c89953f4374346e995fc10a787018c72 (MD5) Previous issue date: 2019-07-16","Embargo set by: Seth Robbins for item 112971 Lift date: 2021-11-26T20:49:41Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 112971 on 2021-11-27T10:15:20Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["The surprising effectiveness of explicit semantic analysis in dataless classification"]}]}],"canonical_facts":{"dc:contributor":["Roth, Dan"],"dc:creator":["Gupta, Shashank"],"dc:date":["2019-11-26T20:49:31Z","2021-11-27T10:15:20Z","2019-07-16","2019-08"],"dc:description":["Organizing textual content into broad labels is one of the most basic tasks that some people carry out on a regular basis. This simple task helps people navigate through large document collections by exposing the labels of the documents, which can then be used for selecting the documents of interest. Currently, the most popular techniques for providing this basic functionality are supervised in nature, wherein someone has to annotate a collection of documents with the labels of interest. However, it might not always be possible to create a sizeable labeled dataset for every scenario or domain of interest. Thus, techniques like “Dataless Classification” have been proposed in the past that are able to bootstrap the creation of a classifier by only requiring semantic descriptions of the labels. However, despite the encouraging performance of Dataless Classification on Text Classification tasks, there is still a room for large improvement. In this thesis, we identify the limitations of ESA-driven Dataless Classification and systematically design techniques for addressing each limitation. In the process, we end up developing 4 new embeddings – EntityESA, Entity2Vec, Topic2Vec and Word2Concept. However, despite our best efforts, we found it difficult to outperform the original Dataless Classification system. For some of the techniques we provide an explanation for this observed behavior, however we also attribute some of these observations to the datasets that are being used for evaluation purposes. We then propose a way to create a new dataset that can used for future Dataless evaluations. The new embedding methods proposed in this work are generic enough that they can be of independent interest as well.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-08-01","The student, Shashank Gupta, accepted the attached license on 2019-07-15 at 16:16.","The student, Shashank Gupta, submitted this Thesis for approval on 2019-07-15 at 16:30.","This Thesis was approved for publication on 2019-07-16 at 09:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14330 on 2019-11-26 at 13:05:59","Made available in DSpace on 2019-11-26T20:49:31Z (GMT). No. of bitstreams: 2 GUPTA-THESIS-2019.pdf: 2559723 bytes, checksum: 627d25e8a76ba33bd36b7f7048a8e3c7 (MD5) LICENSE.txt: 4211 bytes, checksum: c89953f4374346e995fc10a787018c72 (MD5) Previous issue date: 2019-07-16","Embargo set by: Seth Robbins for item 112971 Lift date: 2021-11-26T20:49:41Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 112971 on 2021-11-27T10:15:20Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105826"],"dc:language":["en"],"dc:rights":["Copyright 2019 Shashank Gupta"],"dc:subject":["ESA","Dataless Classification","Embeddings","Unsupervised Learning","EntityESA","Entity2Vec","Topic2Vec","Word2Concept"],"dc:title":["The surprising effectiveness of explicit semantic analysis in dataless classification"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:45Z"}