{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/92851"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/92851","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Semantic modeling of the natural language of Wikipedia annotations","abstract":"Knowledge bases (KB) store relational facts and constitute a significant resource for a variety of natural language processing (NLP) tasks. Improving their coverage and refining the relations is a basic and pressing research effort. In this thesis we propose a novel approach towards this canonical task by using the unstructured Wikipedia corpus: we extract low-dimensional embeddings for title pages of the Wikipedia corpus and show that they can be used to significantly outperform state-of-the-art approaches on a variety of metrics in three concrete tasks: measuring semantic relatedness, solving semantic analogies, and KB completion and refinement. A central feature of our work is a new log-linear discriminative model for the annotations inside a Wikipedia document that we name IBOE (isotropic bag-of-entities): we hypothesize that the parameters of the model satisfy a geometric symmetry property (isotropy). We show that the isotropy property leads to self-normalization allowing for the design of an efficient parameter estimation algorithm that we christen wiki2vec. The self-normalization property of IBOE is validated empirically on the Wikipedia corpus and is also of independent mathematical interest.","abstract_html":"Knowledge bases (KB) store relational facts and constitute a significant resource for a variety of natural language processing (NLP) tasks. Improving their coverage and refining the relations is a basic and pressing research effort. In this thesis we propose a novel approach towards this canonical task by using the unstructured Wikipedia corpus: we extract low-dimensional embeddings for title pages of the Wikipedia corpus and show that they can be used to significantly outperform state-of-the-art approaches on a variety of metrics in three concrete tasks: measuring semantic relatedness, solving semantic analogies, and KB completion and refinement. A central feature of our work is a new log-linear discriminative model for the annotations inside a Wikipedia document that we name IBOE (isotropic bag-of-entities): we hypothesize that the parameters of the model satisfy a geometric symmetry property (isotropy). We show that the isotropy property leads to self-normalization allowing for the design of an efficient parameter estimation algorithm that we christen wiki2vec. The self-normalization property of IBOE is validated empirically on the Wikipedia corpus and is also of independent mathematical interest.","abstract_has_math":false,"creators":["Mu, Jiaqi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Viswanath, Pramod","Bhat, Suma P."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-11-10T17:55:14Z","date_published":"2016-11-10T17:55:14Z","updated_at":"2026-07-22T22:26:35Z","subjects":["entity embedding","knowledge base completion"],"languages":["en"],"rights":["Copyright 2016 Jiaqi Mu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/92851","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Viswanath, Pramod","Bhat, Suma P."]},{"key":"dc:creator","label":"Author","values":["Mu, Jiaqi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-11-10T17:55:14Z","2016-07-15","2016-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["entity embedding","knowledge base completion"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2016 Jiaqi Mu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/92851"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Knowledge bases (KB) store relational facts and constitute a significant resource for a variety of natural language processing (NLP) tasks. Improving their coverage and refining the relations is a basic and pressing research effort. In this thesis we propose a novel approach towards this canonical task by using the unstructured Wikipedia corpus: we extract low-dimensional embeddings for title pages of the Wikipedia corpus and show that they can be used to significantly outperform state-of-the-art approaches on a variety of metrics in three concrete tasks: measuring semantic relatedness, solving semantic analogies, and KB completion and refinement. A central feature of our work is a new log-linear discriminative model for the annotations inside a Wikipedia document that we name IBOE (isotropic bag-of-entities): we hypothesize that the parameters of the model satisfy a geometric symmetry property (isotropy). We show that the isotropy property leads to self-normalization allowing for the design of an efficient parameter estimation algorithm that we christen wiki2vec. The self-normalization property of IBOE is validated empirically on the Wikipedia corpus and is also of independent mathematical interest.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-11-09 without embargo terms","The student, Jiaqi Mu, accepted the attached license on 2016-07-15 at 10:25.","The student, Jiaqi Mu, submitted this Thesis for approval on 2016-07-15 at 10:39.","This Thesis was approved for publication on 2016-07-15 at 13:08.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9961 on 2016-11-09 at 10:25:14","Made available in DSpace on 2016-11-10T17:55:14Z (GMT). No. of bitstreams: 2 MU-THESIS-2016.pdf: 1259589 bytes, checksum: 0c2de48d9460151f7e468008a704635e (MD5) LICENSE.txt: 4205 bytes, checksum: 1db9778bc0832fa860692456466d1b5a (MD5) Previous issue date: 2016-07-15"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Semantic modeling of the natural language of Wikipedia annotations"]}]}],"canonical_facts":{"dc:contributor":["Viswanath, Pramod","Bhat, Suma P."],"dc:creator":["Mu, Jiaqi"],"dc:date":["2016-11-10T17:55:14Z","2016-07-15","2016-08"],"dc:description":["Knowledge bases (KB) store relational facts and constitute a significant resource for a variety of natural language processing (NLP) tasks. Improving their coverage and refining the relations is a basic and pressing research effort. In this thesis we propose a novel approach towards this canonical task by using the unstructured Wikipedia corpus: we extract low-dimensional embeddings for title pages of the Wikipedia corpus and show that they can be used to significantly outperform state-of-the-art approaches on a variety of metrics in three concrete tasks: measuring semantic relatedness, solving semantic analogies, and KB completion and refinement. A central feature of our work is a new log-linear discriminative model for the annotations inside a Wikipedia document that we name IBOE (isotropic bag-of-entities): we hypothesize that the parameters of the model satisfy a geometric symmetry property (isotropy). We show that the isotropy property leads to self-normalization allowing for the design of an efficient parameter estimation algorithm that we christen wiki2vec. The self-normalization property of IBOE is validated empirically on the Wikipedia corpus and is also of independent mathematical interest.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-11-09 without embargo terms","The student, Jiaqi Mu, accepted the attached license on 2016-07-15 at 10:25.","The student, Jiaqi Mu, submitted this Thesis for approval on 2016-07-15 at 10:39.","This Thesis was approved for publication on 2016-07-15 at 13:08.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9961 on 2016-11-09 at 10:25:14","Made available in DSpace on 2016-11-10T17:55:14Z (GMT). No. of bitstreams: 2 MU-THESIS-2016.pdf: 1259589 bytes, checksum: 0c2de48d9460151f7e468008a704635e (MD5) LICENSE.txt: 4205 bytes, checksum: 1db9778bc0832fa860692456466d1b5a (MD5) Previous issue date: 2016-07-15"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/92851"],"dc:language":["en"],"dc:rights":["Copyright 2016 Jiaqi Mu"],"dc:subject":["entity embedding","knowledge base completion"],"dc:title":["Semantic modeling of the natural language of Wikipedia annotations"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:35Z"}