{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/106421"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/106421","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Geometries of word embeddings","abstract":"Real-valued word embeddings have transformed natural language processing (NLP) applications, recognized for their ability to capture linguistic regularities. Popular examples are word2vec, GloVe, GPT and BERT. Both word2vec and GloVe are static whose word representations are independent of its context, while GPT and BERT are contextualized whose representations will change corresponding to the semantics conveyed by its surrounding words. In this dissertation, we study four problems associated with the geometrics of word embeddings. First, we demonstrate a very simple, and yet counter-intuitive, postprocessing technique -- eliminate the common mean vector and a few top dominating directions from the word vectors -- that achieves better performances on a variety of standard benchmarks than the original ones. Sentences, as a sequence of words, are also important semantic units of natural language. We extend the embeddings of words toward representing sentences by the low-rank subspace spanned by its word vectors. Such an unsupervised representation is empirically validated via semantic textual similarity tasks on 19 different datasets, where it outperforms the sophisticated neural network models by 15% on average. Having a good sentence embedding, in turn, helps improve word representations. This is because a single vector does not suffice to model the polysemous nature of many (frequent) words, i.e., words with multiple meanings. We leverage the sentence representations on em unsupervised polysemy modeling, which we call K-Grassmeans. This approach is quantitatively tested on standard sense induction and disambiguation datasets and present new state-of-the-art results. Finally, we study the contextualized word embeddings. Given the rapid growth of computational power, pretrained language models are proposed to capture common-sense knowledge hidden behind large training corpora and have achieved great success in natural language understanding (NLU) tasks. We study these pretrained language models using influence function, which characterizes the influence of different training samples on the prediction result of each test sample. The empirical discoveries suggest an interesting future research direction: designing novel regularizations to penalize these correlations during fine-tuning.","abstract_html":"Real-valued word embeddings have transformed natural language processing (NLP) applications, recognized for their ability to capture linguistic regularities. Popular examples are word2vec, GloVe, GPT and BERT. Both word2vec and GloVe are static whose word representations are independent of its context, while GPT and BERT are contextualized whose representations will change corresponding to the semantics conveyed by its surrounding words. In this dissertation, we study four problems associated with the geometrics of word embeddings. First, we demonstrate a very simple, and yet counter-intuitive, postprocessing technique -- eliminate the common mean vector and a few top dominating directions from the word vectors -- that achieves better performances on a variety of standard benchmarks than the original ones. Sentences, as a sequence of words, are also important semantic units of natural language. We extend the embeddings of words toward representing sentences by the low-rank subspace spanned by its word vectors. Such an unsupervised representation is empirically validated via semantic textual similarity tasks on 19 different datasets, where it outperforms the sophisticated neural network models by 15% on average. Having a good sentence embedding, in turn, helps improve word representations. This is because a single vector does not suffice to model the polysemous nature of many (frequent) words, i.e., words with multiple meanings. We leverage the sentence representations on em unsupervised polysemy modeling, which we call K-Grassmeans. This approach is quantitatively tested on standard sense induction and disambiguation datasets and present new state-of-the-art results. Finally, we study the contextualized word embeddings. Given the rapid growth of computational power, pretrained language models are proposed to capture common-sense knowledge hidden behind large training corpora and have achieved great success in natural language understanding (NLU) tasks. We study these pretrained language models using influence function, which characterizes the influence of different training samples on the prediction result of each test sample. The empirical discoveries suggest an interesting future research direction: designing novel regularizations to penalize these correlations during fine-tuning.","abstract_has_math":false,"creators":["Mu, Jiaqi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Viswanath, Pramod","Srikant, Rayadurgam","Bhat, Suma","Oh, Sewoong","Sun, Ruoyu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-03-02T22:38:35Z","date_published":"2020-03-02T22:38:35Z","updated_at":"2026-07-22T22:24:47Z","subjects":["word embedding","natural language processing","representation learning"],"languages":["en"],"rights":["Copyright 2019 Jiaqi Mu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/106421","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Viswanath, Pramod","Srikant, Rayadurgam","Bhat, Suma","Oh, Sewoong","Sun, Ruoyu"]},{"key":"dc:creator","label":"Author","values":["Mu, Jiaqi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-03-02T22:38:35Z","2022-03-03T10:15:22Z","2019-09-04","2019-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["word embedding","natural language processing","representation learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Jiaqi Mu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/106421"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Real-valued word embeddings have transformed natural language processing (NLP) applications, recognized for their ability to capture linguistic regularities. Popular examples are word2vec, GloVe, GPT and BERT. Both word2vec and GloVe are static whose word representations are independent of its context, while GPT and BERT are contextualized whose representations will change corresponding to the semantics conveyed by its surrounding words. In this dissertation, we study four problems associated with the geometrics of word embeddings. First, we demonstrate a very simple, and yet counter-intuitive, postprocessing technique -- eliminate the common mean vector and a few top dominating directions from the word vectors -- that achieves better performances on a variety of standard benchmarks than the original ones. Sentences, as a sequence of words, are also important semantic units of natural language. We extend the embeddings of words toward representing sentences by the low-rank subspace spanned by its word vectors. Such an unsupervised representation is empirically validated via semantic textual similarity tasks on 19 different datasets, where it outperforms the sophisticated neural network models by 15% on average. Having a good sentence embedding, in turn, helps improve word representations. This is because a single vector does not suffice to model the polysemous nature of many (frequent) words, i.e., words with multiple meanings. We leverage the sentence representations on em unsupervised polysemy modeling, which we call K-Grassmeans. This approach is quantitatively tested on standard sense induction and disambiguation datasets and present new state-of-the-art results. Finally, we study the contextualized word embeddings. Given the rapid growth of computational power, pretrained language models are proposed to capture common-sense knowledge hidden behind large training corpora and have achieved great success in natural language understanding (NLU) tasks. We study these pretrained language models using influence function, which characterizes the influence of different training samples on the prediction result of each test sample. The empirical discoveries suggest an interesting future research direction: designing novel regularizations to penalize these correlations during fine-tuning.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-12-01","The student, Jiaqi Mu, accepted the attached license on 2019-08-28 at 18:14.","The student, Jiaqi Mu, submitted this Dissertation for approval on 2019-08-28 at 18:18.","This Dissertation was approved for publication on 2019-09-04 at 09:57.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14439 on 2020-02-28 at 17:35:14","Made available in DSpace on 2020-03-02T22:38:35Z (GMT). No. of bitstreams: 2 MU-DISSERTATION-2019.pdf: 1668147 bytes, checksum: b24f307395c606b0aacebd7263cac313 (MD5) LICENSE.txt: 4205 bytes, checksum: 0d1fa1c59f6d1634411ba0449e5f4634 (MD5) Previous issue date: 2019-09-04","Embargo set by: Seth Robbins for item 113965 Lift date: 2022-03-02T22:39:04Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 113965 on 2022-03-03T10:15:22Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Geometries of word embeddings"]}]}],"canonical_facts":{"dc:contributor":["Viswanath, Pramod","Srikant, Rayadurgam","Bhat, Suma","Oh, Sewoong","Sun, Ruoyu"],"dc:creator":["Mu, Jiaqi"],"dc:date":["2020-03-02T22:38:35Z","2022-03-03T10:15:22Z","2019-09-04","2019-12"],"dc:description":["Real-valued word embeddings have transformed natural language processing (NLP) applications, recognized for their ability to capture linguistic regularities. Popular examples are word2vec, GloVe, GPT and BERT. Both word2vec and GloVe are static whose word representations are independent of its context, while GPT and BERT are contextualized whose representations will change corresponding to the semantics conveyed by its surrounding words. In this dissertation, we study four problems associated with the geometrics of word embeddings. First, we demonstrate a very simple, and yet counter-intuitive, postprocessing technique -- eliminate the common mean vector and a few top dominating directions from the word vectors -- that achieves better performances on a variety of standard benchmarks than the original ones. Sentences, as a sequence of words, are also important semantic units of natural language. We extend the embeddings of words toward representing sentences by the low-rank subspace spanned by its word vectors. Such an unsupervised representation is empirically validated via semantic textual similarity tasks on 19 different datasets, where it outperforms the sophisticated neural network models by 15% on average. Having a good sentence embedding, in turn, helps improve word representations. This is because a single vector does not suffice to model the polysemous nature of many (frequent) words, i.e., words with multiple meanings. We leverage the sentence representations on em unsupervised polysemy modeling, which we call K-Grassmeans. This approach is quantitatively tested on standard sense induction and disambiguation datasets and present new state-of-the-art results. Finally, we study the contextualized word embeddings. Given the rapid growth of computational power, pretrained language models are proposed to capture common-sense knowledge hidden behind large training corpora and have achieved great success in natural language understanding (NLU) tasks. We study these pretrained language models using influence function, which characterizes the influence of different training samples on the prediction result of each test sample. The empirical discoveries suggest an interesting future research direction: designing novel regularizations to penalize these correlations during fine-tuning.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-12-01","The student, Jiaqi Mu, accepted the attached license on 2019-08-28 at 18:14.","The student, Jiaqi Mu, submitted this Dissertation for approval on 2019-08-28 at 18:18.","This Dissertation was approved for publication on 2019-09-04 at 09:57.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14439 on 2020-02-28 at 17:35:14","Made available in DSpace on 2020-03-02T22:38:35Z (GMT). No. of bitstreams: 2 MU-DISSERTATION-2019.pdf: 1668147 bytes, checksum: b24f307395c606b0aacebd7263cac313 (MD5) LICENSE.txt: 4205 bytes, checksum: 0d1fa1c59f6d1634411ba0449e5f4634 (MD5) Previous issue date: 2019-09-04","Embargo set by: Seth Robbins for item 113965 Lift date: 2022-03-02T22:39:04Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 113965 on 2022-03-03T10:15:22Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/106421"],"dc:language":["en"],"dc:rights":["Copyright 2019 Jiaqi Mu"],"dc:subject":["word embedding","natural language processing","representation learning"],"dc:title":["Geometries of word embeddings"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:47Z"}