{"id":{"repo_id":"carleton","oai_identifier":"oai:carleton.scholaris.ca:20.500.14718/41318"},"canonical_url":"https://search.dev.ndltd.org/etd/carleton/oai:carleton.scholaris.ca:20.500.14718/41318","repository":{"repo_id":"carleton","name":"Carleton University","base_url":"https://carleton.scholaris.ca/server/oai/request"},"display":{"title":"Post-processing Techniques for Word Embedding","abstract":"Word embedding has been a significant breakthrough in natural language processing (NLP). Although word representation has improved remarkably and resulted in better performance in downstream NLP applications, interpretability of word embeddings remains a challenge. Post-processing techniques have been developed to improve the quality of word embeddings by fine-tuning real-valued vectors and eliminating artifacts that arise due to flaws in the training corpus. In this thesis, we propose a simple method to improve the quality of word embeddings and reveal their hidden structure using a post-processing technique. We deploy co-clustering techniques to detect sub-matrices between word meaning and specific dimensions. Furthermore, pre-trained word embeddings based on neural networks often suffer from representation degeneration issues, especially when trained on large datasets. We propose a novel method Orthogonal Auto Encoder with Variational Dropout (OAEVD), which utilizes orthogonal autoencoders and variational dropout techniques to enhance word embedding. The orthogonality constraint encourages more diversity in the latent space, and variational dropout makes the embedding more robust to overfitting. Empirical evaluation on a range of downstream NLP tasks shows that our proposed method effectively improves the quality of pre-trained word embeddings and pre-trained language models. In addition, We introduce a technique to enhance the accuracy of word embeddings by incorporating external knowledge in the form of WordNet pairs to improve the accuracy of word embedding on several NLP tasks.","abstract_html":"Word embedding has been a significant breakthrough in natural language processing (NLP). Although word representation has improved remarkably and resulted in better performance in downstream NLP applications, interpretability of word embeddings remains a challenge. Post-processing techniques have been developed to improve the quality of word embeddings by fine-tuning real-valued vectors and eliminating artifacts that arise due to flaws in the training corpus. In this thesis, we propose a simple method to improve the quality of word embeddings and reveal their hidden structure using a post-processing technique. We deploy co-clustering techniques to detect sub-matrices between word meaning and specific dimensions. Furthermore, pre-trained word embeddings based on neural networks often suffer from representation degeneration issues, especially when trained on large datasets. We propose a novel method Orthogonal Auto Encoder with Variational Dropout (OAEVD), which utilizes orthogonal autoencoders and variational dropout techniques to enhance word embedding. The orthogonality constraint encourages more diversity in the latent space, and variational dropout makes the embedding more robust to overfitting. Empirical evaluation on a range of downstream NLP tasks shows that our proposed method effectively improves the quality of pre-trained word embeddings and pre-trained language models. In addition, We introduce a technique to enhance the accuracy of word embeddings by incorporating external knowledge in the form of WordNet pairs to improve the accuracy of word embedding on several NLP tasks.","abstract_has_math":false,"creators":["Albujasim, Zainab Majeed"],"institution":"Carleton University","degree_name":"Doctor of Philosophy (Ph.D.)","degree_level":"Doctoral","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024","date_published":"2024","updated_at":"2026-07-24T01:34:30Z","subjects":[],"languages":["en"],"rights":["Copyright © 2024 the author(s). Theses may be used for non-commercial research, educational, or related academic purposes only. Such uses include personal study, distribution to students, research and scholarship. Theses may only be shared by linking to the Carleton University Institutional Repository and no part may be copied without proper attribution to the author; no part may be used for commercial purposes directly or indirectly via a for-profit platform; no adaptation or derivative works are permitted without consent from the copyright owner."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.doi","label":"DOI","values":["10.22215/etd/2024-15911"],"render_values":[{"text":"10.22215/etd/2024-15911","href":"https://doi.org/10.22215/etd/2024-15911","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/20.500.14718/41318","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Albujasim, Zainab Majeed"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-04-08T20:16:51Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-04-08T20:16:51Z"]},{"key":"dc:date.issued","label":"Date","values":["2024"]},{"key":"dc:publisher","label":"Institution","values":["Carleton University"]},{"key":"dc:type","label":"Dc Type","values":["thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (Ph.D.)"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright © 2024 the author(s). Theses may be used for non-commercial research, educational, or related academic purposes only. Such uses include personal study, distribution to students, research and scholarship. Theses may only be shared by linking to the Carleton University Institutional Repository and no part may be copied without proper attribution to the author; no part may be used for commercial purposes directly or indirectly via a for-profit platform; no adaptation or derivative works are permitted without consent from the copyright owner."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["10.22215/etd/2024-15911"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/20.500.14718/41318"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Word embedding has been a significant breakthrough in natural language processing (NLP). Although word representation has improved remarkably and resulted in better performance in downstream NLP applications, interpretability of word embeddings remains a challenge. Post-processing techniques have been developed to improve the quality of word embeddings by fine-tuning real-valued vectors and eliminating artifacts that arise due to flaws in the training corpus. In this thesis, we propose a simple method to improve the quality of word embeddings and reveal their hidden structure using a post-processing technique. We deploy co-clustering techniques to detect sub-matrices between word meaning and specific dimensions. Furthermore, pre-trained word embeddings based on neural networks often suffer from representation degeneration issues, especially when trained on large datasets. We propose a novel method Orthogonal Auto Encoder with Variational Dropout (OAEVD), which utilizes orthogonal autoencoders and variational dropout techniques to enhance word embedding. The orthogonality constraint encourages more diversity in the latent space, and variational dropout makes the embedding more robust to overfitting. Empirical evaluation on a range of downstream NLP tasks shows that our proposed method effectively improves the quality of pre-trained word embeddings and pre-trained language models. In addition, We introduce a technique to enhance the accuracy of word embeddings by incorporating external knowledge in the form of WordNet pairs to improve the accuracy of word embedding on several NLP tasks."]},{"key":"dc:title","label":"Title","values":["Post-processing Techniques for Word Embedding"]}]}],"canonical_facts":{"dc:creator":["Albujasim, Zainab Majeed"],"dc:date.accessioned":["2025-04-08T20:16:51Z"],"dc:date.available":["2025-04-08T20:16:51Z"],"dc:date.issued":["2024"],"dc:description.abstract":["Word embedding has been a significant breakthrough in natural language processing (NLP). Although word representation has improved remarkably and resulted in better performance in downstream NLP applications, interpretability of word embeddings remains a challenge. Post-processing techniques have been developed to improve the quality of word embeddings by fine-tuning real-valued vectors and eliminating artifacts that arise due to flaws in the training corpus. In this thesis, we propose a simple method to improve the quality of word embeddings and reveal their hidden structure using a post-processing technique. We deploy co-clustering techniques to detect sub-matrices between word meaning and specific dimensions. Furthermore, pre-trained word embeddings based on neural networks often suffer from representation degeneration issues, especially when trained on large datasets. We propose a novel method Orthogonal Auto Encoder with Variational Dropout (OAEVD), which utilizes orthogonal autoencoders and variational dropout techniques to enhance word embedding. The orthogonality constraint encourages more diversity in the latent space, and variational dropout makes the embedding more robust to overfitting. Empirical evaluation on a range of downstream NLP tasks shows that our proposed method effectively improves the quality of pre-trained word embeddings and pre-trained language models. In addition, We introduce a technique to enhance the accuracy of word embeddings by incorporating external knowledge in the form of WordNet pairs to improve the accuracy of word embedding on several NLP tasks."],"dc:identifier.doi":["10.22215/etd/2024-15911"],"dc:identifier.uri":["https://hdl.handle.net/20.500.14718/41318"],"dc:language.iso":["en"],"dc:publisher":["Carleton University"],"dc:rights":["Copyright © 2024 the author(s). Theses may be used for non-commercial research, educational, or related academic purposes only. Such uses include personal study, distribution to students, research and scholarship. Theses may only be shared by linking to the Carleton University Institutional Repository and no part may be copied without proper attribution to the author; no part may be used for commercial purposes directly or indirectly via a for-profit platform; no adaptation or derivative works are permitted without consent from the copyright owner."],"dc:title":["Post-processing Techniques for Word Embedding"],"dc:type":["thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy (Ph.D.)"]},"updated_at":"2026-07-24T01:34:30Z"}