Abstract
dc:description.abstractWord embedding has been a significant breakthrough in natural language processing (NLP). Although word representation has improved remarkably and resulted in better performance in downstream NLP applications, interpretability of word embeddings remains a challenge. Post-processing techniques have been developed to improve the quality of word embeddings by fine-tuning real-valued vectors and eliminating artifacts that arise due to flaws in the training corpus. In this thesis, we propose a simple method to improve the quality of word embeddings and reveal their hidden structure using a post-processing technique. We deploy co-clustering techniques to detect sub-matrices between word meaning and specific dimensions. Furthermore, pre-trained word embeddings based on neural networks often suffer from representation degeneration issues, especially when trained on large datasets. We propose a novel method Orthogonal Auto Encoder with Variational Dropout (OAEVD), which utilizes orthogonal autoencoders and variational dropout techniques to enhance word embedding. The orthogonality constraint encourages more diversity in the latent space, and variational dropout makes the embedding more robust to overfitting. Empirical evaluation on a range of downstream NLP tasks shows that our proposed method effectively improves the quality of pre-trained word embeddings and pre-trained language models. In addition, We introduce a technique to enhance the accuracy of word embeddings by incorporating external knowledge in the form of WordNet pairs to improve the accuracy of word embedding on several NLP tasks.
Degree
thesis:*- Name thesis:degree_name
- Doctor of Philosophy (Ph.D.)
- Level thesis:degree_level
- Doctoral
- Discipline thesis:degree_discipline
- Computer Science
- Grantor dc:publisher
- Carleton University
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Albujasim, Zainab Majeed
Rights
dc:rights- Statement dc:rights
-
- Copyright © 2024 the author(s). Theses may be used for non-commercial research, educational, or related academic purposes only. Such uses include personal study, distribution to students, research and scholarship. Theses may only be shared by linking to the Carleton University Institutional Repository and no part may be copied without proper attribution to the author; no part may be used for commercial purposes directly or indirectly via a for-profit platform; no adaptation or derivative works are permitted without consent from the copyright owner.
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- OAI identifier oai:identifier
- oai:carleton.scholaris.ca:20.500.14718/41318