{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/347873"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/347873","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"On the Evaluation and Modelling of Context-sensitive Lexical Semantics","abstract":"A word can change its meaning in different contexts. The evaluation and the modelling of such contextual effect on lexical meaning are pivotal to natural language understanding. This thesis sets out to answer the following two-fold research question: (1) how can we design a reliable evaluation framework that accurately reflects the challenges in contextual lexical semantics? And (2) how can we improve contextual word representations in a data-efficient manner? To address the first question, the thesis starts with systematic analysis on the existing benchmark datasets and stresses the importance to assess the complex word-context interaction. I propose that the evaluation of crosslingual correspondence can effectively assess this interaction. Specifically, I introduced two datasets, i.e. BTSR (Bilingual Token-level Sense Retrieval) and AM2ICO (Adversarial and Multilingual Meaning in Context), to evaluate crosslingual word-in-context correspondences that require the accurate crosslingual modelling of both the target words and their context. BTSR and AM2ICO complement each other in task formulations and language/word coverage. Finally, I show that evaluating and modelling crosslingual word-in-context representations have direct benefit for downstream applications. In a case study, I apply the BTSR task formulation to improve machine translation for rare senses. While improving the current state-of-the-art contextual models typically involves labelled data which is not easy to obtain especially for low-resource languages, the second research question of this thesis aims to elicit better word in context representations from state-ofthe-art contextual models without resorting to labelled data. As an outcome, I designed two novel unsupervised methods (MIRRORWIC and STATICTRANSFORM) that improve either within-word or inter-word contextualisation of the pretrained contextual models both monolingually and crosslingually. In sum, the thesis contributes to the field of computational lexical semantics by providing challenging and accurate evaluation frameworks and efficient modelling techniques that avoid the need for labelled data.","abstract_html":"A word can change its meaning in different contexts. The evaluation and the modelling of such contextual effect on lexical meaning are pivotal to natural language understanding. This thesis sets out to answer the following two-fold research question: (1) how can we design a reliable evaluation framework that accurately reflects the challenges in contextual lexical semantics? And (2) how can we improve contextual word representations in a data-efficient manner? To address the first question, the thesis starts with systematic analysis on the existing benchmark datasets and stresses the importance to assess the complex word-context interaction. I propose that the evaluation of crosslingual correspondence can effectively assess this interaction. Specifically, I introduced two datasets, i.e. BTSR (Bilingual Token-level Sense Retrieval) and AM2ICO (Adversarial and Multilingual Meaning in Context), to evaluate crosslingual word-in-context correspondences that require the accurate crosslingual modelling of both the target words and their context. BTSR and AM2ICO complement each other in task formulations and language/word coverage. Finally, I show that evaluating and modelling crosslingual word-in-context representations have direct benefit for downstream applications. In a case study, I apply the BTSR task formulation to improve machine translation for rare senses. While improving the current state-of-the-art contextual models typically involves labelled data which is not easy to obtain especially for low-resource languages, the second research question of this thesis aims to elicit better word in context representations from state-ofthe-art contextual models without resorting to labelled data. As an outcome, I designed two novel unsupervised methods (MIRRORWIC and STATICTRANSFORM) that improve either within-word or inter-word contextualisation of the pretrained contextual models both monolingually and crosslingually. In sum, the thesis contributes to the field of computational lexical semantics by providing challenging and accurate evaluation frameworks and efficient modelling techniques that avoid the need for labelled data.","abstract_has_math":false,"creators":["Liu, Qianchu"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Korhonen, Anna-Leena"],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-31","date_published":"2022-01-31","updated_at":"2026-07-22T22:24:11Z","subjects":["Natural Language Processing"],"languages":["eng"],"rights":[],"rights_urls":["https://www.rioxx.net/licenses/all-rights-reserved/"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.95289","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Korhonen, Anna-Leena"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Peterhouse Research Studentship"]},{"key":"dc:creator","label":"Author","values":["Liu, Qianchu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2022-01-31"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/347873"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Natural Language Processing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.rioxx.net/licenses/all-rights-reserved/"]},{"key":"dc:rights.embargotype","label":"Dc Rights Embargotype","values":["controlled.access"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["10.17863/CAM.95289"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/7b5ca109-48c1-483d-a923-be46218ddde4/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["A word can change its meaning in different contexts. The evaluation and the modelling of such contextual effect on lexical meaning are pivotal to natural language understanding. This thesis sets out to answer the following two-fold research question: (1) how can we design a reliable evaluation framework that accurately reflects the challenges in contextual lexical semantics? And (2) how can we improve contextual word representations in a data-efficient manner? To address the first question, the thesis starts with systematic analysis on the existing benchmark datasets and stresses the importance to assess the complex word-context interaction. I propose that the evaluation of crosslingual correspondence can effectively assess this interaction. Specifically, I introduced two datasets, i.e. BTSR (Bilingual Token-level Sense Retrieval) and AM2ICO (Adversarial and Multilingual Meaning in Context), to evaluate crosslingual word-in-context correspondences that require the accurate crosslingual modelling of both the target words and their context. BTSR and AM2ICO complement each other in task formulations and language/word coverage. Finally, I show that evaluating and modelling crosslingual word-in-context representations have direct benefit for downstream applications. In a case study, I apply the BTSR task formulation to improve machine translation for rare senses. While improving the current state-of-the-art contextual models typically involves labelled data which is not easy to obtain especially for low-resource languages, the second research question of this thesis aims to elicit better word in context representations from state-ofthe-art contextual models without resorting to labelled data. As an outcome, I designed two novel unsupervised methods (MIRRORWIC and STATICTRANSFORM) that improve either within-word or inter-word contextualisation of the pretrained contextual models both monolingually and crosslingually. In sum, the thesis contributes to the field of computational lexical semantics by providing challenging and accurate evaluation frameworks and efficient modelling techniques that avoid the need for labelled data."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["a5dfcf138f3d50d90c5a04be836e92ab"]},{"key":"dc:title","label":"Title","values":["On the Evaluation and Modelling of Context-sensitive Lexical Semantics"]}]}],"canonical_facts":{"dc:contributor.advisor":["Korhonen, Anna-Leena"],"dc:contributor.sponsor":["Peterhouse Research Studentship"],"dc:creator":["Liu, Qianchu"],"dc:date.issued":["2022-01-31"],"dc:description.abstract":["A word can change its meaning in different contexts. The evaluation and the modelling of such contextual effect on lexical meaning are pivotal to natural language understanding. This thesis sets out to answer the following two-fold research question: (1) how can we design a reliable evaluation framework that accurately reflects the challenges in contextual lexical semantics? And (2) how can we improve contextual word representations in a data-efficient manner? To address the first question, the thesis starts with systematic analysis on the existing benchmark datasets and stresses the importance to assess the complex word-context interaction. I propose that the evaluation of crosslingual correspondence can effectively assess this interaction. Specifically, I introduced two datasets, i.e. BTSR (Bilingual Token-level Sense Retrieval) and AM2ICO (Adversarial and Multilingual Meaning in Context), to evaluate crosslingual word-in-context correspondences that require the accurate crosslingual modelling of both the target words and their context. BTSR and AM2ICO complement each other in task formulations and language/word coverage. Finally, I show that evaluating and modelling crosslingual word-in-context representations have direct benefit for downstream applications. In a case study, I apply the BTSR task formulation to improve machine translation for rare senses. While improving the current state-of-the-art contextual models typically involves labelled data which is not easy to obtain especially for low-resource languages, the second research question of this thesis aims to elicit better word in context representations from state-ofthe-art contextual models without resorting to labelled data. As an outcome, I designed two novel unsupervised methods (MIRRORWIC and STATICTRANSFORM) that improve either within-word or inter-word contextualisation of the pretrained contextual models both monolingually and crosslingually. In sum, the thesis contributes to the field of computational lexical semantics by providing challenging and accurate evaluation frameworks and efficient modelling techniques that avoid the need for labelled data."],"dc:format.checksum.md5":["a5dfcf138f3d50d90c5a04be836e92ab"],"dc:identifier.doi":["10.17863/CAM.95289"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/7b5ca109-48c1-483d-a923-be46218ddde4/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/347873"],"dc:rights":["https://www.rioxx.net/licenses/all-rights-reserved/"],"dc:rights.embargotype":["controlled.access"],"dc:subject":["Natural Language Processing"],"dc:title":["On the Evaluation and Modelling of Context-sensitive Lexical Semantics"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:11Z"}