Università degli studi di Trento
Models and Application of Question Retrieval for Natural Language Processing
Abstract
dc:descriptionThis thesis investigates the role of question understanding in Question Answering systems, developing methods that exploit question semantic equivalence at progressively larger scales: from individual question pairs, through equivalence clusters, to entire datasets. The first part addresses question retrieval at scale. We introduce QUADRo, a retrieval framework operating over millions of question-answer pairs, and the Question Ranking Corpus (QRC), a large-scale resource with answer-aware annotations and challenging hard negatives. We demonstrate that incorporating answers during retrieval substantially improves accuracy, as answers serve as a semantic bridge between questions that share little lexical overlap but seek the same information. To reduce annotation costs, we develop Question Ranking Pre-training (QRP), a self-supervised method that learns question equivalence patterns without labeled data, achieving significant improvements while reducing model variance by over 50\%. The second part extends pairwise equivalence to question clusters. We analyze coherence in Large Language Models, finding that a substantial portion of question clusters exhibit incoherent behavior: models answer some phrasings correctly while failing on semantically equivalent alternatives. This reveals that understanding failures, not just knowledge gaps, limit LLM performance. We introduce Question-Augmented Generation (q-RAG), which supplements prompts with retrieved similar questions, improving accuracy by up to 9 percentage points and coherence by up to 28 points. We further show that q-RAG's benefits can be distilled into model parameters through Direct Preference Optimization (DPO) and Supervised Fine-Tuning, producing standalone models with improved coherence that surpass the inference-time approach. For retrieval systems, we apply clusters to train models for consistency: the Coherence Ranking Loss improves ranking coherence by up to 30\% while simultaneously improving relevance. The third part lifts equivalence to the dataset level. We introduce dataset declassification, a framework that replaces proprietary questions with semantically equivalent public alternatives, enabling dataset sharing without exposing sensitive content. Models trained on fully declassified data match baseline performance (WikiQA $\Delta \approx 0$, TrecQA $|\Delta| \leq 1.2$ points), and test set declassification preserves evaluation validity when high-quality mappings exist ($|\Delta| \leq 2$ on standard benchmarks), enabling the release of ``shadow benchmarks'' for evaluation integrity. We identify boundary conditions through experiments on adversarially-constructed benchmarks. Together, these contributions show that question semantic equivalence, systematically exploited at multiple scales, enables substantial improvements to QA system accuracy, consistency, and evaluation integrity.
Degree
thesis:*- Grantor dc:publisher
- Università degli studi di Trento
- Year dc:date
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Campese, Stefano
- Contributors dc:contributor
-
- Moschitti, Alessandro
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- info:eu-repo/semantics/openAccess
- license:Tutti i diritti riservati (All rights reserved)
- license uri:iris.PRI01
- Language dc:language
- eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/11572/483952
- OAI identifier oai:identifier
- oai:iris.unitn.it:11572/483952