{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/116283"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/116283","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Synthetic pre-training for robustness in information retrieval","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_has_math":false,"creators":["Gangi Reddy, Revanth"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08","date_published":"2022-08","updated_at":"2026-07-22T22:24:56Z","subjects":["Neural Information Retrieval","Data Augmentation","Weak supervision"],"languages":["en","eng"],"rights":["Copyright 2022 Revanth Gangi Reddy"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/116283","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng"]},{"key":"dc:creator","label":"Author","values":["Gangi Reddy, Revanth"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-08","2022-07-21"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Neural Information Retrieval","Data Augmentation","Weak supervision"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Revanth Gangi Reddy"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/116283"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Revanth Gangi Reddy, accepted the attached license on 2022-07-21 at 13:30.","The student, Revanth Gangi Reddy, submitted this Thesis for approval on 2022-07-21 at 13:37.","This Thesis was approved for publication on 2022-07-21 at 15:14.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18415 on 2022-11-15 at 18:22:00","Research on neural information retrieval has so far been focused primarily on standard supervised learning settings, where it outperforms traditional term matching baselines. Many practical use cases of such models, however, may involve previously unseen target domains. In this thesis, we first improve the out-of-domain generalization of Dense Passage Retrieval (DPR)—a popular choice for neural information retrieval (IR)—through synthetic data augmentation only in the source domain. We empirically show that pre-training DPR with additional synthetic data in its source domain (Wikipedia), which we generate using a fine-tuned sequence-to-sequence generator, can be a low-cost yet effective first step towards its generalization. Across five different test sets, our augmented model shows more robust performance than DPR in both in-domain and zero-shot out-of-domain evaluation. We then show that supervised neural IR models are prone to learning sparse attention patterns over passage tokens, which can result in key phrases including named entities receiving low attention weights, eventually leading to model under-performance. Using a novel targeted synthetic data generation method that identifies poorly attended entities and conditions the generation episodes on those, we teach neural IR to attend more uniformly and robustly to all entities in a given passage. On two public IR benchmarks, we empirically show that the proposed method helps improve both the model’s attention patterns and retrieval performance, including in zero-shot settings."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Synthetic pre-training for robustness in information retrieval"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng"],"dc:creator":["Gangi Reddy, Revanth"],"dc:date":["2022-08","2022-07-21"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Revanth Gangi Reddy, accepted the attached license on 2022-07-21 at 13:30.","The student, Revanth Gangi Reddy, submitted this Thesis for approval on 2022-07-21 at 13:37.","This Thesis was approved for publication on 2022-07-21 at 15:14.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18415 on 2022-11-15 at 18:22:00","Research on neural information retrieval has so far been focused primarily on standard supervised learning settings, where it outperforms traditional term matching baselines. Many practical use cases of such models, however, may involve previously unseen target domains. In this thesis, we first improve the out-of-domain generalization of Dense Passage Retrieval (DPR)—a popular choice for neural information retrieval (IR)—through synthetic data augmentation only in the source domain. We empirically show that pre-training DPR with additional synthetic data in its source domain (Wikipedia), which we generate using a fine-tuned sequence-to-sequence generator, can be a low-cost yet effective first step towards its generalization. Across five different test sets, our augmented model shows more robust performance than DPR in both in-domain and zero-shot out-of-domain evaluation. We then show that supervised neural IR models are prone to learning sparse attention patterns over passage tokens, which can result in key phrases including named entities receiving low attention weights, eventually leading to model under-performance. Using a novel targeted synthetic data generation method that identifies poorly attended entities and conditions the generation episodes on those, we teach neural IR to attend more uniformly and robustly to all entities in a given passage. On two public IR benchmarks, we empirically show that the proposed method helps improve both the model’s attention patterns and retrieval performance, including in zero-shot settings."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/116283"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Revanth Gangi Reddy"],"dc:subject":["Neural Information Retrieval","Data Augmentation","Weak supervision"],"dc:title":["Synthetic pre-training for robustness in information retrieval"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:56Z"}