{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/31920"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/31920","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"A study of the impact of global statistics in distributed information retrieval","abstract":"Today’s information retrieval systems have to deal with very large data collections and take a distributed approach to achieve scalable retrieval performance. The most widely used approach, called document-partitioning, is to partition the data among multiple search-nodes, which then index their sub-collection independently and are responsible for scoring documents present in their index, against queries. Most of the famous document scoring functions depend on various global (collection-wide) statistics such as document frequency of terms. However, as search-nodes don’t have access to global-statistics and rely on local (sub-collection-wide) statistics for the purpose of scoring, document-partitioning can result in a degraded retrieval performance. In this thesis, we study the impact of the lack of global-statistics on the retrieval performance of a distributed information retrieval (DIR) system. Our experiments show that the performance, as indicated by multiple measures, degrades as the number of search-nodes are increased. We thus conclude that global-statistics are essential to the retrieval performance in a distributed setup. Finally, we present a novel scheme for lazy and adaptive dissemination of global-statistics in a document-partitioned DIR system.","abstract_html":"Today’s information retrieval systems have to deal with very large data collections and take a distributed approach to achieve scalable retrieval performance. The most widely used approach, called document-partitioning, is to partition the data among multiple search-nodes, which then index their sub-collection independently and are responsible for scoring documents present in their index, against queries. Most of the famous document scoring functions depend on various global (collection-wide) statistics such as document frequency of terms. However, as search-nodes don’t have access to global-statistics and rely on local (sub-collection-wide) statistics for the purpose of scoring, document-partitioning can result in a degraded retrieval performance. In this thesis, we study the impact of the lack of global-statistics on the retrieval performance of a distributed information retrieval (DIR) system. Our experiments show that the performance, as indicated by multiple measures, degrades as the number of search-nodes are increased. We thus conclude that global-statistics are essential to the retrieval performance in a distributed setup. Finally, we present a novel scheme for lazy and adaptive dissemination of global-statistics in a document-partitioned DIR system.","abstract_has_math":false,"creators":["Sehrawat, Nipun"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, ChengXiang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-06-27T21:19:22Z","date_published":"2012-06-27T21:19:22Z","updated_at":"2026-07-22T22:25:30Z","subjects":["Distributed Information Retrieval","Global Statistics","Retrieval Performance"],"languages":["en"],"rights":["Copyright 2012 Nipun Sehrawat"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/31920","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, ChengXiang"]},{"key":"dc:creator","label":"Author","values":["Sehrawat, Nipun"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-06-27T21:19:22Z","2014-06-28T10:00:18Z","2012-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Distributed Information Retrieval","Global Statistics","Retrieval Performance"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2012 Nipun Sehrawat"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/31920"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Today’s information retrieval systems have to deal with very large data collections and take a distributed approach to achieve scalable retrieval performance. The most widely used approach, called document-partitioning, is to partition the data among multiple search-nodes, which then index their sub-collection independently and are responsible for scoring documents present in their index, against queries. Most of the famous document scoring functions depend on various global (collection-wide) statistics such as document frequency of terms. However, as search-nodes don’t have access to global-statistics and rely on local (sub-collection-wide) statistics for the purpose of scoring, document-partitioning can result in a degraded retrieval performance. In this thesis, we study the impact of the lack of global-statistics on the retrieval performance of a distributed information retrieval (DIR) system. Our experiments show that the performance, as indicated by multiple measures, degrades as the number of search-nodes are increased. We thus conclude that global-statistics are essential to the retrieval performance in a distributed setup. Finally, we present a novel scheme for lazy and adaptive dissemination of global-statistics in a document-partitioned DIR system.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-04-21T16:54:47Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Sehrawat_Nipun.pdf: 825551 bytes, checksum: 1e3f371affc85a7d9928e18dda11f1be (MD5)","Made available in DSpace on 2012-06-27T21:19:22Z (GMT). No. of bitstreams: 2 Sehrawat_Nipun.pdf: 825547 bytes, checksum: 3a166b054aae15ccce207de91a02dc1c (MD5) license.txt: 4064 bytes, checksum: 22fbe4e60e40b7e09a9c41428eaed23e (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by William Ingram (wingram2@illinois.edu) on 2012-06-27T21:24:31Z Item is restricted until 2014-06-27T21:24:27Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:18Z Item was in collections: Dissertations and Theses - Computer Science (ID: 587) Graduate Theses and Dissertations at Illinois (ID: 204) No. of bitstreams: 2 Sehrawat_Nipun.pdf: 825547 bytes, checksum: 3a166b054aae15ccce207de91a02dc1c (MD5) license.txt: 4064 bytes, checksum: 22fbe4e60e40b7e09a9c41428eaed23e (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:18Z"]},{"key":"dc:title","label":"Title","values":["A study of the impact of global statistics in distributed information retrieval"]}]}],"canonical_facts":{"dc:contributor":["Zhai, ChengXiang"],"dc:creator":["Sehrawat, Nipun"],"dc:date":["2012-06-27T21:19:22Z","2014-06-28T10:00:18Z","2012-05"],"dc:description":["Today’s information retrieval systems have to deal with very large data collections and take a distributed approach to achieve scalable retrieval performance. The most widely used approach, called document-partitioning, is to partition the data among multiple search-nodes, which then index their sub-collection independently and are responsible for scoring documents present in their index, against queries. Most of the famous document scoring functions depend on various global (collection-wide) statistics such as document frequency of terms. However, as search-nodes don’t have access to global-statistics and rely on local (sub-collection-wide) statistics for the purpose of scoring, document-partitioning can result in a degraded retrieval performance. In this thesis, we study the impact of the lack of global-statistics on the retrieval performance of a distributed information retrieval (DIR) system. Our experiments show that the performance, as indicated by multiple measures, degrades as the number of search-nodes are increased. We thus conclude that global-statistics are essential to the retrieval performance in a distributed setup. Finally, we present a novel scheme for lazy and adaptive dissemination of global-statistics in a document-partitioned DIR system.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-04-21T16:54:47Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Sehrawat_Nipun.pdf: 825551 bytes, checksum: 1e3f371affc85a7d9928e18dda11f1be (MD5)","Made available in DSpace on 2012-06-27T21:19:22Z (GMT). No. of bitstreams: 2 Sehrawat_Nipun.pdf: 825547 bytes, checksum: 3a166b054aae15ccce207de91a02dc1c (MD5) license.txt: 4064 bytes, checksum: 22fbe4e60e40b7e09a9c41428eaed23e (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by William Ingram (wingram2@illinois.edu) on 2012-06-27T21:24:31Z Item is restricted until 2014-06-27T21:24:27Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:18Z Item was in collections: Dissertations and Theses - Computer Science (ID: 587) Graduate Theses and Dissertations at Illinois (ID: 204) No. of bitstreams: 2 Sehrawat_Nipun.pdf: 825547 bytes, checksum: 3a166b054aae15ccce207de91a02dc1c (MD5) license.txt: 4064 bytes, checksum: 22fbe4e60e40b7e09a9c41428eaed23e (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:18Z"],"dc:identifier":["http://hdl.handle.net/2142/31920"],"dc:language":["en"],"dc:rights":["Copyright 2012 Nipun Sehrawat"],"dc:subject":["Distributed Information Retrieval","Global Statistics","Retrieval Performance"],"dc:title":["A study of the impact of global statistics in distributed information retrieval"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:30Z"}