{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/49645"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/49645","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Entity finder: A system for entity web page retrieval using pseudo-relevance feedback","abstract":"Collecting all the online information about a particular entity (e.g., a person or a product) is a task commonly needed in many applications. In many cases, there often already exists a database with limited information about interesting entities, but those databases usually suffer from incompleteness and out-of-date problems. But with the increasing amount of information available on the World Wide Web, crawling and searching the web may be an attractive technological approach that can help update a database to make it more complete and up to date. In this thesis, we propose a retrieval system that crawls and searches the web in order to complete and update the information about an entity already existent in a database maintained by an organization. Taking the information stored in the database as input, this system can crawl the web and retrieve the web pages mentioning the entities in the database. We study several approaches to solving this special retrieval problem, and propose a novel pseudo-relevance feedback approach to improve the retrieval accuracy. We evaluate our system over a dataset containing 112 alumni in the College of Engineering of the University of Illinois, and show that our system can effectively retrieve relevant pages of alumni on the web and that the novel pseudo-relevance feedback method outperforms a simple baseline approach.","abstract_html":"Collecting all the online information about a particular entity (e.g., a person or a product) is a task commonly needed in many applications. In many cases, there often already exists a database with limited information about interesting entities, but those databases usually suffer from incompleteness and out-of-date problems. But with the increasing amount of information available on the World Wide Web, crawling and searching the web may be an attractive technological approach that can help update a database to make it more complete and up to date. In this thesis, we propose a retrieval system that crawls and searches the web in order to complete and update the information about an entity already existent in a database maintained by an organization. Taking the information stored in the database as input, this system can crawl the web and retrieve the web pages mentioning the entities in the database. We study several approaches to solving this special retrieval problem, and propose a novel pseudo-relevance feedback approach to improve the retrieval accuracy. We evaluate our system over a dataset containing 112 alumni in the College of Engineering of the University of Illinois, and show that our system can effectively retrieve relevant pages of alumni on the web and that the novel pseudo-relevance feedback method outperforms a simple baseline approach.","abstract_has_math":false,"creators":["Wang, Rui"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, ChengXiang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-05-30T16:53:53Z","date_published":"2014-05-30T16:53:53Z","updated_at":"2026-07-22T22:25:38Z","subjects":["Information Retrieval","Web Search","Relevance Feedback"],"languages":["en"],"rights":["Copyright 2014 Rui Wang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/49645","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, ChengXiang"]},{"key":"dc:creator","label":"Author","values":["Wang, Rui"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014-05-30T16:53:53Z","2014-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Information Retrieval","Web Search","Relevance Feedback"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2014 Rui Wang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/49645"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Collecting all the online information about a particular entity (e.g., a person or a product) is a task commonly needed in many applications. In many cases, there often already exists a database with limited information about interesting entities, but those databases usually suffer from incompleteness and out-of-date problems. But with the increasing amount of information available on the World Wide Web, crawling and searching the web may be an attractive technological approach that can help update a database to make it more complete and up to date. In this thesis, we propose a retrieval system that crawls and searches the web in order to complete and update the information about an entity already existent in a database maintained by an organization. Taking the information stored in the database as input, this system can crawl the web and retrieve the web pages mentioning the entities in the database. We study several approaches to solving this special retrieval problem, and propose a novel pseudo-relevance feedback approach to improve the retrieval accuracy. We evaluate our system over a dataset containing 112 alumni in the College of Engineering of the University of Illinois, and show that our system can effectively retrieve relevant pages of alumni on the web and that the novel pseudo-relevance feedback method outperforms a simple baseline approach.","Item withdrawn by Laura Spradlin (lspradl2@illinois.edu) on 2014-04-30T15:39:21Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Wang_Rui.docx: 118694 bytes, checksum: 8d8c978c175a2c433299df15d93d836d (MD5) Wang_Rui.pdf: 1722326 bytes, checksum: af1f2743797eb268aa65af9cb7a19838 (MD5)","Made available in DSpace on 2014-05-30T16:53:53Z (GMT). No. of bitstreams: 3 Rui_Wang.pdf: 1722326 bytes, checksum: af1f2743797eb268aa65af9cb7a19838 (MD5) Wang_Rui.docx: 118694 bytes, checksum: 8d8c978c175a2c433299df15d93d836d (MD5) license.txt: 4058 bytes, checksum: e8213f365da8d43ba75763ea9f93b443 (MD5)"]},{"key":"dc:title","label":"Title","values":["Entity finder: A system for entity web page retrieval using pseudo-relevance feedback"]}]}],"canonical_facts":{"dc:contributor":["Zhai, ChengXiang"],"dc:creator":["Wang, Rui"],"dc:date":["2014-05-30T16:53:53Z","2014-05"],"dc:description":["Collecting all the online information about a particular entity (e.g., a person or a product) is a task commonly needed in many applications. In many cases, there often already exists a database with limited information about interesting entities, but those databases usually suffer from incompleteness and out-of-date problems. But with the increasing amount of information available on the World Wide Web, crawling and searching the web may be an attractive technological approach that can help update a database to make it more complete and up to date. In this thesis, we propose a retrieval system that crawls and searches the web in order to complete and update the information about an entity already existent in a database maintained by an organization. Taking the information stored in the database as input, this system can crawl the web and retrieve the web pages mentioning the entities in the database. We study several approaches to solving this special retrieval problem, and propose a novel pseudo-relevance feedback approach to improve the retrieval accuracy. We evaluate our system over a dataset containing 112 alumni in the College of Engineering of the University of Illinois, and show that our system can effectively retrieve relevant pages of alumni on the web and that the novel pseudo-relevance feedback method outperforms a simple baseline approach.","Item withdrawn by Laura Spradlin (lspradl2@illinois.edu) on 2014-04-30T15:39:21Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Wang_Rui.docx: 118694 bytes, checksum: 8d8c978c175a2c433299df15d93d836d (MD5) Wang_Rui.pdf: 1722326 bytes, checksum: af1f2743797eb268aa65af9cb7a19838 (MD5)","Made available in DSpace on 2014-05-30T16:53:53Z (GMT). No. of bitstreams: 3 Rui_Wang.pdf: 1722326 bytes, checksum: af1f2743797eb268aa65af9cb7a19838 (MD5) Wang_Rui.docx: 118694 bytes, checksum: 8d8c978c175a2c433299df15d93d836d (MD5) license.txt: 4058 bytes, checksum: e8213f365da8d43ba75763ea9f93b443 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/49645"],"dc:language":["en"],"dc:rights":["Copyright 2014 Rui Wang"],"dc:subject":["Information Retrieval","Web Search","Relevance Feedback"],"dc:title":["Entity finder: A system for entity web page retrieval using pseudo-relevance feedback"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:38Z"}