{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/50493"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/50493","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Integrating multiple conflicting sources by truth discovery and source quality estimation","abstract":"Multiple descriptions about the same entity from different sources will inevitably result in data or information inconsistency. Among conflicting pieces of information, which one is the most trustworthy? How to detect the fraudulence of a rumor? Obviously, it is unrealistic to curate and validate the trustworthiness of every piece of information because of the high cost of human labeling and lack of experts. To find the truth of each entity, much research work has shown that considering the quality of information providers can improve the performance of data integration. Due to different quality of data sources, it is hard to find a general solution that works for every case. Therefore, we start from a general setting of truth analysis at first and narrow down to two basic problems in data integration. We first propose a general framework to deal with numerical data with flexibility of defining loss function. Source quality is represented by a vector to model the source credibility in different error interval. Then we propose a new method called No Truth Truth Model(NTTM) to deal with truth existence problem in low-quality data. Preliminary experiments on real stock data and slot filling data show promising results.","abstract_html":"Multiple descriptions about the same entity from different sources will inevitably result in data or information inconsistency. Among conflicting pieces of information, which one is the most trustworthy? How to detect the fraudulence of a rumor? Obviously, it is unrealistic to curate and validate the trustworthiness of every piece of information because of the high cost of human labeling and lack of experts. To find the truth of each entity, much research work has shown that considering the quality of information providers can improve the performance of data integration. Due to different quality of data sources, it is hard to find a general solution that works for every case. Therefore, we start from a general setting of truth analysis at first and narrow down to two basic problems in data integration. We first propose a general framework to deal with numerical data with flexibility of defining loss function. Source quality is represented by a vector to model the source credibility in different error interval. Then we propose a new method called No Truth Truth Model(NTTM) to deal with truth existence problem in low-quality data. Preliminary experiments on real stock data and slot filling data show promising results.","abstract_has_math":false,"creators":["Zhi, Shi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-09-16T17:18:00Z","date_published":"2014-09-16T17:18:00Z","updated_at":"2026-07-22T22:25:40Z","subjects":["Truth Discovery","Data Integration","Data Quality"],"languages":["en"],"rights":["Copyright 2014 Shi Zhi"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/50493","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei"]},{"key":"dc:creator","label":"Author","values":["Zhi, Shi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014-09-16T17:18:00Z","2016-09-22T20:59:16Z","2014-08","2014-09-16"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Truth Discovery","Data Integration","Data Quality"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2014 Shi Zhi"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/50493"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Multiple descriptions about the same entity from different sources will inevitably result in data or information inconsistency. Among conflicting pieces of information, which one is the most trustworthy? How to detect the fraudulence of a rumor? Obviously, it is unrealistic to curate and validate the trustworthiness of every piece of information because of the high cost of human labeling and lack of experts. To find the truth of each entity, much research work has shown that considering the quality of information providers can improve the performance of data integration. Due to different quality of data sources, it is hard to find a general solution that works for every case. Therefore, we start from a general setting of truth analysis at first and narrow down to two basic problems in data integration. We first propose a general framework to deal with numerical data with flexibility of defining loss function. Source quality is represented by a vector to model the source credibility in different error interval. Then we propose a new method called No Truth Truth Model(NTTM) to deal with truth existence problem in low-quality data. Preliminary experiments on real stock data and slot filling data show promising results.","Item withdrawn by Laura Spradlin (lspradl2@illinois.edu) on 2014-07-24T14:56:06Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Zhi_Shi.pdf: 1401640 bytes, checksum: f7b1d363045a4848183f7d3c539fdd52 (MD5)","Made available in DSpace on 2014-09-16T17:18:00Z (GMT). No. of bitstreams: 2 Shi_Zhi.pdf: 1401640 bytes, checksum: f7b1d363045a4848183f7d3c539fdd52 (MD5) license.txt: 4056 bytes, checksum: 2eeba217c60f8b64020ad2b9b76641fc (MD5)","Embargo set by: Seth Robbins for item 50604 Lift date: 2016-09-16T17:18:17Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 50604 on 2016-09-22T20:59:16Z."]},{"key":"dc:title","label":"Title","values":["Integrating multiple conflicting sources by truth discovery and source quality estimation"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei"],"dc:creator":["Zhi, Shi"],"dc:date":["2014-09-16T17:18:00Z","2016-09-22T20:59:16Z","2014-08","2014-09-16"],"dc:description":["Multiple descriptions about the same entity from different sources will inevitably result in data or information inconsistency. Among conflicting pieces of information, which one is the most trustworthy? How to detect the fraudulence of a rumor? Obviously, it is unrealistic to curate and validate the trustworthiness of every piece of information because of the high cost of human labeling and lack of experts. To find the truth of each entity, much research work has shown that considering the quality of information providers can improve the performance of data integration. Due to different quality of data sources, it is hard to find a general solution that works for every case. Therefore, we start from a general setting of truth analysis at first and narrow down to two basic problems in data integration. We first propose a general framework to deal with numerical data with flexibility of defining loss function. Source quality is represented by a vector to model the source credibility in different error interval. Then we propose a new method called No Truth Truth Model(NTTM) to deal with truth existence problem in low-quality data. Preliminary experiments on real stock data and slot filling data show promising results.","Item withdrawn by Laura Spradlin (lspradl2@illinois.edu) on 2014-07-24T14:56:06Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Zhi_Shi.pdf: 1401640 bytes, checksum: f7b1d363045a4848183f7d3c539fdd52 (MD5)","Made available in DSpace on 2014-09-16T17:18:00Z (GMT). No. of bitstreams: 2 Shi_Zhi.pdf: 1401640 bytes, checksum: f7b1d363045a4848183f7d3c539fdd52 (MD5) license.txt: 4056 bytes, checksum: 2eeba217c60f8b64020ad2b9b76641fc (MD5)","Embargo set by: Seth Robbins for item 50604 Lift date: 2016-09-16T17:18:17Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 50604 on 2016-09-22T20:59:16Z."],"dc:identifier":["http://hdl.handle.net/2142/50493"],"dc:language":["en"],"dc:rights":["Copyright 2014 Shi Zhi"],"dc:subject":["Truth Discovery","Data Integration","Data Quality"],"dc:title":["Integrating multiple conflicting sources by truth discovery and source quality estimation"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:40Z"}