{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113012"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113012","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Mining social sensing data: Representation, modeling, and applications","abstract":"\"Social sensing refers to using humans as \"\"sensors\"\" to collect information about external physical events, such as disasters, protests, and traffic. This is motivated by the observations that today millions of people often share their observations in the physical world online via social media platforms and mobile apps. However, social sensing data are noisy, multi-modal, and unreliable. Errors are highly correlated as many people may retweet the same false information. In addition, social sensing data are very unique as its text content and graph structures vary a lot for different tasks. Hence, it requires human graders to label the data for new domains when using supervised learning, costing a lot of human labor. The main goal of this dissertation is to develop unsupervised solutions that learn latent structures and models from observed social media data to advance social sensing applications, including misinformation detection, link prediction, and polarity detection. The work lies in the exploration of learning systems that include both analytic schemes, where model structure is known ahead of time (but parameters need to be estimated), and data-driven schemes, where the latent model structure itself is to be learned from data without prior knowledge. As a means of exploring the gamut of such unsupervised schemes, we consider a broad range of social sensing applications that call for different modeling complexity. This dissertation first focuses on traditional maximum-likelihood estimators - a solution that estimates unknown parameter values of analytical likelihood models. We consider truth discovery that estimates the veracity of claims made by different users with unknown reliability on social media. We then extend the truth-finding algorithm by considering additional content features to further improve the quality estimation results. Accordingly, we propose two flavors of maximum likelihood estimators, EM-MultiF and PEM-MultiF, that jointly learn the importance of different content features together with the veracity of observations. In order to further improve truth discovery, we develop an optimal source selection model to minimize the expected fusion error when some sources can influence others. In order to better leverage multi-modal data in social sensing, next we consider representation learning that uncovers semantically meaningful latent factors from the observed data without the benefit of known model structure. We first develop a novel maximum-likelihood estimator, ControlVAE, that combines control theory with a Variational Autoencoder (VAE) to disentangle the latent factors from unstructured data. The proposed ControlVAE can not only disentangle the latent factors but also solve the posterior collapse problem. Finally, we apply ControlVAE and representation learning model to different social sensing applications, including misinformation detection and polarity analysis. A trade-off is discussed between the different modeling approaches in terms of applicability, accuracy, and robustness to inform the design of future social sensing systems.\"","abstract_html":"&quot;Social sensing refers to using humans as &quot;&quot;sensors&quot;&quot; to collect information about external physical events, such as disasters, protests, and traffic. This is motivated by the observations that today millions of people often share their observations in the physical world online via social media platforms and mobile apps. However, social sensing data are noisy, multi-modal, and unreliable. Errors are highly correlated as many people may retweet the same false information. In addition, social sensing data are very unique as its text content and graph structures vary a lot for different tasks. Hence, it requires human graders to label the data for new domains when using supervised learning, costing a lot of human labor. The main goal of this dissertation is to develop unsupervised solutions that learn latent structures and models from observed social media data to advance social sensing applications, including misinformation detection, link prediction, and polarity detection. The work lies in the exploration of learning systems that include both analytic schemes, where model structure is known ahead of time (but parameters need to be estimated), and data-driven schemes, where the latent model structure itself is to be learned from data without prior knowledge. As a means of exploring the gamut of such unsupervised schemes, we consider a broad range of social sensing applications that call for different modeling complexity. This dissertation first focuses on traditional maximum-likelihood estimators - a solution that estimates unknown parameter values of analytical likelihood models. We consider truth discovery that estimates the veracity of claims made by different users with unknown reliability on social media. We then extend the truth-finding algorithm by considering additional content features to further improve the quality estimation results. Accordingly, we propose two flavors of maximum likelihood estimators, EM-MultiF and PEM-MultiF, that jointly learn the importance of different content features together with the veracity of observations. In order to further improve truth discovery, we develop an optimal source selection model to minimize the expected fusion error when some sources can influence others. In order to better leverage multi-modal data in social sensing, next we consider representation learning that uncovers semantically meaningful latent factors from the observed data without the benefit of known model structure. We first develop a novel maximum-likelihood estimator, ControlVAE, that combines control theory with a Variational Autoencoder (VAE) to disentangle the latent factors from unstructured data. The proposed ControlVAE can not only disentangle the latent factors but also solve the posterior collapse problem. Finally, we apply ControlVAE and representation learning model to different social sensing applications, including misinformation detection and polarity analysis. A trade-off is discussed between the different modeling approaches in terms of applicability, accuracy, and robustness to inform the design of future social sensing systems.&quot;","abstract_has_math":false,"creators":["Shao, Huajie"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Abdelzaher, Tarek","Han, Jiawei","Ji, Heng","Kaplan, Lance"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T21:45:33Z","date_published":"2022-01-12T21:45:33Z","updated_at":"2026-07-22T22:24:52Z","subjects":["Truth Discovery","Social Sensing","Maximum-likelihood estimator","Robustness","Unsupervised learning"],"languages":["en"],"rights":["Copyright 2021 Huajie Shao"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113012","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Abdelzaher, Tarek","Han, Jiawei","Ji, Heng","Kaplan, Lance"]},{"key":"dc:creator","label":"Author","values":["Shao, Huajie"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T21:45:33Z","2021-07-12","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["Thesis","text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Truth Discovery","Social Sensing","Maximum-likelihood estimator","Robustness","Unsupervised learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Huajie Shao"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113012"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"Social sensing refers to using humans as \"\"sensors\"\" to collect information about external physical events, such as disasters, protests, and traffic. This is motivated by the observations that today millions of people often share their observations in the physical world online via social media platforms and mobile apps. However, social sensing data are noisy, multi-modal, and unreliable. Errors are highly correlated as many people may retweet the same false information. In addition, social sensing data are very unique as its text content and graph structures vary a lot for different tasks. Hence, it requires human graders to label the data for new domains when using supervised learning, costing a lot of human labor. The main goal of this dissertation is to develop unsupervised solutions that learn latent structures and models from observed social media data to advance social sensing applications, including misinformation detection, link prediction, and polarity detection. The work lies in the exploration of learning systems that include both analytic schemes, where model structure is known ahead of time (but parameters need to be estimated), and data-driven schemes, where the latent model structure itself is to be learned from data without prior knowledge. As a means of exploring the gamut of such unsupervised schemes, we consider a broad range of social sensing applications that call for different modeling complexity. This dissertation first focuses on traditional maximum-likelihood estimators - a solution that estimates unknown parameter values of analytical likelihood models. We consider truth discovery that estimates the veracity of claims made by different users with unknown reliability on social media. We then extend the truth-finding algorithm by considering additional content features to further improve the quality estimation results. Accordingly, we propose two flavors of maximum likelihood estimators, EM-MultiF and PEM-MultiF, that jointly learn the importance of different content features together with the veracity of observations. In order to further improve truth discovery, we develop an optimal source selection model to minimize the expected fusion error when some sources can influence others. In order to better leverage multi-modal data in social sensing, next we consider representation learning that uncovers semantically meaningful latent factors from the observed data without the benefit of known model structure. We first develop a novel maximum-likelihood estimator, ControlVAE, that combines control theory with a Variational Autoencoder (VAE) to disentangle the latent factors from unstructured data. The proposed ControlVAE can not only disentangle the latent factors but also solve the posterior collapse problem. Finally, we apply ControlVAE and representation learning model to different social sensing applications, including misinformation detection and polarity analysis. A trade-off is discussed between the different modeling approaches in terms of applicability, accuracy, and robustness to inform the design of future social sensing systems.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Huajie Shao, accepted the attached license on 2021-07-09 at 19:12.","The student, Huajie Shao, submitted this Dissertation for approval on 2021-07-09 at 19:19.","This Dissertation was approved for publication on 2021-07-12 at 10:07.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16842 on 2022-01-12 at 12:44:48","Made available in DSpace on 2022-01-12T21:45:33Z (GMT). No. of bitstreams: 3 SHAO-DISSERTATION-2021.pdf: 2022188 bytes, checksum: fd44156c181533be4f18a73d9463f6db (MD5) LICENSE.txt: 4208 bytes, checksum: 5318b51fee1b8cb45092168089256357 (MD5) PROQUEST_LICENSE.txt: 4554 bytes, checksum: 68b093eedeeccee6988a617d93506047 (MD5) Previous issue date: 2021-07-12"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Mining social sensing data: Representation, modeling, and applications"]}]}],"canonical_facts":{"dc:contributor":["Abdelzaher, Tarek","Han, Jiawei","Ji, Heng","Kaplan, Lance"],"dc:creator":["Shao, Huajie"],"dc:date":["2022-01-12T21:45:33Z","2021-07-12","2021-08"],"dc:description":["\"Social sensing refers to using humans as \"\"sensors\"\" to collect information about external physical events, such as disasters, protests, and traffic. This is motivated by the observations that today millions of people often share their observations in the physical world online via social media platforms and mobile apps. However, social sensing data are noisy, multi-modal, and unreliable. Errors are highly correlated as many people may retweet the same false information. In addition, social sensing data are very unique as its text content and graph structures vary a lot for different tasks. Hence, it requires human graders to label the data for new domains when using supervised learning, costing a lot of human labor. The main goal of this dissertation is to develop unsupervised solutions that learn latent structures and models from observed social media data to advance social sensing applications, including misinformation detection, link prediction, and polarity detection. The work lies in the exploration of learning systems that include both analytic schemes, where model structure is known ahead of time (but parameters need to be estimated), and data-driven schemes, where the latent model structure itself is to be learned from data without prior knowledge. As a means of exploring the gamut of such unsupervised schemes, we consider a broad range of social sensing applications that call for different modeling complexity. This dissertation first focuses on traditional maximum-likelihood estimators - a solution that estimates unknown parameter values of analytical likelihood models. We consider truth discovery that estimates the veracity of claims made by different users with unknown reliability on social media. We then extend the truth-finding algorithm by considering additional content features to further improve the quality estimation results. Accordingly, we propose two flavors of maximum likelihood estimators, EM-MultiF and PEM-MultiF, that jointly learn the importance of different content features together with the veracity of observations. In order to further improve truth discovery, we develop an optimal source selection model to minimize the expected fusion error when some sources can influence others. In order to better leverage multi-modal data in social sensing, next we consider representation learning that uncovers semantically meaningful latent factors from the observed data without the benefit of known model structure. We first develop a novel maximum-likelihood estimator, ControlVAE, that combines control theory with a Variational Autoencoder (VAE) to disentangle the latent factors from unstructured data. The proposed ControlVAE can not only disentangle the latent factors but also solve the posterior collapse problem. Finally, we apply ControlVAE and representation learning model to different social sensing applications, including misinformation detection and polarity analysis. A trade-off is discussed between the different modeling approaches in terms of applicability, accuracy, and robustness to inform the design of future social sensing systems.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Huajie Shao, accepted the attached license on 2021-07-09 at 19:12.","The student, Huajie Shao, submitted this Dissertation for approval on 2021-07-09 at 19:19.","This Dissertation was approved for publication on 2021-07-12 at 10:07.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16842 on 2022-01-12 at 12:44:48","Made available in DSpace on 2022-01-12T21:45:33Z (GMT). No. of bitstreams: 3 SHAO-DISSERTATION-2021.pdf: 2022188 bytes, checksum: fd44156c181533be4f18a73d9463f6db (MD5) LICENSE.txt: 4208 bytes, checksum: 5318b51fee1b8cb45092168089256357 (MD5) PROQUEST_LICENSE.txt: 4554 bytes, checksum: 68b093eedeeccee6988a617d93506047 (MD5) Previous issue date: 2021-07-12"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113012"],"dc:language":["en"],"dc:rights":["Copyright 2021 Huajie Shao"],"dc:subject":["Truth Discovery","Social Sensing","Maximum-likelihood estimator","Robustness","Unsupervised learning"],"dc:title":["Mining social sensing data: Representation, modeling, and applications"],"dc:type":["Thesis","text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}