{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/110469"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/110469","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Information sampling from online social networks","abstract":"The student, Suhansanu Kumar, accepted the attached license on 2021-04-14 at 11:22.","abstract_html":"The student, Suhansanu Kumar, accepted the attached license on 2021-04-14 at 11:22.","abstract_has_math":false,"creators":["Kumar, Suhansanu"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Sundaram, Hari","Tong, Hanghang","Koyejo, Sanmi","Jiang, Meng"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-09-17T01:10:49Z","date_published":"2021-09-17T01:10:49Z","updated_at":"2026-07-22T22:24:50Z","subjects":["sampling","network","graph","online social network","reinforcement learning","hidden population","content"],"languages":["en"],"rights":["Copyright 2021 Suhansanu Kumar"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/110469","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Sundaram, Hari","Tong, Hanghang","Koyejo, Sanmi","Jiang, Meng"]},{"key":"dc:creator","label":"Author","values":["Kumar, Suhansanu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-09-17T01:10:49Z","2021-04-14","2021-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["sampling","network","graph","online social network","reinforcement learning","hidden population","content"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Suhansanu Kumar"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/110469"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The student, Suhansanu Kumar, accepted the attached license on 2021-04-14 at 11:22.","The student, Suhansanu Kumar, submitted this Dissertation for approval on 2021-04-14 at 11:24.","This Dissertation was approved for publication on 2021-04-14 at 14:43.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16315 on 2021-09-16 at 16:41:14","Made available in DSpace on 2021-09-17T01:10:49Z (GMT). No. of bitstreams: 3 KUMAR-DISSERTATION-2021.pdf: 3836895 bytes, checksum: aefb4b3574532f026087ef6a674b6685 (MD5) LICENSE.txt: 4212 bytes, checksum: 0eb7dec2721bf4c56cdfc4f2712b86b9 (MD5) PROQUEST_LICENSE.txt: 4558 bytes, checksum: 8f55cc1fbb040f96b5d4f52575c18bae (MD5) Previous issue date: 2021-04-14","Data sampling from online social networks is a pre-requisite step for several downstream applications. Further, the massive size of the online social networks coupled with several API limitations and restrictions to the social information makes sampling a challenging problem. This thesis addresses some of the sampling challenges by proposing novel samplers for sampling attributes (content), hidden attributes (population), and networks from online social networks. Specifically, we first propose an information-based sampler in Chapter 3 for sampling content from online social networks. We leverage the surprise of content to direct our sampler towards informative content. The surprise-based sampling strategy allows us to sample the cluster shape and boundary of content clusters efficiently, which is crucial for several data-mining tasks, including clustering, classification, regression, and attribute discovery. We demonstrate our proposed sampler's efficacy on a suite of thirty real-world networks and four data-mining tasks. We further show through empirical counterfactual analysis that network structure does not hinder the performance of surprise-based link-trace samplers in many real-world datasets. Next in Chapter 4, we propose a novel attributed search-based sampler to sample hidden populations. We use a decision-tree-based search strategy to query the attribute-search space systematically. Our proposed decision-tree Thompson sampler follows the exploration and exploitation strategy to sample hidden populations from social networks. We demonstrate our sampler's efficacy over a suite of fourteen sampling tasks on three online social sites and five offline datasets. Furthermore, we show the impact of several factors, like page size, missing information, and noise, affecting hidden population sampling in real-world social networks. Finally, in Chapter 5, we propose a novel framework for learning network samplers. First, we show through theoretical and empirical proof that there exists no universal network sampler that can preserve all the topological properties of the underlying graph in the sample. To address the non-existence issue, we propose a reinforcement learning framework that learns high-quality sampling policies according to application needs. We demonstrate the efficacy of our proposed sampling framework through extensive experiments across ten different graph families and seven diverse tasks. In summary, this thesis develops several sampling strategies for sampling information (attribute, hidden attribute, network) from online social networks while being cognizant of API restrictions' constraints. We propose adaptive samplers that can cater to different application needs.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-09-16 without embargo terms"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Information sampling from online social networks"]}]}],"canonical_facts":{"dc:contributor":["Sundaram, Hari","Tong, Hanghang","Koyejo, Sanmi","Jiang, Meng"],"dc:creator":["Kumar, Suhansanu"],"dc:date":["2021-09-17T01:10:49Z","2021-04-14","2021-05"],"dc:description":["The student, Suhansanu Kumar, accepted the attached license on 2021-04-14 at 11:22.","The student, Suhansanu Kumar, submitted this Dissertation for approval on 2021-04-14 at 11:24.","This Dissertation was approved for publication on 2021-04-14 at 14:43.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16315 on 2021-09-16 at 16:41:14","Made available in DSpace on 2021-09-17T01:10:49Z (GMT). No. of bitstreams: 3 KUMAR-DISSERTATION-2021.pdf: 3836895 bytes, checksum: aefb4b3574532f026087ef6a674b6685 (MD5) LICENSE.txt: 4212 bytes, checksum: 0eb7dec2721bf4c56cdfc4f2712b86b9 (MD5) PROQUEST_LICENSE.txt: 4558 bytes, checksum: 8f55cc1fbb040f96b5d4f52575c18bae (MD5) Previous issue date: 2021-04-14","Data sampling from online social networks is a pre-requisite step for several downstream applications. Further, the massive size of the online social networks coupled with several API limitations and restrictions to the social information makes sampling a challenging problem. This thesis addresses some of the sampling challenges by proposing novel samplers for sampling attributes (content), hidden attributes (population), and networks from online social networks. Specifically, we first propose an information-based sampler in Chapter 3 for sampling content from online social networks. We leverage the surprise of content to direct our sampler towards informative content. The surprise-based sampling strategy allows us to sample the cluster shape and boundary of content clusters efficiently, which is crucial for several data-mining tasks, including clustering, classification, regression, and attribute discovery. We demonstrate our proposed sampler's efficacy on a suite of thirty real-world networks and four data-mining tasks. We further show through empirical counterfactual analysis that network structure does not hinder the performance of surprise-based link-trace samplers in many real-world datasets. Next in Chapter 4, we propose a novel attributed search-based sampler to sample hidden populations. We use a decision-tree-based search strategy to query the attribute-search space systematically. Our proposed decision-tree Thompson sampler follows the exploration and exploitation strategy to sample hidden populations from social networks. We demonstrate our sampler's efficacy over a suite of fourteen sampling tasks on three online social sites and five offline datasets. Furthermore, we show the impact of several factors, like page size, missing information, and noise, affecting hidden population sampling in real-world social networks. Finally, in Chapter 5, we propose a novel framework for learning network samplers. First, we show through theoretical and empirical proof that there exists no universal network sampler that can preserve all the topological properties of the underlying graph in the sample. To address the non-existence issue, we propose a reinforcement learning framework that learns high-quality sampling policies according to application needs. We demonstrate the efficacy of our proposed sampling framework through extensive experiments across ten different graph families and seven diverse tasks. In summary, this thesis develops several sampling strategies for sampling information (attribute, hidden attribute, network) from online social networks while being cognizant of API restrictions' constraints. We propose adaptive samplers that can cater to different application needs.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-09-16 without embargo terms"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/110469"],"dc:language":["en"],"dc:rights":["Copyright 2021 Suhansanu Kumar"],"dc:subject":["sampling","network","graph","online social network","reinforcement learning","hidden population","content"],"dc:title":["Information sampling from online social networks"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:50Z"}