Back to search

University of Illinois at Urbana-Champaign

Information sampling from online social networks

Abstract

dc:description

Data sampling from online social networks is a pre-requisite step for several downstream applications. Further, the massive size of the online social networks coupled with several API limitations and restrictions to the social information makes sampling a challenging problem. This thesis addresses some of the sampling challenges by proposing novel samplers for sampling attributes (content), hidden attributes (population), and networks from online social networks. Specifically, we first propose an information-based sampler in Chapter 3 for sampling content from online social networks. We leverage the surprise of content to direct our sampler towards informative content. The surprise-based sampling strategy allows us to sample the cluster shape and boundary of content clusters efficiently, which is crucial for several data-mining tasks, including clustering, classification, regression, and attribute discovery. We demonstrate our proposed sampler's efficacy on a suite of thirty real-world networks and four data-mining tasks. We further show through empirical counterfactual analysis that network structure does not hinder the performance of surprise-based link-trace samplers in many real-world datasets. Next in Chapter 4, we propose a novel attributed search-based sampler to sample hidden populations. We use a decision-tree-based search strategy to query the attribute-search space systematically. Our proposed decision-tree Thompson sampler follows the exploration and exploitation strategy to sample hidden populations from social networks. We demonstrate our sampler's efficacy over a suite of fourteen sampling tasks on three online social sites and five offline datasets. Furthermore, we show the impact of several factors, like page size, missing information, and noise, affecting hidden population sampling in real-world social networks. Finally, in Chapter 5, we propose a novel framework for learning network samplers. First, we show through theoretical and empirical proof that there exists no universal network sampler that can preserve all the topological properties of the underlying graph in the sample. To address the non-existence issue, we propose a reinforcement learning framework that learns high-quality sampling policies according to application needs. We demonstrate the efficacy of our proposed sampling framework through extensive experiments across ten different graph families and seven diverse tasks. In summary, this thesis develops several sampling strategies for sampling information (attribute, hidden attribute, network) from online social networks while being cognizant of API restrictions' constraints. We propose adaptive samplers that can cater to different application needs.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2021

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Kumar, Suhansanu
Contributors dc:contributor
  • Sundaram, Hari
  • Tong, Hanghang
  • Koyejo, Sanmi
  • Jiang, Meng

Subjects

dc:subject × 7

Rights

dc:rights
Statement dc:rights
  • Copyright 2021 Suhansanu Kumar
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/110469

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Kumar, Suhansanu. Information sampling from online social networks. Dissertation thesis, University of Illinois at Urbana-Champaign, 2021. http://hdl.handle.net/2142/110469