Back to results

University of Illinois at Urbana-Champaign

Subsampling based inference for network data

Abstract

dc:description

Contemporary systems often comprise interactions between numerous agents, typically represented using networks. Network data is widespread across disciplines such as social sciences, biological sciences, information technology, and computer sciences. As technology rapidly advances, networks arising from such fields are becoming increasingly large and complex. Effectively analyzing these networks poses challenges concerning computational feasibility and the choice of suitable analytical models. This dissertation addresses two such problems in this area. Large networks are becoming widespread in scientific fields. Performing statistical analysis on such large networks is challenging due to high computation time and memory requirements. In the second chapter of this dissertation, we introduce a subsampling-based divide-and-conquer algorithm, SONNET, for detecting communities in large networks. The algorithm divides the original network into several subnetworks with an overlap part and applies a community detection algorithm to each subnetwork. The results from each subnetwork are combined using a label matching approach to determine the final community labels. This method significantly reduces both memory and computation costs since it only requires processing and storing the smaller subnetworks. It is also parallelizable, enhancing its speed. Theoretical and numerical performance of the algorithm is also presented in this chapter. Complex and extensive networks are increasingly common in scientific applications across various fields. Despite the availability of numerous network models and methodologies, cross-validation on networks is still difficult due to the unique structure of network data. In the third chapter, we propose a general cross-validation procedure, CROISSANT, based on subsampling for networks. The proposed algorithm splits the original network into multiple subnetworks with a shared overlap, creating a training set comprising the subnetworks and a test set with the node pairs between the subnetworks. This train-test split forms the basis for a network cross-validation procedure that can be used for a broad range of model selection and parameter tuning problems for network data. The method is computationally efficient for large networks, as it utilizes smaller subnetworks for the training process. It is also adaptable for specific network model selection and parameter tuning, with theoretical justifications provided as well. Numerical results show that the proposed algorithm accurately performs model selection and parameter tuning on various simulated and real networks from diverse models. They also indicate that the method is faster than existing network cross-validation methods.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Statistics
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Chakrabarty, Sayan
Contributors dc:contributor
  • Chen, Yuguo
  • Sengupta, Srijan
  • Shao, Xiaofeng
  • Simpson, Douglas

Subjects

dc:subject × 7

Rights

dc:rights
Statement dc:rights
  • Copyright 2024 Sayan Chakrabarty
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/125814

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Chakrabarty, Sayan. Subsampling based inference for network data. Dissertation thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/125814