{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/125814"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/125814","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Subsampling based inference for network data","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2026-08-01","abstract_has_math":false,"creators":["Chakrabarty, Sayan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Chen, Yuguo","Sengupta, Srijan","Shao, Xiaofeng","Simpson, Douglas"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-07-11","date_published":"2024-07-11","updated_at":"2026-07-22T22:25:02Z","subjects":["Blockmodels","Community Detection","Large Networks","Model Selection","Network Cross-validation","Network Subsampling","Random Dot Product Graph"],"languages":["en","eng"],"rights":["Copyright 2024 Sayan Chakrabarty"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/125814","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chen, Yuguo","Sengupta, Srijan","Shao, Xiaofeng","Simpson, Douglas"]},{"key":"dc:creator","label":"Author","values":["Chakrabarty, Sayan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-07-11","2024-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Blockmodels","Community Detection","Large Networks","Model Selection","Network Cross-validation","Network Subsampling","Random Dot Product Graph"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Sayan Chakrabarty"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/125814"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-08-01","The student, Sayan Chakrabarty, accepted the attached license on 2024-07-10 at 10:10.","The student, Sayan Chakrabarty, submitted this Dissertation for approval on 2024-07-10 at 10:21.","This Dissertation was approved for publication on 2024-07-11 at 17:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21039 on 2025-02-04 at 21:25:48","Contemporary systems often comprise interactions between numerous agents, typically represented using networks. Network data is widespread across disciplines such as social sciences, biological sciences, information technology, and computer sciences. As technology rapidly advances, networks arising from such fields are becoming increasingly large and complex. Effectively analyzing these networks poses challenges concerning computational feasibility and the choice of suitable analytical models. This dissertation addresses two such problems in this area. Large networks are becoming widespread in scientific fields. Performing statistical analysis on such large networks is challenging due to high computation time and memory requirements. In the second chapter of this dissertation, we introduce a subsampling-based divide-and-conquer algorithm, SONNET, for detecting communities in large networks. The algorithm divides the original network into several subnetworks with an overlap part and applies a community detection algorithm to each subnetwork. The results from each subnetwork are combined using a label matching approach to determine the final community labels. This method significantly reduces both memory and computation costs since it only requires processing and storing the smaller subnetworks. It is also parallelizable, enhancing its speed. Theoretical and numerical performance of the algorithm is also presented in this chapter. Complex and extensive networks are increasingly common in scientific applications across various fields. Despite the availability of numerous network models and methodologies, cross-validation on networks is still difficult due to the unique structure of network data. In the third chapter, we propose a general cross-validation procedure, CROISSANT, based on subsampling for networks. The proposed algorithm splits the original network into multiple subnetworks with a shared overlap, creating a training set comprising the subnetworks and a test set with the node pairs between the subnetworks. This train-test split forms the basis for a network cross-validation procedure that can be used for a broad range of model selection and parameter tuning problems for network data. The method is computationally efficient for large networks, as it utilizes smaller subnetworks for the training process. It is also adaptable for specific network model selection and parameter tuning, with theoretical justifications provided as well. Numerical results show that the proposed algorithm accurately performs model selection and parameter tuning on various simulated and real networks from diverse models. They also indicate that the method is faster than existing network cross-validation methods."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Subsampling based inference for network data"]}]}],"canonical_facts":{"dc:contributor":["Chen, Yuguo","Sengupta, Srijan","Shao, Xiaofeng","Simpson, Douglas"],"dc:creator":["Chakrabarty, Sayan"],"dc:date":["2024-07-11","2024-08"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-08-01","The student, Sayan Chakrabarty, accepted the attached license on 2024-07-10 at 10:10.","The student, Sayan Chakrabarty, submitted this Dissertation for approval on 2024-07-10 at 10:21.","This Dissertation was approved for publication on 2024-07-11 at 17:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21039 on 2025-02-04 at 21:25:48","Contemporary systems often comprise interactions between numerous agents, typically represented using networks. Network data is widespread across disciplines such as social sciences, biological sciences, information technology, and computer sciences. As technology rapidly advances, networks arising from such fields are becoming increasingly large and complex. Effectively analyzing these networks poses challenges concerning computational feasibility and the choice of suitable analytical models. This dissertation addresses two such problems in this area. Large networks are becoming widespread in scientific fields. Performing statistical analysis on such large networks is challenging due to high computation time and memory requirements. In the second chapter of this dissertation, we introduce a subsampling-based divide-and-conquer algorithm, SONNET, for detecting communities in large networks. The algorithm divides the original network into several subnetworks with an overlap part and applies a community detection algorithm to each subnetwork. The results from each subnetwork are combined using a label matching approach to determine the final community labels. This method significantly reduces both memory and computation costs since it only requires processing and storing the smaller subnetworks. It is also parallelizable, enhancing its speed. Theoretical and numerical performance of the algorithm is also presented in this chapter. Complex and extensive networks are increasingly common in scientific applications across various fields. Despite the availability of numerous network models and methodologies, cross-validation on networks is still difficult due to the unique structure of network data. In the third chapter, we propose a general cross-validation procedure, CROISSANT, based on subsampling for networks. The proposed algorithm splits the original network into multiple subnetworks with a shared overlap, creating a training set comprising the subnetworks and a test set with the node pairs between the subnetworks. This train-test split forms the basis for a network cross-validation procedure that can be used for a broad range of model selection and parameter tuning problems for network data. The method is computationally efficient for large networks, as it utilizes smaller subnetworks for the training process. It is also adaptable for specific network model selection and parameter tuning, with theoretical justifications provided as well. Numerical results show that the proposed algorithm accurately performs model selection and parameter tuning on various simulated and real networks from diverse models. They also indicate that the method is faster than existing network cross-validation methods."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/125814"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Sayan Chakrabarty"],"dc:subject":["Blockmodels","Community Detection","Large Networks","Model Selection","Network Cross-validation","Network Subsampling","Random Dot Product Graph"],"dc:title":["Subsampling based inference for network data"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}