{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/92763"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/92763","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Statistical analysis of networks with community structure and bootstrap methods for big data","abstract":"This dissertation is divided into two parts, concerning two areas of statistical methodology. The first part of this dissertation concerns statistical analysis of networks with community structure. The second part of this dissertation concerns bootstrap methods for big data. Statistical analysis of networks with community structure: Networks are ubiquitous in today's world --- network data appears from varied fields such as scientific studies, sociology, technology, social media and the Internet, to name a few. An interesting aspect of many real-world networks is the presence of community structure and the problem of detecting this community structure. In the first chapter, we consider heterogeneous networks which seems to have not been considered in the statistical community detection literature. We propose a blockmodel for heterogeneous networks with community structure, and introduce a heterogeneous spectral clustering algorithm for community detection in heterogeneous networks. Theoretical properties of the clustering algorithm under the proposed model are studied, along with simulation study and data analysis. A network feature that is closely associated with community structure is the popularity of nodes in different communities. Neither the classical stochastic blockmodel nor its degree-corrected extension can satisfactorily capture the dynamics of node popularity. In the second chapter, we propose a popularity-adjusted blockmodel for flexible modeling of node popularity. We establish consistency of likelihood modularity for community detection under the proposed model, and illustrate the improved empirical insights that can be gained through this methodology by analyzing the political blogs network and the British MP network, as well as in simulation studies. Bootstrap methods for big data: Resampling methods provide a powerful method of evaluating the precision of a wide variety of statistical inference methods. The complexity and massive size of big data makes it infeasible to apply traditional resampling methods for big data. In the first chapter, we consider the problem of resampling for irregularly spaced dependent data. Traditional block-based resampling or subsampling schemes for stationary data are difficult to implement when the data are irregularly spaced, as it takes careful programming effort to partition the sampling region into complete and incomplete blocks. We develop a resampling method called Dependent Random Weighting (DRW) for irregularly spaced dependent data, where instead of using blocks we use random weights to resample the data. By allowing the random weights to be dependent, the dependency structure of the data can be preserved in the resamples. We study the theoretical properties of this resampling methods as well as its numerical performance in simulations. In the second chapter, we consider the problem of resampling in massive data, where traditional methods like bootstrap (for independent data) or moving block bootstrap (for dependent data) can be computationally infeasible since each resample has effective size of the same order as the sample. We develop a new resampling method called subsampled double bootstrap (SDB) for both independent and stationary data. SDB works by choosing small random subsets of the massive data, and then constructing a single resample from that subset using bootstrap (for independent data) or moving block bootstrap (for stationary data). We study theoretical properties of SDB as well as its numerical performance in simulated data and real data. Extending the underlying ideas of the second chapter, we introduce two new resampling strategies for big data in Chapter 3. The first strategy is called aggregation of little bootstraps or ALB, a generalized resampling technique that includes the SDB as a special case. The second strategy is called subsampled residual bootstrap or SRB, a fast version of residual bootstrap intended for massive regression models. We study both methods through simulations.","abstract_html":"This dissertation is divided into two parts, concerning two areas of statistical methodology. The first part of this dissertation concerns statistical analysis of networks with community structure. The second part of this dissertation concerns bootstrap methods for big data. Statistical analysis of networks with community structure: Networks are ubiquitous in today&#x27;s world --- network data appears from varied fields such as scientific studies, sociology, technology, social media and the Internet, to name a few. An interesting aspect of many real-world networks is the presence of community structure and the problem of detecting this community structure. In the first chapter, we consider heterogeneous networks which seems to have not been considered in the statistical community detection literature. We propose a blockmodel for heterogeneous networks with community structure, and introduce a heterogeneous spectral clustering algorithm for community detection in heterogeneous networks. Theoretical properties of the clustering algorithm under the proposed model are studied, along with simulation study and data analysis. A network feature that is closely associated with community structure is the popularity of nodes in different communities. Neither the classical stochastic blockmodel nor its degree-corrected extension can satisfactorily capture the dynamics of node popularity. In the second chapter, we propose a popularity-adjusted blockmodel for flexible modeling of node popularity. We establish consistency of likelihood modularity for community detection under the proposed model, and illustrate the improved empirical insights that can be gained through this methodology by analyzing the political blogs network and the British MP network, as well as in simulation studies. Bootstrap methods for big data: Resampling methods provide a powerful method of evaluating the precision of a wide variety of statistical inference methods. The complexity and massive size of big data makes it infeasible to apply traditional resampling methods for big data. In the first chapter, we consider the problem of resampling for irregularly spaced dependent data. Traditional block-based resampling or subsampling schemes for stationary data are difficult to implement when the data are irregularly spaced, as it takes careful programming effort to partition the sampling region into complete and incomplete blocks. We develop a resampling method called Dependent Random Weighting (DRW) for irregularly spaced dependent data, where instead of using blocks we use random weights to resample the data. By allowing the random weights to be dependent, the dependency structure of the data can be preserved in the resamples. We study the theoretical properties of this resampling methods as well as its numerical performance in simulations. In the second chapter, we consider the problem of resampling in massive data, where traditional methods like bootstrap (for independent data) or moving block bootstrap (for dependent data) can be computationally infeasible since each resample has effective size of the same order as the sample. We develop a new resampling method called subsampled double bootstrap (SDB) for both independent and stationary data. SDB works by choosing small random subsets of the massive data, and then constructing a single resample from that subset using bootstrap (for independent data) or moving block bootstrap (for stationary data). We study theoretical properties of SDB as well as its numerical performance in simulated data and real data. Extending the underlying ideas of the second chapter, we introduce two new resampling strategies for big data in Chapter 3. The first strategy is called aggregation of little bootstraps or ALB, a generalized resampling technique that includes the SDB as a special case. The second strategy is called subsampled residual bootstrap or SRB, a fast version of residual bootstrap intended for massive regression models. We study both methods through simulations.","abstract_has_math":false,"creators":["Sengupta, Srijan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Chen, Yuguo","Shao, Xiaofeng","Simpson, Douglas G.","Marden, John I."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-11-10T17:50:10Z","date_published":"2016-11-10T17:50:10Z","updated_at":"2026-07-22T22:26:35Z","subjects":["network data","resampling","community structure","big data"],"languages":["en"],"rights":["Copyright 2016 Srijan Sengupta"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/92763","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chen, Yuguo","Shao, Xiaofeng","Simpson, Douglas G.","Marden, John I."]},{"key":"dc:creator","label":"Author","values":["Sengupta, Srijan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-11-10T17:50:10Z","2016-07-08","2016-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["network data","resampling","community structure","big data"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2016 Srijan Sengupta"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/92763"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This dissertation is divided into two parts, concerning two areas of statistical methodology. The first part of this dissertation concerns statistical analysis of networks with community structure. The second part of this dissertation concerns bootstrap methods for big data. Statistical analysis of networks with community structure: Networks are ubiquitous in today's world --- network data appears from varied fields such as scientific studies, sociology, technology, social media and the Internet, to name a few. An interesting aspect of many real-world networks is the presence of community structure and the problem of detecting this community structure. In the first chapter, we consider heterogeneous networks which seems to have not been considered in the statistical community detection literature. We propose a blockmodel for heterogeneous networks with community structure, and introduce a heterogeneous spectral clustering algorithm for community detection in heterogeneous networks. Theoretical properties of the clustering algorithm under the proposed model are studied, along with simulation study and data analysis. A network feature that is closely associated with community structure is the popularity of nodes in different communities. Neither the classical stochastic blockmodel nor its degree-corrected extension can satisfactorily capture the dynamics of node popularity. In the second chapter, we propose a popularity-adjusted blockmodel for flexible modeling of node popularity. We establish consistency of likelihood modularity for community detection under the proposed model, and illustrate the improved empirical insights that can be gained through this methodology by analyzing the political blogs network and the British MP network, as well as in simulation studies. Bootstrap methods for big data: Resampling methods provide a powerful method of evaluating the precision of a wide variety of statistical inference methods. The complexity and massive size of big data makes it infeasible to apply traditional resampling methods for big data. In the first chapter, we consider the problem of resampling for irregularly spaced dependent data. Traditional block-based resampling or subsampling schemes for stationary data are difficult to implement when the data are irregularly spaced, as it takes careful programming effort to partition the sampling region into complete and incomplete blocks. We develop a resampling method called Dependent Random Weighting (DRW) for irregularly spaced dependent data, where instead of using blocks we use random weights to resample the data. By allowing the random weights to be dependent, the dependency structure of the data can be preserved in the resamples. We study the theoretical properties of this resampling methods as well as its numerical performance in simulations. In the second chapter, we consider the problem of resampling in massive data, where traditional methods like bootstrap (for independent data) or moving block bootstrap (for dependent data) can be computationally infeasible since each resample has effective size of the same order as the sample. We develop a new resampling method called subsampled double bootstrap (SDB) for both independent and stationary data. SDB works by choosing small random subsets of the massive data, and then constructing a single resample from that subset using bootstrap (for independent data) or moving block bootstrap (for stationary data). We study theoretical properties of SDB as well as its numerical performance in simulated data and real data. Extending the underlying ideas of the second chapter, we introduce two new resampling strategies for big data in Chapter 3. The first strategy is called aggregation of little bootstraps or ALB, a generalized resampling technique that includes the SDB as a special case. The second strategy is called subsampled residual bootstrap or SRB, a fast version of residual bootstrap intended for massive regression models. We study both methods through simulations.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-11-09 without embargo terms","The student, Srijan Sengupta, accepted the attached license on 2016-07-06 at 13:49.","The student, Srijan Sengupta, submitted this Dissertation for approval on 2016-07-07 at 11:35.","This Dissertation was approved for publication on 2016-07-08 at 15:42.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9783 on 2016-11-09 at 10:22:53","Made available in DSpace on 2016-11-10T17:50:10Z (GMT). No. of bitstreams: 21 SENGUPTA-DISSERTATION-2016.pdf: 1339541 bytes, checksum: fd8fcc19303b1422bc8b98a52199d854 (MD5) ALB.tex: 16768 bytes, checksum: 31fac3ffab24a66daebe42a08b31c977 (MD5) DRW.tex: 70525 bytes, checksum: abbf3f7719a8a1869320debda6bee1e5 (MD5) HetClus.tex: 178539 bytes, checksum: f2d7d9e8b8683e6348a26fcda3bb6d8a (MD5) HetClusfiles.zip: 49530 bytes, checksum: 1370470c47d37e7d68e620423ead0b77 (MD5) IIDplots.zip: 140921 bytes, checksum: 644eeee7ba5e9c20605eb750e7051755 (MD5) PABM2.tex: 86864 bytes, checksum: 581f9673b9b3b7421cf14fc203dcc109 (MD5) SDB_RB.tex: 23243 bytes, checksum: 833b493a691f4ad3c456543385f14403 (MD5) SDB_revised.tex: 110034 bytes, checksum: c7e5795a012af4980c2554d512c39e2b (MD5) SDB_supp.tex: 10472 bytes, checksum: b5cd1121f2a22adfe9d754b255305d88 (MD5) SRB1110dgf2d10.eps: 8281 bytes, checksum: 990af1ef1ba8010d89f3ada2f2fd910a (MD5) abs.tex: 4233 bytes, checksum: 57c4a961c26fd363b81806a5489e1ae8 (MD5) ack.tex: 637 bytes, checksum: 738e17d866b7dc5d440f9626012ce68c (MD5) images.zip: 96671 bytes, checksum: 673a84c728501c33a7b12a021b27a89a (MD5) intro1.tex: 4606 bytes, checksum: c156fec92c20c3f057755c0ecbf03515 (MD5) intro2.tex: 5012 bytes, checksum: 79486107430e548d83768b2e20cf0de0 (MD5) thesis.tex: 8794 bytes, checksum: fcf2f2fe79afcf30f1144bb9a99bb997 (MD5) thesisrefs.bib: 119495 bytes, checksum: de7bb4421dc541de48d8d78279844101 (MD5) tsplots.zip: 1491356 bytes, checksum: 74c57e19cd5342054bafc4e2fb4ba886 (MD5) LICENSE.txt: 4212 bytes, checksum: a463d9bc00f8fbf7bb751cddb3790ad2 (MD5) PROQUEST_LICENSE.txt: 4558 bytes, checksum: 1a745ae8500a0c2b1af0656b4307e113 (MD5) Previous issue date: 2016-07-08"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Statistical analysis of networks with community structure and bootstrap methods for big data"]}]}],"canonical_facts":{"dc:contributor":["Chen, Yuguo","Shao, Xiaofeng","Simpson, Douglas G.","Marden, John I."],"dc:creator":["Sengupta, Srijan"],"dc:date":["2016-11-10T17:50:10Z","2016-07-08","2016-08"],"dc:description":["This dissertation is divided into two parts, concerning two areas of statistical methodology. The first part of this dissertation concerns statistical analysis of networks with community structure. The second part of this dissertation concerns bootstrap methods for big data. Statistical analysis of networks with community structure: Networks are ubiquitous in today's world --- network data appears from varied fields such as scientific studies, sociology, technology, social media and the Internet, to name a few. An interesting aspect of many real-world networks is the presence of community structure and the problem of detecting this community structure. In the first chapter, we consider heterogeneous networks which seems to have not been considered in the statistical community detection literature. We propose a blockmodel for heterogeneous networks with community structure, and introduce a heterogeneous spectral clustering algorithm for community detection in heterogeneous networks. Theoretical properties of the clustering algorithm under the proposed model are studied, along with simulation study and data analysis. A network feature that is closely associated with community structure is the popularity of nodes in different communities. Neither the classical stochastic blockmodel nor its degree-corrected extension can satisfactorily capture the dynamics of node popularity. In the second chapter, we propose a popularity-adjusted blockmodel for flexible modeling of node popularity. We establish consistency of likelihood modularity for community detection under the proposed model, and illustrate the improved empirical insights that can be gained through this methodology by analyzing the political blogs network and the British MP network, as well as in simulation studies. Bootstrap methods for big data: Resampling methods provide a powerful method of evaluating the precision of a wide variety of statistical inference methods. The complexity and massive size of big data makes it infeasible to apply traditional resampling methods for big data. In the first chapter, we consider the problem of resampling for irregularly spaced dependent data. Traditional block-based resampling or subsampling schemes for stationary data are difficult to implement when the data are irregularly spaced, as it takes careful programming effort to partition the sampling region into complete and incomplete blocks. We develop a resampling method called Dependent Random Weighting (DRW) for irregularly spaced dependent data, where instead of using blocks we use random weights to resample the data. By allowing the random weights to be dependent, the dependency structure of the data can be preserved in the resamples. We study the theoretical properties of this resampling methods as well as its numerical performance in simulations. In the second chapter, we consider the problem of resampling in massive data, where traditional methods like bootstrap (for independent data) or moving block bootstrap (for dependent data) can be computationally infeasible since each resample has effective size of the same order as the sample. We develop a new resampling method called subsampled double bootstrap (SDB) for both independent and stationary data. SDB works by choosing small random subsets of the massive data, and then constructing a single resample from that subset using bootstrap (for independent data) or moving block bootstrap (for stationary data). We study theoretical properties of SDB as well as its numerical performance in simulated data and real data. Extending the underlying ideas of the second chapter, we introduce two new resampling strategies for big data in Chapter 3. The first strategy is called aggregation of little bootstraps or ALB, a generalized resampling technique that includes the SDB as a special case. The second strategy is called subsampled residual bootstrap or SRB, a fast version of residual bootstrap intended for massive regression models. We study both methods through simulations.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-11-09 without embargo terms","The student, Srijan Sengupta, accepted the attached license on 2016-07-06 at 13:49.","The student, Srijan Sengupta, submitted this Dissertation for approval on 2016-07-07 at 11:35.","This Dissertation was approved for publication on 2016-07-08 at 15:42.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9783 on 2016-11-09 at 10:22:53","Made available in DSpace on 2016-11-10T17:50:10Z (GMT). No. of bitstreams: 21 SENGUPTA-DISSERTATION-2016.pdf: 1339541 bytes, checksum: fd8fcc19303b1422bc8b98a52199d854 (MD5) ALB.tex: 16768 bytes, checksum: 31fac3ffab24a66daebe42a08b31c977 (MD5) DRW.tex: 70525 bytes, checksum: abbf3f7719a8a1869320debda6bee1e5 (MD5) HetClus.tex: 178539 bytes, checksum: f2d7d9e8b8683e6348a26fcda3bb6d8a (MD5) HetClusfiles.zip: 49530 bytes, checksum: 1370470c47d37e7d68e620423ead0b77 (MD5) IIDplots.zip: 140921 bytes, checksum: 644eeee7ba5e9c20605eb750e7051755 (MD5) PABM2.tex: 86864 bytes, checksum: 581f9673b9b3b7421cf14fc203dcc109 (MD5) SDB_RB.tex: 23243 bytes, checksum: 833b493a691f4ad3c456543385f14403 (MD5) SDB_revised.tex: 110034 bytes, checksum: c7e5795a012af4980c2554d512c39e2b (MD5) SDB_supp.tex: 10472 bytes, checksum: b5cd1121f2a22adfe9d754b255305d88 (MD5) SRB1110dgf2d10.eps: 8281 bytes, checksum: 990af1ef1ba8010d89f3ada2f2fd910a (MD5) abs.tex: 4233 bytes, checksum: 57c4a961c26fd363b81806a5489e1ae8 (MD5) ack.tex: 637 bytes, checksum: 738e17d866b7dc5d440f9626012ce68c (MD5) images.zip: 96671 bytes, checksum: 673a84c728501c33a7b12a021b27a89a (MD5) intro1.tex: 4606 bytes, checksum: c156fec92c20c3f057755c0ecbf03515 (MD5) intro2.tex: 5012 bytes, checksum: 79486107430e548d83768b2e20cf0de0 (MD5) thesis.tex: 8794 bytes, checksum: fcf2f2fe79afcf30f1144bb9a99bb997 (MD5) thesisrefs.bib: 119495 bytes, checksum: de7bb4421dc541de48d8d78279844101 (MD5) tsplots.zip: 1491356 bytes, checksum: 74c57e19cd5342054bafc4e2fb4ba886 (MD5) LICENSE.txt: 4212 bytes, checksum: a463d9bc00f8fbf7bb751cddb3790ad2 (MD5) PROQUEST_LICENSE.txt: 4558 bytes, checksum: 1a745ae8500a0c2b1af0656b4307e113 (MD5) Previous issue date: 2016-07-08"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/92763"],"dc:language":["en"],"dc:rights":["Copyright 2016 Srijan Sengupta"],"dc:subject":["network data","resampling","community structure","big data"],"dc:title":["Statistical analysis of networks with community structure and bootstrap methods for big data"],"dc:type":["text"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:35Z"}