{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/137278"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/137278","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Recent Advances on Statistical Network Analysis and Multi-task Learning for Complex Data","abstract":"Real-world data are increasingly complex, arising from diverse domains that demand sophisticated analytical approaches. This dissertation focuses on developing advanced statistical methods to address the challenges in analyzing social network data and clinical trial data. First, I propose a novel exponential random graph model (ERGM) to study the common knowledge (CK) phenomenon in Facebook social networks. Unlike traditional contagion models, CK allows individuals to coordinate their activation as a group, thereby facilitating both the initiation and propagation of information. To investigate how network structure influences CK-based contagion, I develop an ERGM to generate networks while controlling for bicliques, which are the characterizing graph substructures for generating CK. Second, according to FDA guidance, prognostic variables—baseline covariates associated with clinical trial study outcomes—must be pre-specified at the study design stage to improve the precision of treatment effect estimation. To support this, I develop a multi-task learning approach that leverages historical trials of the treatment being studied to identify prognostic variables, which can guide the design and analysis of new studies. The performance is validated through simulations and demonstrated using real-world clinical trial data. In addition, I propose a frequentist dynamic borrowing approach that borrows information from the control arms of historical trials similar to the current study. This approach augments the control arm, improving the precision of estimating the treatment effect and the efficiency of conducting randomized controlled trials.","abstract_html":"Real-world data are increasingly complex, arising from diverse domains that demand sophisticated analytical approaches. This dissertation focuses on developing advanced statistical methods to address the challenges in analyzing social network data and clinical trial data. First, I propose a novel exponential random graph model (ERGM) to study the common knowledge (CK) phenomenon in Facebook social networks. Unlike traditional contagion models, CK allows individuals to coordinate their activation as a group, thereby facilitating both the initiation and propagation of information. To investigate how network structure influences CK-based contagion, I develop an ERGM to generate networks while controlling for bicliques, which are the characterizing graph substructures for generating CK. Second, according to FDA guidance, prognostic variables—baseline covariates associated with clinical trial study outcomes—must be pre-specified at the study design stage to improve the precision of treatment effect estimation. To support this, I develop a multi-task learning approach that leverages historical trials of the treatment being studied to identify prognostic variables, which can guide the design and analysis of new studies. The performance is validated through simulations and demonstrated using real-world clinical trial data. In addition, I propose a frequentist dynamic borrowing approach that borrows information from the control arms of historical trials similar to the current study. This approach augments the control arm, improving the precision of estimating the treatment effect and the efficiency of conducting randomized controlled trials.","abstract_has_math":false,"creators":["Liu, Xueying"],"institution":"Virginia Tech","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"Statistics","degree_department":"Statistics","school":null,"contributors":[],"advisors":[],"committee_chairs":["Deng, Xinwei"],"committee_members":["Du, Pang","Hong, Yili","Liu, Meimei","Kuhlman, Christopher James"],"year":2025,"date_issued":"2025-08-08","date_published":"2025-08-08","updated_at":"2026-07-22T22:19:24Z","subjects":["Social network data","Clinical trial data","Multi-task learning","Dynamic borrowing"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44421"],"render_values":[{"text":"vt_gsexam:44421","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/137278","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Deng, Xinwei"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Du, Pang","Hong, Yili","Liu, Meimei","Kuhlman, Christopher James"]},{"key":"dc:contributor.department","label":"Department","values":["Statistics"]},{"key":"dc:creator","label":"Author","values":["Liu, Xueying"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-08-09T08:00:17Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-08-09T08:00:17Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-08-08"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Social network data","Clinical trial data","Multi-task learning","Dynamic borrowing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44421"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/137278"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Real-world data are increasingly complex, arising from diverse domains that demand sophisticated analytical approaches. This dissertation focuses on developing advanced statistical methods to address the challenges in analyzing social network data and clinical trial data. First, I propose a novel exponential random graph model (ERGM) to study the common knowledge (CK) phenomenon in Facebook social networks. Unlike traditional contagion models, CK allows individuals to coordinate their activation as a group, thereby facilitating both the initiation and propagation of information. To investigate how network structure influences CK-based contagion, I develop an ERGM to generate networks while controlling for bicliques, which are the characterizing graph substructures for generating CK. Second, according to FDA guidance, prognostic variables—baseline covariates associated with clinical trial study outcomes—must be pre-specified at the study design stage to improve the precision of treatment effect estimation. To support this, I develop a multi-task learning approach that leverages historical trials of the treatment being studied to identify prognostic variables, which can guide the design and analysis of new studies. The performance is validated through simulations and demonstrated using real-world clinical trial data. In addition, I propose a frequentist dynamic borrowing approach that borrows information from the control arms of historical trials similar to the current study. This approach augments the control arm, improving the precision of estimating the treatment effect and the efficiency of conducting randomized controlled trials."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Advanced statistical methods are very useful to handle the complexity of real-world data. First, in social networks, common knowledge (CK) is a phenomenon where each individual within a group knows the same information and everyone knows that everyone knows the information, infinitely recursively. To investigate how network structure influences the emergence and spread of CK on Facebook networks, I develop a network-generating model that can control the number and size of bicliques, which are the characterizing substructures for generating CK. Second, in clinical trials, I leverage historical trials to help design and analyze new studies. According to FDA guidance, prognostic variables—baseline covariates associated with clinical trial study outcomes—must be pre-specified at the study design stage to improve the precision of treatment effect estimation. To support this, I propose an approach to select prognostic variables from multiple historical trials. These identified variables can then be used as stratification factors in new trial designs. In addition, I propose an approach to borrow information from the control arm of historical trials to augment the control arm in the current study, thereby improving study efficiency."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Recent Advances on Statistical Network Analysis and Multi-task Learning for Complex Data"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Deng, Xinwei"],"dc:contributor.committeemember":["Du, Pang","Hong, Yili","Liu, Meimei","Kuhlman, Christopher James"],"dc:contributor.department":["Statistics"],"dc:creator":["Liu, Xueying"],"dc:date.accessioned":["2025-08-09T08:00:17Z"],"dc:date.available":["2025-08-09T08:00:17Z"],"dc:date.issued":["2025-08-08"],"dc:description.abstract":["Real-world data are increasingly complex, arising from diverse domains that demand sophisticated analytical approaches. This dissertation focuses on developing advanced statistical methods to address the challenges in analyzing social network data and clinical trial data. First, I propose a novel exponential random graph model (ERGM) to study the common knowledge (CK) phenomenon in Facebook social networks. Unlike traditional contagion models, CK allows individuals to coordinate their activation as a group, thereby facilitating both the initiation and propagation of information. To investigate how network structure influences CK-based contagion, I develop an ERGM to generate networks while controlling for bicliques, which are the characterizing graph substructures for generating CK. Second, according to FDA guidance, prognostic variables—baseline covariates associated with clinical trial study outcomes—must be pre-specified at the study design stage to improve the precision of treatment effect estimation. To support this, I develop a multi-task learning approach that leverages historical trials of the treatment being studied to identify prognostic variables, which can guide the design and analysis of new studies. The performance is validated through simulations and demonstrated using real-world clinical trial data. In addition, I propose a frequentist dynamic borrowing approach that borrows information from the control arms of historical trials similar to the current study. This approach augments the control arm, improving the precision of estimating the treatment effect and the efficiency of conducting randomized controlled trials."],"dc:description.abstractgeneral":["Advanced statistical methods are very useful to handle the complexity of real-world data. First, in social networks, common knowledge (CK) is a phenomenon where each individual within a group knows the same information and everyone knows that everyone knows the information, infinitely recursively. To investigate how network structure influences the emergence and spread of CK on Facebook networks, I develop a network-generating model that can control the number and size of bicliques, which are the characterizing substructures for generating CK. Second, in clinical trials, I leverage historical trials to help design and analyze new studies. According to FDA guidance, prognostic variables—baseline covariates associated with clinical trial study outcomes—must be pre-specified at the study design stage to improve the precision of treatment effect estimation. To support this, I propose an approach to select prognostic variables from multiple historical trials. These identified variables can then be used as stratification factors in new trial designs. In addition, I propose an approach to borrow information from the control arm of historical trials to augment the control arm in the current study, thereby improving study efficiency."],"dc:description.degree":["Doctor of Philosophy"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:44421"],"dc:identifier.uri":["https://hdl.handle.net/10919/137278"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Social network data","Clinical trial data","Multi-task learning","Dynamic borrowing"],"dc:title":["Recent Advances on Statistical Network Analysis and Multi-task Learning for Complex Data"],"dc:type":["Dissertation"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:24Z"}