{"id":{"repo_id":"purdue-thes","oai_identifier":"oai:docs.lib.purdue.edu:open_access_dissertations-1318"},"canonical_url":"https://search.dev.ndltd.org/etd/purdue-thes/oai:docs.lib.purdue.edu:open_access_dissertations-1318","repository":{"repo_id":"purdue-thes","name":"Purdue University","base_url":"https://docs.lib.purdue.edu/do/oai/"},"display":{"title":"Divide and recombine for large complex data: The subset likelihood modeling approach to recombination","abstract":"<p>Divide and recombine (D&R) is a statistical framework for the analysis of large complex data. The data are divided into subsets. Numeric and visualization methods, which collectively are analytic methods, are applied to each subset. For each analytic method, the outputs of the application of the method to the subsets are recombined. So each analytic method has associated with it a division method and a recombination method. Here we study D&R methods for likelihood-based model fitting. We introduce a notion of likelihood analysis and modeling. We divide the data and fit a likelihood model on each subset. The fitted model is characterized by a set of parameters much smaller than the subset data size, but retains as much information as possible about the true subset likelihood. Analysis of subset likelihoods and their fitted models consists of visualizations on an appropriate scale and region. These visualizations allow the analyst to verify the choice and fit of the model. The fitted models are recombined across subsets to form a model of the the all-data likelihood, which we maximize to obtain a likelihood modeling estimate (LME). We present simulation results demonstrating the performance of our method compared with the all-data maximum likelihood estimate (MLE) for the case of logistic regression.</p>","abstract_html":"&lt;p&gt;Divide and recombine (D&amp;R) is a statistical framework for the analysis of large complex data. The data are divided into subsets. Numeric and visualization methods, which collectively are analytic methods, are applied to each subset. For each analytic method, the outputs of the application of the method to the subsets are recombined. So each analytic method has associated with it a division method and a recombination method. Here we study D&amp;R methods for likelihood-based model fitting. We introduce a notion of likelihood analysis and modeling. We divide the data and fit a likelihood model on each subset. The fitted model is characterized by a set of parameters much smaller than the subset data size, but retains as much information as possible about the true subset likelihood. Analysis of subset likelihoods and their fitted models consists of visualizations on an appropriate scale and region. These visualizations allow the analyst to verify the choice and fit of the model. The fitted models are recombined across subsets to form a model of the the all-data likelihood, which we maximize to obtain a likelihood modeling estimate (LME). We present simulation results demonstrating the performance of our method compared with the all-data maximum likelihood estimate (MLE) for the case of logistic regression.&lt;/p&gt;","abstract_has_math":false,"creators":["Gautier, Philip"],"institution":null,"degree_name":"Doctor of Philosophy (PhD)","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["William S. Cleveland","Chuanhai Liu","Bowei Xi","Lingsong Zhang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-04-01T07:00:00Z","date_published":"2015-04-01T07:00:00Z","updated_at":"2026-07-24T03:53:28Z","subjects":["Statistics and Probability"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://docs.lib.purdue.edu/open_access_dissertations/458","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["William S. Cleveland","Chuanhai Liu","Bowei Xi","Lingsong Zhang"]},{"key":"dc:creator","label":"Author","values":["Gautier, Philip"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Statistics and Probability"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://docs.lib.purdue.edu/open_access_dissertations/458"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Divide and recombine (D&R) is a statistical framework for the analysis of large complex data. The data are divided into subsets. Numeric and visualization methods, which collectively are analytic methods, are applied to each subset. For each analytic method, the outputs of the application of the method to the subsets are recombined. So each analytic method has associated with it a division method and a recombination method. Here we study D&R methods for likelihood-based model fitting. We introduce a notion of likelihood analysis and modeling. We divide the data and fit a likelihood model on each subset. The fitted model is characterized by a set of parameters much smaller than the subset data size, but retains as much information as possible about the true subset likelihood. Analysis of subset likelihoods and their fitted models consists of visualizations on an appropriate scale and region. These visualizations allow the analyst to verify the choice and fit of the model. The fitted models are recombined across subsets to form a model of the the all-data likelihood, which we maximize to obtain a likelihood modeling estimate (LME). We present simulation results demonstrating the performance of our method compared with the all-data maximum likelihood estimate (MLE) for the case of logistic regression.</p>"]},{"key":"dc:title","label":"Title","values":["Divide and recombine for large complex data: The subset likelihood modeling approach to recombination"]}]}],"canonical_facts":{"dc:contributor":["William S. Cleveland","Chuanhai Liu","Bowei Xi","Lingsong Zhang"],"dc:creator":["Gautier, Philip"],"dc:description.abstract":["<p>Divide and recombine (D&R) is a statistical framework for the analysis of large complex data. The data are divided into subsets. Numeric and visualization methods, which collectively are analytic methods, are applied to each subset. For each analytic method, the outputs of the application of the method to the subsets are recombined. So each analytic method has associated with it a division method and a recombination method. Here we study D&R methods for likelihood-based model fitting. We introduce a notion of likelihood analysis and modeling. We divide the data and fit a likelihood model on each subset. The fitted model is characterized by a set of parameters much smaller than the subset data size, but retains as much information as possible about the true subset likelihood. Analysis of subset likelihoods and their fitted models consists of visualizations on an appropriate scale and region. These visualizations allow the analyst to verify the choice and fit of the model. The fitted models are recombined across subsets to form a model of the the all-data likelihood, which we maximize to obtain a likelihood modeling estimate (LME). We present simulation results demonstrating the performance of our method compared with the all-data maximum likelihood estimate (MLE) for the case of logistic regression.</p>"],"dc:identifier":["https://docs.lib.purdue.edu/open_access_dissertations/458"],"dc:subject":["Statistics and Probability"],"dc:title":["Divide and recombine for large complex data: The subset likelihood modeling approach to recombination"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-24T03:53:28Z"}