{"id":{"repo_id":"ku","oai_identifier":"oai:kuscholarworks.ku.edu:1808/36346"},"canonical_url":"https://search.dev.ndltd.org/etd/ku/oai:kuscholarworks.ku.edu:1808/36346","repository":{"repo_id":"ku","name":"University of Kansas","base_url":"https://kuscholarworks.ku.edu/server/oai/request"},"display":{"title":"Essays on Model Selection Uncertainty and Model Averaging: Computational and Empirical Work with Beta Regression, Multiple Linear Regression with ARMA Innovations, and the Minimum Description Length Principle","abstract":"Uncertainty in model selection is under-explored and frequently resolved non-rigorously through beliefs about generalizability, practical usefulness, and computational ease. This is problematic as model selection routinely admits multiple models which imposes extra uncertainty on all post-selection conclusions. This research emphasized integrated model averaging and selection methods to enhance the characterization and handling of model selection uncertainty. In Chapter 2, regression models based on the beta distribution were investigated from the standpoint of multi-model uncertainty and model averaging. A tool that can combine model selection uncertainty and beta regression modeling was developed. It combined bootstrap-based resampling, model averaging, model selection, and asymptotic theory to yield a procedure that can perform joint modeling of the mean and precision parameters, capture more sources of variability, and achieve more accurate claims of estimate precision, variable importance, and model stability. Utility of the tool was demonstrated through a study of model selection consistency and variable importance in clickstream data. In Chapter 3, model averaging to improve model qualityin the context of time series data was investigated. Serial dependence in the data demanded that the bootstrap-based procedure from Chapter 2 be modified to use subsampling as the resampling method. Time series analysis with multiple linear regression with ARMA innovation models was believed to benefit from model combination with the developed tool through diversification of individual model misspecification bias. Model combination also offered an exogenous variable and ARMA order ranking and selection routine that differed from conventional methods. The mean interval score was employed as a novel index to quantify prediction interval quality. Highly detailed simulations indicated that the model combination tool yielded superior or competitive prediction performance compared to alternative forecast techniques. Utility was exhibited in the modeling of national labor force participation rate dynamics as a function of social and demographic variables. In Chapter 4, concerns about individual model selection criteria were investigated through the lens of information and coding theory, specifically, the minimum description length principle. This principle combines aspects of the principle of parsimony and goodness-of-fit to automatically control overfitting and model generalization. One- and two-part minimum description length selection criteria for use in beta regression model selection were derived. The two-part criterion enabled simultaneous selection and parameter estimation of mean and dispersion sub-models through a novel formula. The one-part criterion encoded structural complexity of beta regression models into model selection which was separate from complexity from the number of estimable parameters. Monte Carlo integration with importance sampling was shown to help approximate the one-part criterion. Simulations assessed model selection under various beta regression settings for the two-part criterion. Results indicated that the two-part criterion was competitive with conventional criteria like AIC with respect to correct variable inclusion, L1 norm of the estimated coefficients, and out-of-sample prediction quality measures. The MDL criterion surpassed the conventional criteria in that it allowed greater control over model selection through adjustable parameters.","abstract_html":"Uncertainty in model selection is under-explored and frequently resolved non-rigorously through beliefs about generalizability, practical usefulness, and computational ease. This is problematic as model selection routinely admits multiple models which imposes extra uncertainty on all post-selection conclusions. This research emphasized integrated model averaging and selection methods to enhance the characterization and handling of model selection uncertainty. In Chapter 2, regression models based on the beta distribution were investigated from the standpoint of multi-model uncertainty and model averaging. A tool that can combine model selection uncertainty and beta regression modeling was developed. It combined bootstrap-based resampling, model averaging, model selection, and asymptotic theory to yield a procedure that can perform joint modeling of the mean and precision parameters, capture more sources of variability, and achieve more accurate claims of estimate precision, variable importance, and model stability. Utility of the tool was demonstrated through a study of model selection consistency and variable importance in clickstream data. In Chapter 3, model averaging to improve model qualityin the context of time series data was investigated. Serial dependence in the data demanded that the bootstrap-based procedure from Chapter 2 be modified to use subsampling as the resampling method. Time series analysis with multiple linear regression with ARMA innovation models was believed to benefit from model combination with the developed tool through diversification of individual model misspecification bias. Model combination also offered an exogenous variable and ARMA order ranking and selection routine that differed from conventional methods. The mean interval score was employed as a novel index to quantify prediction interval quality. Highly detailed simulations indicated that the model combination tool yielded superior or competitive prediction performance compared to alternative forecast techniques. Utility was exhibited in the modeling of national labor force participation rate dynamics as a function of social and demographic variables. In Chapter 4, concerns about individual model selection criteria were investigated through the lens of information and coding theory, specifically, the minimum description length principle. This principle combines aspects of the principle of parsimony and goodness-of-fit to automatically control overfitting and model generalization. One- and two-part minimum description length selection criteria for use in beta regression model selection were derived. The two-part criterion enabled simultaneous selection and parameter estimation of mean and dispersion sub-models through a novel formula. The one-part criterion encoded structural complexity of beta regression models into model selection which was separate from complexity from the number of estimable parameters. Monte Carlo integration with importance sampling was shown to help approximate the one-part criterion. Simulations assessed model selection under various beta regression settings for the two-part criterion. Results indicated that the two-part criterion was competitive with conventional criteria like AIC with respect to correct variable inclusion, L1 norm of the estimated coefficients, and out-of-sample prediction quality measures. The MDL criterion surpassed the conventional criteria in that it allowed greater control over model selection through adjustable parameters.","abstract_has_math":false,"creators":["Allenbrand, Corban"],"institution":"University of Kansas","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Sherwood, Ben"],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-05-31","date_published":"2023-05-31","updated_at":"2026-07-24T02:47:03Z","subjects":["Applied mathematics","Information science","Operations research","Applied Statistics","Bootstrap and Subsampling","Computational Statistics","Information Theory and Statistics","Model Combination","Time Series Analysis"],"languages":["en"],"rights":["Copyright held by the author."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["http://dissertations.umi.com/ku:18928"],"render_values":[{"text":"http://dissertations.umi.com/ku:18928","href":"http://dissertations.umi.com/ku:18928","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/1808/36346","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Sherwood, Ben"]},{"key":"dc:creator","label":"Author","values":["Allenbrand, Corban"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-02-25T21:56:29Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-02-25T21:56:29Z"]},{"key":"dc:date.issued","label":"Date","values":["2023-05-31"]},{"key":"dc:publisher","label":"Institution","values":["University of Kansas"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Applied mathematics","Information science","Operations research","Applied Statistics","Bootstrap and Subsampling","Computational Statistics","Information Theory and Statistics","Model Combination","Time Series Analysis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright held by the author."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["http://dissertations.umi.com/ku:18928"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1808/36346"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This item is protected by copyright and unless otherwise specified the copyright of this thesis/dissertation is held by the author."]},{"key":"dc:description.abstract","label":"Abstract","values":["Uncertainty in model selection is under-explored and frequently resolved non-rigorously through beliefs about generalizability, practical usefulness, and computational ease. This is problematic as model selection routinely admits multiple models which imposes extra uncertainty on all post-selection conclusions. This research emphasized integrated model averaging and selection methods to enhance the characterization and handling of model selection uncertainty. In Chapter 2, regression models based on the beta distribution were investigated from the standpoint of multi-model uncertainty and model averaging. A tool that can combine model selection uncertainty and beta regression modeling was developed. It combined bootstrap-based resampling, model averaging, model selection, and asymptotic theory to yield a procedure that can perform joint modeling of the mean and precision parameters, capture more sources of variability, and achieve more accurate claims of estimate precision, variable importance, and model stability. Utility of the tool was demonstrated through a study of model selection consistency and variable importance in clickstream data. In Chapter 3, model averaging to improve model qualityin the context of time series data was investigated. Serial dependence in the data demanded that the bootstrap-based procedure from Chapter 2 be modified to use subsampling as the resampling method. Time series analysis with multiple linear regression with ARMA innovation models was believed to benefit from model combination with the developed tool through diversification of individual model misspecification bias. Model combination also offered an exogenous variable and ARMA order ranking and selection routine that differed from conventional methods. The mean interval score was employed as a novel index to quantify prediction interval quality. Highly detailed simulations indicated that the model combination tool yielded superior or competitive prediction performance compared to alternative forecast techniques. Utility was exhibited in the modeling of national labor force participation rate dynamics as a function of social and demographic variables. In Chapter 4, concerns about individual model selection criteria were investigated through the lens of information and coding theory, specifically, the minimum description length principle. This principle combines aspects of the principle of parsimony and goodness-of-fit to automatically control overfitting and model generalization. One- and two-part minimum description length selection criteria for use in beta regression model selection were derived. The two-part criterion enabled simultaneous selection and parameter estimation of mean and dispersion sub-models through a novel formula. The one-part criterion encoded structural complexity of beta regression models into model selection which was separate from complexity from the number of estimable parameters. Monte Carlo integration with importance sampling was shown to help approximate the one-part criterion. Simulations assessed model selection under various beta regression settings for the two-part criterion. Results indicated that the two-part criterion was competitive with conventional criteria like AIC with respect to correct variable inclusion, L1 norm of the estimated coefficients, and out-of-sample prediction quality measures. The MDL criterion surpassed the conventional criteria in that it allowed greater control over model selection through adjustable parameters."]},{"key":"dc:title","label":"Title","values":["Essays on Model Selection Uncertainty and Model Averaging: Computational and Empirical Work with Beta Regression, Multiple Linear Regression with ARMA Innovations, and the Minimum Description Length Principle"]}]}],"canonical_facts":{"dc:contributor.advisor":["Sherwood, Ben"],"dc:creator":["Allenbrand, Corban"],"dc:date.accessioned":["2026-02-25T21:56:29Z"],"dc:date.available":["2026-02-25T21:56:29Z"],"dc:date.issued":["2023-05-31"],"dc:description":["This item is protected by copyright and unless otherwise specified the copyright of this thesis/dissertation is held by the author."],"dc:description.abstract":["Uncertainty in model selection is under-explored and frequently resolved non-rigorously through beliefs about generalizability, practical usefulness, and computational ease. This is problematic as model selection routinely admits multiple models which imposes extra uncertainty on all post-selection conclusions. This research emphasized integrated model averaging and selection methods to enhance the characterization and handling of model selection uncertainty. In Chapter 2, regression models based on the beta distribution were investigated from the standpoint of multi-model uncertainty and model averaging. A tool that can combine model selection uncertainty and beta regression modeling was developed. It combined bootstrap-based resampling, model averaging, model selection, and asymptotic theory to yield a procedure that can perform joint modeling of the mean and precision parameters, capture more sources of variability, and achieve more accurate claims of estimate precision, variable importance, and model stability. Utility of the tool was demonstrated through a study of model selection consistency and variable importance in clickstream data. In Chapter 3, model averaging to improve model qualityin the context of time series data was investigated. Serial dependence in the data demanded that the bootstrap-based procedure from Chapter 2 be modified to use subsampling as the resampling method. Time series analysis with multiple linear regression with ARMA innovation models was believed to benefit from model combination with the developed tool through diversification of individual model misspecification bias. Model combination also offered an exogenous variable and ARMA order ranking and selection routine that differed from conventional methods. The mean interval score was employed as a novel index to quantify prediction interval quality. Highly detailed simulations indicated that the model combination tool yielded superior or competitive prediction performance compared to alternative forecast techniques. Utility was exhibited in the modeling of national labor force participation rate dynamics as a function of social and demographic variables. In Chapter 4, concerns about individual model selection criteria were investigated through the lens of information and coding theory, specifically, the minimum description length principle. This principle combines aspects of the principle of parsimony and goodness-of-fit to automatically control overfitting and model generalization. One- and two-part minimum description length selection criteria for use in beta regression model selection were derived. The two-part criterion enabled simultaneous selection and parameter estimation of mean and dispersion sub-models through a novel formula. The one-part criterion encoded structural complexity of beta regression models into model selection which was separate from complexity from the number of estimable parameters. Monte Carlo integration with importance sampling was shown to help approximate the one-part criterion. Simulations assessed model selection under various beta regression settings for the two-part criterion. Results indicated that the two-part criterion was competitive with conventional criteria like AIC with respect to correct variable inclusion, L1 norm of the estimated coefficients, and out-of-sample prediction quality measures. The MDL criterion surpassed the conventional criteria in that it allowed greater control over model selection through adjustable parameters."],"dc:identifier.other":["http://dissertations.umi.com/ku:18928"],"dc:identifier.uri":["https://hdl.handle.net/1808/36346"],"dc:language.iso":["en"],"dc:publisher":["University of Kansas"],"dc:rights":["Copyright held by the author."],"dc:subject":["Applied mathematics","Information science","Operations research","Applied Statistics","Bootstrap and Subsampling","Computational Statistics","Information Theory and Statistics","Model Combination","Time Series Analysis"],"dc:title":["Essays on Model Selection Uncertainty and Model Averaging: Computational and Empirical Work with Beta Regression, Multiple Linear Regression with ARMA Innovations, and the Minimum Description Length Principle"],"dc:type":["Dissertation"]},"updated_at":"2026-07-24T02:47:03Z"}