{"id":{"repo_id":"ttu","oai_identifier":"oai:ttu-ir.tdl.org:2346/90236"},"canonical_url":"https://search.dev.ndltd.org/etd/ttu/oai:ttu-ir.tdl.org:2346/90236","repository":{"repo_id":"ttu","name":"Texas Technology University","base_url":"https://ttu-ir.tdl.org/server/oai/request"},"display":{"title":"Data-Driven sequential decision making with learning under ambiguity","abstract":"Markov decision processes are often used to model sequential decision-making problems in uncertain dynamic environments, such as equipment maintenance and replacement problems, and inventory control problems. The objective of these problems is to find a policy or strategy, which is a prescription of which action to choose, to maximize a predetermined reward function. Determining an optimal policy for a Markov decision process model, first and foremost, requires the full knowledge of the reward and transition parameters, which are, however, fundamentally unknown in many applications. In conventional approach, these parameters are estimated from historical data, and treated as known quantities in the decision-making process. This approach ignores parameter uncertainty introduced by data inadequacy (e.g., limited availability of data, data contained with noise), making the robustness of decisions questionable. Two critical research questions arise in sequential decision making: 1) How can we incorporate the parameter uncertainty into the decision-making process to obtain more data-driven decisions? 2) What are the asymptotic behaviors of the decision-making process as more observations that reflect unknown parameters become available? This dissertation provides answers to these questions by developing novel data-driven sequential decision-making models in the presence of parameter uncertainty, and investigating the performance of the proposed decision-making models in maintenance and remanufacturing planning problems based on real-world cases. We first develop a novel distributionally robust Markov decision process model that incorporates a decision maker’s prior probabilistic information to model uncertain transition probabilities in a Bayesian framework. We first construct a set of prior distributions for unknown transition probabilities. With the evidence of historical data, the prior distributions in the ambiguity set are reevaluated using a likelihood-ratio test, and only admitted prior distributions are updated using the standard Bayesian updating rule. The objective of the proposed model is to maximize the expected total discounted reward over the worst posterior distribution. We investigate the asymptotic properties of the proposed multiple-priors model with a continuous support of prior distributions. We further develop an efficient approximation method to solve the resulting optimization model, and provide theoretical analysis on the performance of such an approximation. The utility of the proposed decision framework is demonstrated using a remanufacturing planning example. We then consider two sequential maintenance planning problems for both single- and multi-component systems. We first consider the problem of optimally maintaining a periodically inspected system with multi-level preventive maintenance whose effects are complex. We formulate the problem as an infinite-horizon Markov decision process with the objective of minimizing the expected total discounted inspection and maintenance cost. Sufficient conditions are established to ensure the existence of an optimal monotone control-limit type policy with respect to the system’s deterioration level and age. We further consider an integrated budget allocation and preventive maintenance optimization problem for multi-facility deteriorating transportation infrastructure systems, where preventive maintenance has multiple types and each maintenance type introduces complex maintenance effects. We formulate the problem as a sum of multiple Markov decision process models with budget constraints. A priority-based twostage method is developed to find optimal maintenance decisions for large-scale problems. Real-world deterioration data are used to demonstrate the effectiveness and efficiency of the proposed models and solution algorithms.","abstract_html":"Markov decision processes are often used to model sequential decision-making problems in uncertain dynamic environments, such as equipment maintenance and replacement problems, and inventory control problems. The objective of these problems is to find a policy or strategy, which is a prescription of which action to choose, to maximize a predetermined reward function. Determining an optimal policy for a Markov decision process model, first and foremost, requires the full knowledge of the reward and transition parameters, which are, however, fundamentally unknown in many applications. In conventional approach, these parameters are estimated from historical data, and treated as known quantities in the decision-making process. This approach ignores parameter uncertainty introduced by data inadequacy (e.g., limited availability of data, data contained with noise), making the robustness of decisions questionable. Two critical research questions arise in sequential decision making: 1) How can we incorporate the parameter uncertainty into the decision-making process to obtain more data-driven decisions? 2) What are the asymptotic behaviors of the decision-making process as more observations that reflect unknown parameters become available? This dissertation provides answers to these questions by developing novel data-driven sequential decision-making models in the presence of parameter uncertainty, and investigating the performance of the proposed decision-making models in maintenance and remanufacturing planning problems based on real-world cases. We first develop a novel distributionally robust Markov decision process model that incorporates a decision maker’s prior probabilistic information to model uncertain transition probabilities in a Bayesian framework. We first construct a set of prior distributions for unknown transition probabilities. With the evidence of historical data, the prior distributions in the ambiguity set are reevaluated using a likelihood-ratio test, and only admitted prior distributions are updated using the standard Bayesian updating rule. The objective of the proposed model is to maximize the expected total discounted reward over the worst posterior distribution. We investigate the asymptotic properties of the proposed multiple-priors model with a continuous support of prior distributions. We further develop an efficient approximation method to solve the resulting optimization model, and provide theoretical analysis on the performance of such an approximation. The utility of the proposed decision framework is demonstrated using a remanufacturing planning example. We then consider two sequential maintenance planning problems for both single- and multi-component systems. We first consider the problem of optimally maintaining a periodically inspected system with multi-level preventive maintenance whose effects are complex. We formulate the problem as an infinite-horizon Markov decision process with the objective of minimizing the expected total discounted inspection and maintenance cost. Sufficient conditions are established to ensure the existence of an optimal monotone control-limit type policy with respect to the system’s deterioration level and age. We further consider an integrated budget allocation and preventive maintenance optimization problem for multi-facility deteriorating transportation infrastructure systems, where preventive maintenance has multiple types and each maintenance type introduces complex maintenance effects. We formulate the problem as a sum of multiple Markov decision process models with budget constraints. A priority-based twostage method is developed to find optimal maintenance decisions for large-scale problems. Real-world deterioration data are used to demonstrate the effectiveness and efficiency of the proposed models and solution algorithms.","abstract_has_math":false,"creators":["Shi, Yue"],"institution":"Texas Tech University","degree_name":"Doctor of Philosophy","degree_level":"Doctoral","degree_discipline":"Industrial Engineering","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":["Xiang, Yisha"],"committee_members":["Du, Dongping","Matis, Timothy","Coit, David W."],"year":2022,"date_issued":"2022-08","date_published":"2022-08","updated_at":"2026-07-24T05:04:54Z","subjects":["Sequential Decision Making","Data-Driven Method","Bayesian Learning Method","Maintenance Optimization"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2346/90236","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Xiang, Yisha"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Du, Dongping","Matis, Timothy","Coit, David W."]},{"key":"dc:creator","label":"Author","values":["Shi, Yue"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2022-09-13T14:17:56Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2022-09-13T14:17:56Z"]},{"key":"dc:date.issued","label":"Date","values":["2022-08"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Industrial Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Texas Tech University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Sequential Decision Making","Data-Driven Method","Bayesian Learning Method","Maintenance Optimization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/2346/90236"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Markov decision processes are often used to model sequential decision-making problems in uncertain dynamic environments, such as equipment maintenance and replacement problems, and inventory control problems. The objective of these problems is to find a policy or strategy, which is a prescription of which action to choose, to maximize a predetermined reward function. Determining an optimal policy for a Markov decision process model, first and foremost, requires the full knowledge of the reward and transition parameters, which are, however, fundamentally unknown in many applications. In conventional approach, these parameters are estimated from historical data, and treated as known quantities in the decision-making process. This approach ignores parameter uncertainty introduced by data inadequacy (e.g., limited availability of data, data contained with noise), making the robustness of decisions questionable. Two critical research questions arise in sequential decision making: 1) How can we incorporate the parameter uncertainty into the decision-making process to obtain more data-driven decisions? 2) What are the asymptotic behaviors of the decision-making process as more observations that reflect unknown parameters become available? This dissertation provides answers to these questions by developing novel data-driven sequential decision-making models in the presence of parameter uncertainty, and investigating the performance of the proposed decision-making models in maintenance and remanufacturing planning problems based on real-world cases. We first develop a novel distributionally robust Markov decision process model that incorporates a decision maker’s prior probabilistic information to model uncertain transition probabilities in a Bayesian framework. We first construct a set of prior distributions for unknown transition probabilities. With the evidence of historical data, the prior distributions in the ambiguity set are reevaluated using a likelihood-ratio test, and only admitted prior distributions are updated using the standard Bayesian updating rule. The objective of the proposed model is to maximize the expected total discounted reward over the worst posterior distribution. We investigate the asymptotic properties of the proposed multiple-priors model with a continuous support of prior distributions. We further develop an efficient approximation method to solve the resulting optimization model, and provide theoretical analysis on the performance of such an approximation. The utility of the proposed decision framework is demonstrated using a remanufacturing planning example. We then consider two sequential maintenance planning problems for both single- and multi-component systems. We first consider the problem of optimally maintaining a periodically inspected system with multi-level preventive maintenance whose effects are complex. We formulate the problem as an infinite-horizon Markov decision process with the objective of minimizing the expected total discounted inspection and maintenance cost. Sufficient conditions are established to ensure the existence of an optimal monotone control-limit type policy with respect to the system’s deterioration level and age. We further consider an integrated budget allocation and preventive maintenance optimization problem for multi-facility deteriorating transportation infrastructure systems, where preventive maintenance has multiple types and each maintenance type introduces complex maintenance effects. We formulate the problem as a sum of multiple Markov decision process models with budget constraints. A priority-based twostage method is developed to find optimal maintenance decisions for large-scale problems. Real-world deterioration data are used to demonstrate the effectiveness and efficiency of the proposed models and solution algorithms.","Embargo status: Restricted until 09/2027. To request the author grant access, click on the PDF link to the left."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Data-Driven sequential decision making with learning under ambiguity"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Xiang, Yisha"],"dc:contributor.committeemember":["Du, Dongping","Matis, Timothy","Coit, David W."],"dc:creator":["Shi, Yue"],"dc:date.accessioned":["2022-09-13T14:17:56Z"],"dc:date.available":["2022-09-13T14:17:56Z"],"dc:date.issued":["2022-08"],"dc:description.abstract":["Markov decision processes are often used to model sequential decision-making problems in uncertain dynamic environments, such as equipment maintenance and replacement problems, and inventory control problems. The objective of these problems is to find a policy or strategy, which is a prescription of which action to choose, to maximize a predetermined reward function. Determining an optimal policy for a Markov decision process model, first and foremost, requires the full knowledge of the reward and transition parameters, which are, however, fundamentally unknown in many applications. In conventional approach, these parameters are estimated from historical data, and treated as known quantities in the decision-making process. This approach ignores parameter uncertainty introduced by data inadequacy (e.g., limited availability of data, data contained with noise), making the robustness of decisions questionable. Two critical research questions arise in sequential decision making: 1) How can we incorporate the parameter uncertainty into the decision-making process to obtain more data-driven decisions? 2) What are the asymptotic behaviors of the decision-making process as more observations that reflect unknown parameters become available? This dissertation provides answers to these questions by developing novel data-driven sequential decision-making models in the presence of parameter uncertainty, and investigating the performance of the proposed decision-making models in maintenance and remanufacturing planning problems based on real-world cases. We first develop a novel distributionally robust Markov decision process model that incorporates a decision maker’s prior probabilistic information to model uncertain transition probabilities in a Bayesian framework. We first construct a set of prior distributions for unknown transition probabilities. With the evidence of historical data, the prior distributions in the ambiguity set are reevaluated using a likelihood-ratio test, and only admitted prior distributions are updated using the standard Bayesian updating rule. The objective of the proposed model is to maximize the expected total discounted reward over the worst posterior distribution. We investigate the asymptotic properties of the proposed multiple-priors model with a continuous support of prior distributions. We further develop an efficient approximation method to solve the resulting optimization model, and provide theoretical analysis on the performance of such an approximation. The utility of the proposed decision framework is demonstrated using a remanufacturing planning example. We then consider two sequential maintenance planning problems for both single- and multi-component systems. We first consider the problem of optimally maintaining a periodically inspected system with multi-level preventive maintenance whose effects are complex. We formulate the problem as an infinite-horizon Markov decision process with the objective of minimizing the expected total discounted inspection and maintenance cost. Sufficient conditions are established to ensure the existence of an optimal monotone control-limit type policy with respect to the system’s deterioration level and age. We further consider an integrated budget allocation and preventive maintenance optimization problem for multi-facility deteriorating transportation infrastructure systems, where preventive maintenance has multiple types and each maintenance type introduces complex maintenance effects. We formulate the problem as a sum of multiple Markov decision process models with budget constraints. A priority-based twostage method is developed to find optimal maintenance decisions for large-scale problems. Real-world deterioration data are used to demonstrate the effectiveness and efficiency of the proposed models and solution algorithms.","Embargo status: Restricted until 09/2027. To request the author grant access, click on the PDF link to the left."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/2346/90236"],"dc:language.iso":["eng"],"dc:subject":["Sequential Decision Making","Data-Driven Method","Bayesian Learning Method","Maintenance Optimization"],"dc:title":["Data-Driven sequential decision making with learning under ambiguity"],"dc:type":["Dissertation"],"thesis:degree_discipline":["Industrial Engineering"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Texas Tech University"]},"updated_at":"2026-07-24T05:04:54Z"}