{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/388332"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/388332","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Dynamic Discrete Choice Estimation using Reinforcement Learning with Applications in Online Food Markets","abstract":"Dynamic Discrete Choice (DDC) models are widely used to analyze sequential decision-making in economics and marketing. However, their estimation remains computationally challenging, especially as state spaces expand, limiting their application to large-scale consumer datasets. This thesis develops Reinforcement Learning (RL)-based estimation methods to improve the speed and scalability of DDC estimation. The second chapter establishes a theoretical foundation for integrating RL with DDC estimation, emphasizing the shared mathematical structure of Markov Decision Processes (MDPs) in both frameworks. It reviews key tabular RL methods including Dynamic Programming (DP), Monte Carlo (MC), Temporal Difference (TD) and Q-learning, and draws conceptual parallels with DDC methods including Nested Fixed-Point Algorithm (NFXP), Conditional Choice Simulation (CCS), and Nested Pseudo-Likelihood (NPL). The chapter discusses how incorporating both tabular RL methods and those using function or policy approximation into the DDC estimation process can reduce computation time and improve scalability in high-dimensional settings. It also discusses additional RL approaches such as Prioritized Sweeping (PS), Inverse Reinforcement Learning (IRL), and Learning-to-Optimize (L2O) as potential tools for further improving computational efficiency. Building on this, the third chapter introduces the Reinforcement Learning Temporal Difference based Condition Choice Simulation (RLTD-CCS) algorithm. The algorithm leverages forward simulations (CCS) and stepwise TD learning to update value functions more efficiently than CCS. Monte Carlo studies on a small state space machine replacement model and a large state space prototype food choice model confirm that RLTD-CCS achieves estimation accuracy comparable to CCS while being up to 14 times faster, making it a viable approach for scalable DDC estimation. Additionally, in the fourth chapter, RLTD-CCS is embedded within the Expectation-Maximization (EM) algorithm to accommodate models with time-invariant persistent unobservables. In the fifth chapter, the proposed RL-based approach is applied to real-world consumer decision-making using weekly purchase data from a UK recipe box provider. The empirical analysis explores how nutritional attributes (protein and carbohydrate content) and habit persistence shape recipe box selection over time. Results indicate that nutritional factors are the primary drivers of choice, with consumers strongly preferring high-protein meals while avoiding carbohydrate-heavy options. A time-segmented static analysis further reveals a post-COVID shift: demand for high-protein meals increased, while carbohydrate aversion intensified among loyal consumers, suggesting lasting changes in dietary preferences. Counterfactual simulations assess the impact of nutritional composition adjustments and habit persistence modifications, offering actionable insights for business strategy and consumer retention. Based on these findings, recommendations are provided for optimizing menu offerings, encouraging repeat purchases, and improving personalization strategies. This thesis advances the integration of machine learning and structural econometrics by demonstrating the potential of RL-based methods in DDC estimation. By highlighting opportunities to further incorporate RL into DDC modeling, it paves the way for more scalable and computationally efficient approaches to estimating complex structural choice models.","abstract_html":"Dynamic Discrete Choice (DDC) models are widely used to analyze sequential decision-making in economics and marketing. However, their estimation remains computationally challenging, especially as state spaces expand, limiting their application to large-scale consumer datasets. This thesis develops Reinforcement Learning (RL)-based estimation methods to improve the speed and scalability of DDC estimation. The second chapter establishes a theoretical foundation for integrating RL with DDC estimation, emphasizing the shared mathematical structure of Markov Decision Processes (MDPs) in both frameworks. It reviews key tabular RL methods including Dynamic Programming (DP), Monte Carlo (MC), Temporal Difference (TD) and Q-learning, and draws conceptual parallels with DDC methods including Nested Fixed-Point Algorithm (NFXP), Conditional Choice Simulation (CCS), and Nested Pseudo-Likelihood (NPL). The chapter discusses how incorporating both tabular RL methods and those using function or policy approximation into the DDC estimation process can reduce computation time and improve scalability in high-dimensional settings. It also discusses additional RL approaches such as Prioritized Sweeping (PS), Inverse Reinforcement Learning (IRL), and Learning-to-Optimize (L2O) as potential tools for further improving computational efficiency. Building on this, the third chapter introduces the Reinforcement Learning Temporal Difference based Condition Choice Simulation (RLTD-CCS) algorithm. The algorithm leverages forward simulations (CCS) and stepwise TD learning to update value functions more efficiently than CCS. Monte Carlo studies on a small state space machine replacement model and a large state space prototype food choice model confirm that RLTD-CCS achieves estimation accuracy comparable to CCS while being up to 14 times faster, making it a viable approach for scalable DDC estimation. Additionally, in the fourth chapter, RLTD-CCS is embedded within the Expectation-Maximization (EM) algorithm to accommodate models with time-invariant persistent unobservables. In the fifth chapter, the proposed RL-based approach is applied to real-world consumer decision-making using weekly purchase data from a UK recipe box provider. The empirical analysis explores how nutritional attributes (protein and carbohydrate content) and habit persistence shape recipe box selection over time. Results indicate that nutritional factors are the primary drivers of choice, with consumers strongly preferring high-protein meals while avoiding carbohydrate-heavy options. A time-segmented static analysis further reveals a post-COVID shift: demand for high-protein meals increased, while carbohydrate aversion intensified among loyal consumers, suggesting lasting changes in dietary preferences. Counterfactual simulations assess the impact of nutritional composition adjustments and habit persistence modifications, offering actionable insights for business strategy and consumer retention. Based on these findings, recommendations are provided for optimizing menu offerings, encouraging repeat purchases, and improving personalization strategies. This thesis advances the integration of machine learning and structural econometrics by demonstrating the potential of RL-based methods in DDC estimation. By highlighting opportunities to further incorporate RL into DDC modeling, it paves the way for more scalable and computationally efficient approaches to estimating complex structural choice models.","abstract_has_math":false,"creators":["Srivastava, Sonal"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Prabhu, Jaideep"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-02-18","date_published":"2025-02-18","updated_at":"2026-07-22T22:24:31Z","subjects":["Conditional Choice Simulation","Dynamic Discrete Choice","Food Choice Modeling","Reinforcement Learning","Two-step Estimation"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/cd66ceb7-09e4-437e-9c58-b3ca7d2e5ad9/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.120730","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Prabhu, Jaideep"]},{"key":"dc:creator","label":"Author","values":["Srivastava, Sonal"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-02-18"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/388332"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Conditional Choice Simulation","Dynamic Discrete Choice","Food Choice Modeling","Reinforcement Learning","Two-step Estimation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/cd66ceb7-09e4-437e-9c58-b3ca7d2e5ad9/download","http://purl.org/NET/rdflicense/allrightsreserved"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.120730"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/b7a44324-428d-47bb-81bd-b3d784e1223d/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Dynamic Discrete Choice (DDC) models are widely used to analyze sequential decision-making in economics and marketing. However, their estimation remains computationally challenging, especially as state spaces expand, limiting their application to large-scale consumer datasets. This thesis develops Reinforcement Learning (RL)-based estimation methods to improve the speed and scalability of DDC estimation. The second chapter establishes a theoretical foundation for integrating RL with DDC estimation, emphasizing the shared mathematical structure of Markov Decision Processes (MDPs) in both frameworks. It reviews key tabular RL methods including Dynamic Programming (DP), Monte Carlo (MC), Temporal Difference (TD) and Q-learning, and draws conceptual parallels with DDC methods including Nested Fixed-Point Algorithm (NFXP), Conditional Choice Simulation (CCS), and Nested Pseudo-Likelihood (NPL). The chapter discusses how incorporating both tabular RL methods and those using function or policy approximation into the DDC estimation process can reduce computation time and improve scalability in high-dimensional settings. It also discusses additional RL approaches such as Prioritized Sweeping (PS), Inverse Reinforcement Learning (IRL), and Learning-to-Optimize (L2O) as potential tools for further improving computational efficiency. Building on this, the third chapter introduces the Reinforcement Learning Temporal Difference based Condition Choice Simulation (RLTD-CCS) algorithm. The algorithm leverages forward simulations (CCS) and stepwise TD learning to update value functions more efficiently than CCS. Monte Carlo studies on a small state space machine replacement model and a large state space prototype food choice model confirm that RLTD-CCS achieves estimation accuracy comparable to CCS while being up to 14 times faster, making it a viable approach for scalable DDC estimation. Additionally, in the fourth chapter, RLTD-CCS is embedded within the Expectation-Maximization (EM) algorithm to accommodate models with time-invariant persistent unobservables. In the fifth chapter, the proposed RL-based approach is applied to real-world consumer decision-making using weekly purchase data from a UK recipe box provider. The empirical analysis explores how nutritional attributes (protein and carbohydrate content) and habit persistence shape recipe box selection over time. Results indicate that nutritional factors are the primary drivers of choice, with consumers strongly preferring high-protein meals while avoiding carbohydrate-heavy options. A time-segmented static analysis further reveals a post-COVID shift: demand for high-protein meals increased, while carbohydrate aversion intensified among loyal consumers, suggesting lasting changes in dietary preferences. Counterfactual simulations assess the impact of nutritional composition adjustments and habit persistence modifications, offering actionable insights for business strategy and consumer retention. Based on these findings, recommendations are provided for optimizing menu offerings, encouraging repeat purchases, and improving personalization strategies. This thesis advances the integration of machine learning and structural econometrics by demonstrating the potential of RL-based methods in DDC estimation. By highlighting opportunities to further incorporate RL into DDC modeling, it paves the way for more scalable and computationally efficient approaches to estimating complex structural choice models."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["d16a4f8aefb01f10c543e4dc95aba817","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Dynamic Discrete Choice Estimation using Reinforcement Learning with Applications in Online Food Markets"]}]}],"canonical_facts":{"dc:contributor.advisor":["Prabhu, Jaideep"],"dc:creator":["Srivastava, Sonal"],"dc:date.issued":["2025-02-18"],"dc:description.abstract":["Dynamic Discrete Choice (DDC) models are widely used to analyze sequential decision-making in economics and marketing. However, their estimation remains computationally challenging, especially as state spaces expand, limiting their application to large-scale consumer datasets. This thesis develops Reinforcement Learning (RL)-based estimation methods to improve the speed and scalability of DDC estimation. The second chapter establishes a theoretical foundation for integrating RL with DDC estimation, emphasizing the shared mathematical structure of Markov Decision Processes (MDPs) in both frameworks. It reviews key tabular RL methods including Dynamic Programming (DP), Monte Carlo (MC), Temporal Difference (TD) and Q-learning, and draws conceptual parallels with DDC methods including Nested Fixed-Point Algorithm (NFXP), Conditional Choice Simulation (CCS), and Nested Pseudo-Likelihood (NPL). The chapter discusses how incorporating both tabular RL methods and those using function or policy approximation into the DDC estimation process can reduce computation time and improve scalability in high-dimensional settings. It also discusses additional RL approaches such as Prioritized Sweeping (PS), Inverse Reinforcement Learning (IRL), and Learning-to-Optimize (L2O) as potential tools for further improving computational efficiency. Building on this, the third chapter introduces the Reinforcement Learning Temporal Difference based Condition Choice Simulation (RLTD-CCS) algorithm. The algorithm leverages forward simulations (CCS) and stepwise TD learning to update value functions more efficiently than CCS. Monte Carlo studies on a small state space machine replacement model and a large state space prototype food choice model confirm that RLTD-CCS achieves estimation accuracy comparable to CCS while being up to 14 times faster, making it a viable approach for scalable DDC estimation. Additionally, in the fourth chapter, RLTD-CCS is embedded within the Expectation-Maximization (EM) algorithm to accommodate models with time-invariant persistent unobservables. In the fifth chapter, the proposed RL-based approach is applied to real-world consumer decision-making using weekly purchase data from a UK recipe box provider. The empirical analysis explores how nutritional attributes (protein and carbohydrate content) and habit persistence shape recipe box selection over time. Results indicate that nutritional factors are the primary drivers of choice, with consumers strongly preferring high-protein meals while avoiding carbohydrate-heavy options. A time-segmented static analysis further reveals a post-COVID shift: demand for high-protein meals increased, while carbohydrate aversion intensified among loyal consumers, suggesting lasting changes in dietary preferences. Counterfactual simulations assess the impact of nutritional composition adjustments and habit persistence modifications, offering actionable insights for business strategy and consumer retention. Based on these findings, recommendations are provided for optimizing menu offerings, encouraging repeat purchases, and improving personalization strategies. This thesis advances the integration of machine learning and structural econometrics by demonstrating the potential of RL-based methods in DDC estimation. By highlighting opportunities to further incorporate RL into DDC modeling, it paves the way for more scalable and computationally efficient approaches to estimating complex structural choice models."],"dc:format.checksum.md5":["d16a4f8aefb01f10c543e4dc95aba817","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.120730"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/b7a44324-428d-47bb-81bd-b3d784e1223d/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/388332"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/cd66ceb7-09e4-437e-9c58-b3ca7d2e5ad9/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:subject":["Conditional Choice Simulation","Dynamic Discrete Choice","Food Choice Modeling","Reinforcement Learning","Two-step Estimation"],"dc:title":["Dynamic Discrete Choice Estimation using Reinforcement Learning with Applications in Online Food Markets"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:31Z"}