Back to results

University of Cambridge

Dynamic Discrete Choice Estimation using Reinforcement Learning with Applications in Online Food Markets

Abstract

dc:description.abstract

Dynamic Discrete Choice (DDC) models are widely used to analyze sequential decision-making in economics and marketing. However, their estimation remains computationally challenging, especially as state spaces expand, limiting their application to large-scale consumer datasets. This thesis develops Reinforcement Learning (RL)-based estimation methods to improve the speed and scalability of DDC estimation. The second chapter establishes a theoretical foundation for integrating RL with DDC estimation, emphasizing the shared mathematical structure of Markov Decision Processes (MDPs) in both frameworks. It reviews key tabular RL methods including Dynamic Programming (DP), Monte Carlo (MC), Temporal Difference (TD) and Q-learning, and draws conceptual parallels with DDC methods including Nested Fixed-Point Algorithm (NFXP), Conditional Choice Simulation (CCS), and Nested Pseudo-Likelihood (NPL). The chapter discusses how incorporating both tabular RL methods and those using function or policy approximation into the DDC estimation process can reduce computation time and improve scalability in high-dimensional settings. It also discusses additional RL approaches such as Prioritized Sweeping (PS), Inverse Reinforcement Learning (IRL), and Learning-to-Optimize (L2O) as potential tools for further improving computational efficiency. Building on this, the third chapter introduces the Reinforcement Learning Temporal Difference based Condition Choice Simulation (RLTD-CCS) algorithm. The algorithm leverages forward simulations (CCS) and stepwise TD learning to update value functions more efficiently than CCS. Monte Carlo studies on a small state space machine replacement model and a large state space prototype food choice model confirm that RLTD-CCS achieves estimation accuracy comparable to CCS while being up to 14 times faster, making it a viable approach for scalable DDC estimation. Additionally, in the fourth chapter, RLTD-CCS is embedded within the Expectation-Maximization (EM) algorithm to accommodate models with time-invariant persistent unobservables. In the fifth chapter, the proposed RL-based approach is applied to real-world consumer decision-making using weekly purchase data from a UK recipe box provider. The empirical analysis explores how nutritional attributes (protein and carbohydrate content) and habit persistence shape recipe box selection over time. Results indicate that nutritional factors are the primary drivers of choice, with consumers strongly preferring high-protein meals while avoiding carbohydrate-heavy options. A time-segmented static analysis further reveals a post-COVID shift: demand for high-protein meals increased, while carbohydrate aversion intensified among loyal consumers, suggesting lasting changes in dietary preferences. Counterfactual simulations assess the impact of nutritional composition adjustments and habit persistence modifications, offering actionable insights for business strategy and consumer retention. Based on these findings, recommendations are provided for optimizing menu offerings, encouraging repeat purchases, and improving personalization strategies. This thesis advances the integration of machine learning and structural econometrics by demonstrating the potential of RL-based methods in DDC estimation. By highlighting opportunities to further incorporate RL into DDC modeling, it paves the way for more scalable and computationally efficient approaches to estimating complex structural choice models.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Srivastava, Sonal
Advisor dc:contributor.advisor
  • Prabhu, Jaideep

Subjects

dc:subject × 5

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.120730
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/388332

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Srivastava, Sonal. Dynamic Discrete Choice Estimation using Reinforcement Learning with Applications in Online Food Markets. Doctoral thesis, University of Cambridge, 2025. https://doi.org/10.17863/CAM.120730