{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/361432"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/361432","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Advances in Reinforcement Learning for Decision Support","abstract":"On the level of decision support, most algorithmic problems encountered in machine learning are instances of pure prediction or pure automation tasks. This dissertation takes a holistic view of decision support, and begins by identifying four important problem classes that lie between the two extremes: exploration, mediation, interpretation, and generation---specifically with an eye towards the role of reinforcement learning in helping humans 'close the loop' in sequential decision-making. In particular, we focus on the problems of: exploring new environments without guidance, interpreting observed behavior from data, generating synthetic time series, and mediating between humans and machines. For each of these, we proffer novel mathematical formalisms, propose algorithmic solutions, and present empirical illustration of their utility. In the first instance, we refine our notion of curiosity-driven exploration to separate epistemic knowledge from aleatoric variation in hindsight, propose an algorithmic framework that yields a simple and scalable generalization of curiosity that is robust to all types of stochasticity, and demonstrate state-of-the-art results in a popular benchmark. In the second instance, we formalize a unifying perspective on inverse decision modeling that generalizes existing work on imitation learning and reward learning while opening up a broader class of research problems in behavior representation, and instantiate an example for learning interpretable representations of boundedly rational decision-making. In the third instance, we propose a probabilistic generative model of time-series data that optimizes a local transition policy by reinforcement from a global energy model learned by contrastive estimation, draw a rich analogy between synthetic generation and sequential imitation, and verify that it yields useful samples on real-world datasets. In the fourth instance, we formalize the sequential problem of online decision mediation with abstentive feedback, propose an effective solution that seeks to trade off immediate loss terms against future improvements in generalization error, and illustrate its efficacy relative to applicable benchmark algorithms on a variety of metrics. Like so, this dissertation contributes and advances a broader perspective on machine learning for augmenting decision-making processes.","abstract_html":"On the level of decision support, most algorithmic problems encountered in machine learning are instances of pure prediction or pure automation tasks. This dissertation takes a holistic view of decision support, and begins by identifying four important problem classes that lie between the two extremes: exploration, mediation, interpretation, and generation---specifically with an eye towards the role of reinforcement learning in helping humans &#x27;close the loop&#x27; in sequential decision-making. In particular, we focus on the problems of: exploring new environments without guidance, interpreting observed behavior from data, generating synthetic time series, and mediating between humans and machines. For each of these, we proffer novel mathematical formalisms, propose algorithmic solutions, and present empirical illustration of their utility. In the first instance, we refine our notion of curiosity-driven exploration to separate epistemic knowledge from aleatoric variation in hindsight, propose an algorithmic framework that yields a simple and scalable generalization of curiosity that is robust to all types of stochasticity, and demonstrate state-of-the-art results in a popular benchmark. In the second instance, we formalize a unifying perspective on inverse decision modeling that generalizes existing work on imitation learning and reward learning while opening up a broader class of research problems in behavior representation, and instantiate an example for learning interpretable representations of boundedly rational decision-making. In the third instance, we propose a probabilistic generative model of time-series data that optimizes a local transition policy by reinforcement from a global energy model learned by contrastive estimation, draw a rich analogy between synthetic generation and sequential imitation, and verify that it yields useful samples on real-world datasets. In the fourth instance, we formalize the sequential problem of online decision mediation with abstentive feedback, propose an effective solution that seeks to trade off immediate loss terms against future improvements in generalization error, and illustrate its efficacy relative to applicable benchmark algorithms on a variety of metrics. Like so, this dissertation contributes and advances a broader perspective on machine learning for augmenting decision-making processes.","abstract_has_math":false,"creators":["Jarrett, Daniel"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["van der Schaar, Mihaela"],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-02-01","date_published":"2023-02-01","updated_at":"2026-07-22T22:23:59Z","subjects":["Decision Support","Generative Modeling","Reinforcement Learning"],"languages":["eng"],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/01166aa4-d999-45b9-8892-0e47066b0745/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.104260","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["van der Schaar, Mihaela"]},{"key":"dc:creator","label":"Author","values":["Jarrett, Daniel"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2023-02-01"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/361432"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Decision Support","Generative Modeling","Reinforcement Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/01166aa4-d999-45b9-8892-0e47066b0745/download","https://www.rioxx.net/licenses/all-rights-reserved/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.104260"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/06d3f868-20ed-41e0-b677-432e0743e321/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["On the level of decision support, most algorithmic problems encountered in machine learning are instances of pure prediction or pure automation tasks. This dissertation takes a holistic view of decision support, and begins by identifying four important problem classes that lie between the two extremes: exploration, mediation, interpretation, and generation---specifically with an eye towards the role of reinforcement learning in helping humans 'close the loop' in sequential decision-making. In particular, we focus on the problems of: exploring new environments without guidance, interpreting observed behavior from data, generating synthetic time series, and mediating between humans and machines. For each of these, we proffer novel mathematical formalisms, propose algorithmic solutions, and present empirical illustration of their utility. In the first instance, we refine our notion of curiosity-driven exploration to separate epistemic knowledge from aleatoric variation in hindsight, propose an algorithmic framework that yields a simple and scalable generalization of curiosity that is robust to all types of stochasticity, and demonstrate state-of-the-art results in a popular benchmark. In the second instance, we formalize a unifying perspective on inverse decision modeling that generalizes existing work on imitation learning and reward learning while opening up a broader class of research problems in behavior representation, and instantiate an example for learning interpretable representations of boundedly rational decision-making. In the third instance, we propose a probabilistic generative model of time-series data that optimizes a local transition policy by reinforcement from a global energy model learned by contrastive estimation, draw a rich analogy between synthetic generation and sequential imitation, and verify that it yields useful samples on real-world datasets. In the fourth instance, we formalize the sequential problem of online decision mediation with abstentive feedback, propose an effective solution that seeks to trade off immediate loss terms against future improvements in generalization error, and illustrate its efficacy relative to applicable benchmark algorithms on a variety of metrics. Like so, this dissertation contributes and advances a broader perspective on machine learning for augmenting decision-making processes."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["f0fbfcd807609dce921d5504a8c5eaae","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Advances in Reinforcement Learning for Decision Support"]}]}],"canonical_facts":{"dc:contributor.advisor":["van der Schaar, Mihaela"],"dc:creator":["Jarrett, Daniel"],"dc:date.issued":["2023-02-01"],"dc:description.abstract":["On the level of decision support, most algorithmic problems encountered in machine learning are instances of pure prediction or pure automation tasks. This dissertation takes a holistic view of decision support, and begins by identifying four important problem classes that lie between the two extremes: exploration, mediation, interpretation, and generation---specifically with an eye towards the role of reinforcement learning in helping humans 'close the loop' in sequential decision-making. In particular, we focus on the problems of: exploring new environments without guidance, interpreting observed behavior from data, generating synthetic time series, and mediating between humans and machines. For each of these, we proffer novel mathematical formalisms, propose algorithmic solutions, and present empirical illustration of their utility. In the first instance, we refine our notion of curiosity-driven exploration to separate epistemic knowledge from aleatoric variation in hindsight, propose an algorithmic framework that yields a simple and scalable generalization of curiosity that is robust to all types of stochasticity, and demonstrate state-of-the-art results in a popular benchmark. In the second instance, we formalize a unifying perspective on inverse decision modeling that generalizes existing work on imitation learning and reward learning while opening up a broader class of research problems in behavior representation, and instantiate an example for learning interpretable representations of boundedly rational decision-making. In the third instance, we propose a probabilistic generative model of time-series data that optimizes a local transition policy by reinforcement from a global energy model learned by contrastive estimation, draw a rich analogy between synthetic generation and sequential imitation, and verify that it yields useful samples on real-world datasets. In the fourth instance, we formalize the sequential problem of online decision mediation with abstentive feedback, propose an effective solution that seeks to trade off immediate loss terms against future improvements in generalization error, and illustrate its efficacy relative to applicable benchmark algorithms on a variety of metrics. Like so, this dissertation contributes and advances a broader perspective on machine learning for augmenting decision-making processes."],"dc:format.checksum.md5":["f0fbfcd807609dce921d5504a8c5eaae","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.104260"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/06d3f868-20ed-41e0-b677-432e0743e321/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/361432"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/01166aa4-d999-45b9-8892-0e47066b0745/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"dc:subject":["Decision Support","Generative Modeling","Reinforcement Learning"],"dc:title":["Advances in Reinforcement Learning for Decision Support"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:23:59Z"}