{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/394801"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/394801","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Aligning Models for Human-Centric Decision Systems","abstract":"When employed for important decisions, machine learning models do not operate in a vacuum. While they have achieved impressive results across a variety of tasks, they are by no means faultless in their decisions and, without any responsibility or accountability on the part of algorithms, human oversight is needed to ensure their safe deployment. In this thesis we explore the development and alignment of machine learning models for this larger decision system, where careful consideration must be given to the interplay between machine learning models and the humans that will use them. We consider a number of core pressing problems for developing such systems: First, behavioural inference over human decision makers in order to understand their goals, diagnose their weaknesses, and inform support policies. In particular, in Chapter 2 we develop the first method for scalable Bayesian inverse reinforcement learning that can be used in large and complicated environments such as medical decision support. Second, the design of systems that moderate the information flow between predictive machine learning models and human users in order to achieve the best overall policy on a task. In Chapter 3 we consider how best to personalise explanations of a machine learning model to a human user in order to optimise overall task performance, while in Chapter 4 we instil models with a sense of local cultural nuance in order to improve their ability to detect violations in a content moderation setting, proposing a larger reasoning model use them as tools to better assist a human moderator. Third, improving the trustworthiness and alignment of language models in particular, as powerful generative models that can form the backbone of human-centric decision systems given their particularly natural way of interfacing with humans and impressive general abilities. We consider in Chapter 5 the problem of detecting when the output of a language model is not consistent with an internal notion of “truth”, which can occur inconspicuously and is clearly damaging for important decisions. Finally, in Chapter 6 we work on the problem of preference alignment in large language models, diagnosing instabilities in the common reinforcement learning from human feedback paradigm and presenting a new method for improved credit assignment that stabilises and accelerates training. In each case we conduct an investigation of the topic, provide algorithmic solutions for the challenge at hand, and validate proposals through experiments on both simulations and real-world data. Our results demonstrate consistent improvements over the contemporary state-of-the-art, and highlights the importance of taking a holistic approach to decision systems rather than simply focusing on the machine learning methods in isolation.","abstract_html":"When employed for important decisions, machine learning models do not operate in a vacuum. While they have achieved impressive results across a variety of tasks, they are by no means faultless in their decisions and, without any responsibility or accountability on the part of algorithms, human oversight is needed to ensure their safe deployment. In this thesis we explore the development and alignment of machine learning models for this larger decision system, where careful consideration must be given to the interplay between machine learning models and the humans that will use them. We consider a number of core pressing problems for developing such systems: First, behavioural inference over human decision makers in order to understand their goals, diagnose their weaknesses, and inform support policies. In particular, in Chapter 2 we develop the first method for scalable Bayesian inverse reinforcement learning that can be used in large and complicated environments such as medical decision support. Second, the design of systems that moderate the information flow between predictive machine learning models and human users in order to achieve the best overall policy on a task. In Chapter 3 we consider how best to personalise explanations of a machine learning model to a human user in order to optimise overall task performance, while in Chapter 4 we instil models with a sense of local cultural nuance in order to improve their ability to detect violations in a content moderation setting, proposing a larger reasoning model use them as tools to better assist a human moderator. Third, improving the trustworthiness and alignment of language models in particular, as powerful generative models that can form the backbone of human-centric decision systems given their particularly natural way of interfacing with humans and impressive general abilities. We consider in Chapter 5 the problem of detecting when the output of a language model is not consistent with an internal notion of “truth”, which can occur inconspicuously and is clearly damaging for important decisions. Finally, in Chapter 6 we work on the problem of preference alignment in large language models, diagnosing instabilities in the common reinforcement learning from human feedback paradigm and presenting a new method for improved credit assignment that stabilises and accelerates training. In each case we conduct an investigation of the topic, provide algorithmic solutions for the challenge at hand, and validate proposals through experiments on both simulations and real-world data. Our results demonstrate consistent improvements over the contemporary state-of-the-art, and highlights the importance of taking a holistic approach to decision systems rather than simply focusing on the machine learning methods in isolation.","abstract_has_math":false,"creators":["Chan, Alexander"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["van der Schaar, Mihaela"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-03-01","date_published":"2024-03-01","updated_at":"2026-07-22T22:23:59Z","subjects":["Machine Learning","Reinforcement Learning","Human-Centric","Imitation Learning","Alignment"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/cd638358-eeb3-459d-8d2c-90fb0d19b2e4/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.124575","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["van der Schaar, Mihaela"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["EPSRC iCASE Scholarship with Microsoft Research"]},{"key":"dc:creator","label":"Author","values":["Chan, Alexander"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-03-01"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/394801"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","Reinforcement Learning","Human-Centric","Imitation Learning","Alignment"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/cd638358-eeb3-459d-8d2c-90fb0d19b2e4/download","http://purl.org/NET/rdflicense/allrightsreserved"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.124575"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/2803fe5c-f46e-467a-8b47-885fa15d4924/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["When employed for important decisions, machine learning models do not operate in a vacuum. While they have achieved impressive results across a variety of tasks, they are by no means faultless in their decisions and, without any responsibility or accountability on the part of algorithms, human oversight is needed to ensure their safe deployment. In this thesis we explore the development and alignment of machine learning models for this larger decision system, where careful consideration must be given to the interplay between machine learning models and the humans that will use them. We consider a number of core pressing problems for developing such systems: First, behavioural inference over human decision makers in order to understand their goals, diagnose their weaknesses, and inform support policies. In particular, in Chapter 2 we develop the first method for scalable Bayesian inverse reinforcement learning that can be used in large and complicated environments such as medical decision support. Second, the design of systems that moderate the information flow between predictive machine learning models and human users in order to achieve the best overall policy on a task. In Chapter 3 we consider how best to personalise explanations of a machine learning model to a human user in order to optimise overall task performance, while in Chapter 4 we instil models with a sense of local cultural nuance in order to improve their ability to detect violations in a content moderation setting, proposing a larger reasoning model use them as tools to better assist a human moderator. Third, improving the trustworthiness and alignment of language models in particular, as powerful generative models that can form the backbone of human-centric decision systems given their particularly natural way of interfacing with humans and impressive general abilities. We consider in Chapter 5 the problem of detecting when the output of a language model is not consistent with an internal notion of “truth”, which can occur inconspicuously and is clearly damaging for important decisions. Finally, in Chapter 6 we work on the problem of preference alignment in large language models, diagnosing instabilities in the common reinforcement learning from human feedback paradigm and presenting a new method for improved credit assignment that stabilises and accelerates training. In each case we conduct an investigation of the topic, provide algorithmic solutions for the challenge at hand, and validate proposals through experiments on both simulations and real-world data. Our results demonstrate consistent improvements over the contemporary state-of-the-art, and highlights the importance of taking a holistic approach to decision systems rather than simply focusing on the machine learning methods in isolation."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["9d9c229718176c81cf507b6fd5cd6795","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Aligning Models for Human-Centric Decision Systems"]}]}],"canonical_facts":{"dc:contributor.advisor":["van der Schaar, Mihaela"],"dc:contributor.sponsor":["EPSRC iCASE Scholarship with Microsoft Research"],"dc:creator":["Chan, Alexander"],"dc:date.issued":["2024-03-01"],"dc:description.abstract":["When employed for important decisions, machine learning models do not operate in a vacuum. While they have achieved impressive results across a variety of tasks, they are by no means faultless in their decisions and, without any responsibility or accountability on the part of algorithms, human oversight is needed to ensure their safe deployment. In this thesis we explore the development and alignment of machine learning models for this larger decision system, where careful consideration must be given to the interplay between machine learning models and the humans that will use them. We consider a number of core pressing problems for developing such systems: First, behavioural inference over human decision makers in order to understand their goals, diagnose their weaknesses, and inform support policies. In particular, in Chapter 2 we develop the first method for scalable Bayesian inverse reinforcement learning that can be used in large and complicated environments such as medical decision support. Second, the design of systems that moderate the information flow between predictive machine learning models and human users in order to achieve the best overall policy on a task. In Chapter 3 we consider how best to personalise explanations of a machine learning model to a human user in order to optimise overall task performance, while in Chapter 4 we instil models with a sense of local cultural nuance in order to improve their ability to detect violations in a content moderation setting, proposing a larger reasoning model use them as tools to better assist a human moderator. Third, improving the trustworthiness and alignment of language models in particular, as powerful generative models that can form the backbone of human-centric decision systems given their particularly natural way of interfacing with humans and impressive general abilities. We consider in Chapter 5 the problem of detecting when the output of a language model is not consistent with an internal notion of “truth”, which can occur inconspicuously and is clearly damaging for important decisions. Finally, in Chapter 6 we work on the problem of preference alignment in large language models, diagnosing instabilities in the common reinforcement learning from human feedback paradigm and presenting a new method for improved credit assignment that stabilises and accelerates training. In each case we conduct an investigation of the topic, provide algorithmic solutions for the challenge at hand, and validate proposals through experiments on both simulations and real-world data. Our results demonstrate consistent improvements over the contemporary state-of-the-art, and highlights the importance of taking a holistic approach to decision systems rather than simply focusing on the machine learning methods in isolation."],"dc:format.checksum.md5":["9d9c229718176c81cf507b6fd5cd6795","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.124575"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/2803fe5c-f46e-467a-8b47-885fa15d4924/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/394801"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/cd638358-eeb3-459d-8d2c-90fb0d19b2e4/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:subject":["Machine Learning","Reinforcement Learning","Human-Centric","Imitation Learning","Alignment"],"dc:title":["Aligning Models for Human-Centric Decision Systems"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:23:59Z"}