National University of Singapore
TOWARDS HUMAN-CENTRIC AI: INVERSE REINFORCEMENT LEARNING MEETS ALGORITHMIC FAIRNESS
Abstract
dc:description.abstractThis thesis focuses on two aspects of human-centric AI - value alignment and fairness. Our first work explores Inverse reinforcement learning (IRL), a potential solution to value alignment. We introduce BO-IRL, an IRL algorithm that uses a novel kernel to explore the reward function space efficiently. The second work introduces SCALES, a framework that translates various fairness principles into fair decisions by translating them to a combination of utility, non-causal and causal components which are, in turn, mapped to elements of a Constrained Markov Decision Process. Our final work unifies the concepts of value alignment and fairness by extending the IRL problem to scenarios where the expert agent is fairness abiding. We propose FAIR-BOIRL, a new BO-based IRL algorithm that searches for solutions across both the reward function space and fairness principles. FAIR-BOIRL uses IM-GPTS, a novel acquisition function that uses an implicit multi-arm bandit strategy, to perform an efficient search.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- SREEJITH BALAKRISHNAN