Back to results

University of Maryland

Aligning AI with Human Values: A Path Towards Trustworthy Machine Learning Systems

Abstract

dc:description.abstract

Machine learning has become a powerful tool for harnessing vast amounts of data across diverse applications. However, as artificial intelligence (AI) technologies advance and become more deeply integrated into daily life, they also introduce risks such as malicious exploitation, misinformation, and unfair decision-making, which can undermine their reliability and ethical integrity. Given AI’s growing influence, ensuring that these systems are trustworthy and aligned with human values is essential for their responsible and safe deployment. To address these challenges, this dissertation investigates trustworthiness across the AI pipeline, focusing on training-time vulnerabilities, inference-time robustness and alignment, and the long-term impacts of decision-making models. At the training stage, it examines how manipulated training data can compromise vision-language models, facilitating the spread of coherent misinformation. At the inference stage, it develops methods to enhance adversarial robustness in image classifiers and align frozen language models with human values at test time through reward guidance. For the long-term impact, it formulates fairness in sequential decision-making and proposes strategies to mitigate bias accumulation over time. Together, this dissertation aims to provide a holistic framework for improving AI reliability, safety, and fairness, fostering more trustworthy and responsible AI deployment.

Degree

thesis:*
Department dc:contributor.department
Applied Mathematics and Scientific Computation
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Xu, Yuancheng
Advisor dc:contributor.advisor
  • Huang, Furong

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:drum.lib.umd.edu:1903/34142

Chain of custody

source
Harvested from
University of Maryland
Base URL
api.drum.lib.umd.edu/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Xu, Yuancheng. Aligning AI with Human Values: A Path Towards Trustworthy Machine Learning Systems. 2025. http://hdl.handle.net/1903/34142