{"id":{"repo_id":"maryland","oai_identifier":"oai:drum.lib.umd.edu:1903/34142"},"canonical_url":"https://search.dev.ndltd.org/etd/maryland/oai:drum.lib.umd.edu:1903/34142","repository":{"repo_id":"maryland","name":"University of Maryland","base_url":"https://api.drum.lib.umd.edu/server/oai/request"},"display":{"title":"Aligning AI with Human Values: A Path Towards Trustworthy Machine Learning Systems","abstract":"Machine learning has become a powerful tool for harnessing vast amounts of data across diverse applications. However, as artificial intelligence (AI) technologies advance and become more deeply integrated into daily life, they also introduce risks such as malicious exploitation, misinformation, and unfair decision-making, which can undermine their reliability and ethical integrity. Given AI’s growing influence, ensuring that these systems are trustworthy and aligned with human values is essential for their responsible and safe deployment. To address these challenges, this dissertation investigates trustworthiness across the AI pipeline, focusing on training-time vulnerabilities, inference-time robustness and alignment, and the long-term impacts of decision-making models. At the training stage, it examines how manipulated training data can compromise vision-language models, facilitating the spread of coherent misinformation. At the inference stage, it develops methods to enhance adversarial robustness in image classifiers and align frozen language models with human values at test time through reward guidance. For the long-term impact, it formulates fairness in sequential decision-making and proposes strategies to mitigate bias accumulation over time. Together, this dissertation aims to provide a holistic framework for improving AI reliability, safety, and fairness, fostering more trustworthy and responsible AI deployment.","abstract_html":"Machine learning has become a powerful tool for harnessing vast amounts of data across diverse applications. However, as artificial intelligence (AI) technologies advance and become more deeply integrated into daily life, they also introduce risks such as malicious exploitation, misinformation, and unfair decision-making, which can undermine their reliability and ethical integrity. Given AI’s growing influence, ensuring that these systems are trustworthy and aligned with human values is essential for their responsible and safe deployment. To address these challenges, this dissertation investigates trustworthiness across the AI pipeline, focusing on training-time vulnerabilities, inference-time robustness and alignment, and the long-term impacts of decision-making models. At the training stage, it examines how manipulated training data can compromise vision-language models, facilitating the spread of coherent misinformation. At the inference stage, it develops methods to enhance adversarial robustness in image classifiers and align frozen language models with human values at test time through reward guidance. For the long-term impact, it formulates fairness in sequential decision-making and proposes strategies to mitigate bias accumulation over time. Together, this dissertation aims to provide a holistic framework for improving AI reliability, safety, and fairness, fostering more trustworthy and responsible AI deployment.","abstract_has_math":false,"creators":["Xu, Yuancheng"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Applied Mathematics and Scientific Computation","school":null,"contributors":[],"advisors":["Huang, Furong"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-24T03:02:06Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://doi.org/10.13016/ygyc-ex7e"],"render_values":[{"text":"https://doi.org/10.13016/ygyc-ex7e","href":"https://doi.org/10.13016/ygyc-ex7e","code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/1903/34142","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Huang, Furong"]},{"key":"dc:contributor.department","label":"Department","values":["Applied Mathematics and Scientific Computation"]},{"key":"dc:creator","label":"Author","values":["Xu, Yuancheng"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-08-08T11:53:18Z"]},{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://doi.org/10.13016/ygyc-ex7e"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1903/34142"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Machine learning has become a powerful tool for harnessing vast amounts of data across diverse applications. However, as artificial intelligence (AI) technologies advance and become more deeply integrated into daily life, they also introduce risks such as malicious exploitation, misinformation, and unfair decision-making, which can undermine their reliability and ethical integrity. Given AI’s growing influence, ensuring that these systems are trustworthy and aligned with human values is essential for their responsible and safe deployment. To address these challenges, this dissertation investigates trustworthiness across the AI pipeline, focusing on training-time vulnerabilities, inference-time robustness and alignment, and the long-term impacts of decision-making models. At the training stage, it examines how manipulated training data can compromise vision-language models, facilitating the spread of coherent misinformation. At the inference stage, it develops methods to enhance adversarial robustness in image classifiers and align frozen language models with human values at test time through reward guidance. For the long-term impact, it formulates fairness in sequential decision-making and proposes strategies to mitigate bias accumulation over time. Together, this dissertation aims to provide a holistic framework for improving AI reliability, safety, and fairness, fostering more trustworthy and responsible AI deployment."]},{"key":"dc:title","label":"Title","values":["Aligning AI with Human Values: A Path Towards Trustworthy Machine Learning Systems"]}]}],"canonical_facts":{"dc:contributor.advisor":["Huang, Furong"],"dc:contributor.department":["Applied Mathematics and Scientific Computation"],"dc:creator":["Xu, Yuancheng"],"dc:date.accessioned":["2025-08-08T11:53:18Z"],"dc:date.issued":["2025"],"dc:description.abstract":["Machine learning has become a powerful tool for harnessing vast amounts of data across diverse applications. However, as artificial intelligence (AI) technologies advance and become more deeply integrated into daily life, they also introduce risks such as malicious exploitation, misinformation, and unfair decision-making, which can undermine their reliability and ethical integrity. Given AI’s growing influence, ensuring that these systems are trustworthy and aligned with human values is essential for their responsible and safe deployment. To address these challenges, this dissertation investigates trustworthiness across the AI pipeline, focusing on training-time vulnerabilities, inference-time robustness and alignment, and the long-term impacts of decision-making models. At the training stage, it examines how manipulated training data can compromise vision-language models, facilitating the spread of coherent misinformation. At the inference stage, it develops methods to enhance adversarial robustness in image classifiers and align frozen language models with human values at test time through reward guidance. For the long-term impact, it formulates fairness in sequential decision-making and proposes strategies to mitigate bias accumulation over time. Together, this dissertation aims to provide a holistic framework for improving AI reliability, safety, and fairness, fostering more trustworthy and responsible AI deployment."],"dc:identifier":["https://doi.org/10.13016/ygyc-ex7e"],"dc:identifier.uri":["http://hdl.handle.net/1903/34142"],"dc:language.iso":["en"],"dc:title":["Aligning AI with Human Values: A Path Towards Trustworthy Machine Learning Systems"],"dc:type":["Dissertation"]},"updated_at":"2026-07-24T03:02:06Z"}