Back to results

Brock University

Aligning Language Models Using Multi-Objective Deep Reinforcement Learning

Abstract

dc:description.abstract

Large Language Models (LLMs) have been a significant landmark of Artificial Intelligence (AI) advancement. Aligning LLMs to be helpful and harmless is a booming trend in Natural Language Processing (NLP). One of the dominant alignment techniques is reinforcement learning from human feedback (RLHF). RLHF aims to optimize one objective based on human preferences. However, the cost of high-quality human feedback is enormous. Having all human annotators consistent in their opinions on desirable behaviors is also challenging. LLM alignment is intrinsically a multi-objective optimization task since the goal is to train models to be helpful and harmless. It is found that helpfulness and harmlessness sometimes have problems in trade-offs, making it difficult for a model trained toward the optimization of one objective to perform well on both. Therefore, to address the highly potentially conflicting or dominating learning signal problem underlying LLM alignment, a multi-objective deep reinforcement learning (MODRL) methodology is proposed. The MODRL algorithm is based on an adapted deep reinforcement learning Advantage-Induced Policy Alignment (APA) algorithm and the Aligned-MTL approach for multi-task learning. From the overall perspective of helpfulness and harmlessness, language models trained via MODRL perform better than those trained using single-objective deep reinforcement learning methods that consider both objectives.

Degree

thesis:*
Name thesis:degree_name
M.Sc. Computer Science
Level thesis:degree_level
Masters
Discipline thesis:degree_discipline
Faculty of Mathematics and Science
Department dc:contributor.department
Department of Computer Science
Grantor
Brock University
Year dc:date.issued
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zhang, Yage

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • Attribution-NonCommercial-NoDerivatives 4.0 International
Language dc:language.iso
eng

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/10464/18214
OAI identifier oai:identifier
oai:brocku.scholaris.ca:10464/18214

Chain of custody

source
Harvested from
Brock University
Base URL
brocku.scholaris.ca/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Zhang, Yage. Aligning Language Models Using Multi-Objective Deep Reinforcement Learning. Masters thesis, Brock University, 2023. http://hdl.handle.net/10464/18214