Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 8 of 8 for “"Rlhf"”.
-
Aligning Language Models Using Multi-Objective Deep Reinforcement Learning
… is reinforcement learning from human feedback (RLHF). RLHF aims to optimize one objective based on human preferences. However, the cost of high-quality human feedback is enormous. Having all human annotators consistent in their opinions on desirable behaviors is also challenging. LLM alignment …
-
Inverse Constitutional AI
… Empirical evaluation is conducted on the hh-rlhf dataset for training helpful and harmless AI assistants, as well as a synthetic dataset constructed by relabeling hh-rlhf samples with predefined principles. Results demonstrate promising capabilities in clustering semantically coherent topics …
-
Offline Reward Learning from Human Demonstrations and Feedback: A Linear Programming Approach
… and reinforcement learning from human feedback (RLHF). Despite the successful application of these reward learning techniques across a wide range of tasks, a significant gap between theory and practice persists. This work aims to bridge this gap by introducing a novel linear programming (LP) …
-
Adversarial Prompt Transformation for Systematic Jailbreaks of LLMs
… Reinforcement Learning from Human Feedback (RLHF) to transform unsuccessful adversarial prompts into a successful jailbreak. Thus it learns a policy based on relation to existing jailbreak prompts that informs the generator LLM of what makes an adversarial prompt successful. This was …
-
Steerable Alignment with Conditional Multiobjective Preference Optimization
… as Reinforcement Learning from Human Feedback (RLHF) have provided useful paradigms for finetuning LLMs to produce outputs that are more consistent with human preferences. These approaches, however, assume that preferences are formed by a single, underlying reward model, which is likely …
-
Prompt Injection Generation Using Small Language Models with Reinforcement Learning with Artificial Intelligence Feedback
… like Reinforcement Learning with Human Feedback (RLHF) and keyword filtering have contributed to improving the robustness of these models, but these approaches are very resource-intensive and the models can still be vulnerable to malicious attacks like prompt injections and jailbreaking. One …
-
Goal Inference from Open-Ended Dialog
… diverse user goals. While offline methods like RLHF can represent various goals but require large datasets, our approach achieves similar flexibility with online efficiency. We extract natural language goal representations from conversations with Large Language Models (LLMs). We prompt an LLM to …
-
Training a massively multimodal transformer on YouTube data: pre-training and parameter efficient fine-tuning on HPC infrastructure
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-09-01 without embargo terms