Abstract
dc:description.abstractThis thesis is about defending large language models of code (Code-LLMs) against trojan attacks. Code-LLMs are widely used in software development for tasks such as vulnerability detection, clone detection, code completion, and code summarization. Popular platforms that use AI-assisted software development include GitHub’s Copilot and ChatGPT, which are based on versions of the GPT large language model. Several studies have shown the susceptibility of LLMs to backdoor attacks (or trojans), where LLMs generate malicious output (e.g., suggesting seemingly correct but vulnerable code when presented with certain trigger words in their input prompt, but otherwise behave normally). Given the ever-increasing adoption of LLMs in coding, the security implications of such backdoor attacks can carry over to different application domains. Also, triggers need not be unique. Studies have shown stealthy triggers, e.g., the name of a developer in a comment, to mislead LLMs. Further complexity is introduced by the massiveness of code models trained on huge code datasets, which are often unknown to the end users. In this thesis, we tackle this challenge from two angles: (1) we built benchmarks and frameworks to support research in the area, and (2) we built trojan detection techniques for LLMs of code. Towards (1) we built a diverse repository of clean and trojaned code models, along with a poisoning framework that allows researchers to evaluate poisoning techniques on benchmark datasets for coding tasks. In addition, we provided a unifying taxonomy of triggers for Trojan AI for Code. Towards (2), we built and evaluated two trojan detection techniques – a black-box approach, OSeqL, for detecting input triggers to trojaned code models, and a white-box approach for detecting trojan signatures from such models. Finally, we evaluated the effects of quantization on the trojan behaviour of models. Our results indicate that OSeqL is effective in detecting and identifying certain triggers, detecting trojans from solely weight-based analysis is a hard problem, and that the effects of quantization can vary for different models. As trojans become more complex, mitigation techniques would need to account for impacts of trojans on model internals – towards which we provide insights-based future directions.
Degree
thesis:*- Name thesis:degree_name
- Doctor of Philosophy
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Houston
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Hussain, Aftab
- Advisor dc:contributor.advisor
-
- Alipour, Mohammad Amin
- Committee members dc:contributor.committeemember
-
- Huang, Stephen
- Lin, Sen
- Gnawali, Omprakash
- Hellendoorn, Vincent J.
- Xu, Bowen
Rights
- Language dc:language.iso
- English
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/10657/18302
- OAI identifier oai:identifier
- oai:uh-ir.tdl.org:10657/18302