Back to results

University of Houston

Trojan Detection in Large Language Models of Code

Abstract

dc:description.abstract

This thesis is about defending large language models of code (Code-LLMs) against trojan attacks. Code-LLMs are widely used in software development for tasks such as vulnerability detection, clone detection, code completion, and code summarization. Popular platforms that use AI-assisted software development include GitHub’s Copilot and ChatGPT, which are based on versions of the GPT large language model. Several studies have shown the susceptibility of LLMs to backdoor attacks (or trojans), where LLMs generate malicious output (e.g., suggesting seemingly correct but vulnerable code when presented with certain trigger words in their input prompt, but otherwise behave normally). Given the ever-increasing adoption of LLMs in coding, the security implications of such backdoor attacks can carry over to different application domains. Also, triggers need not be unique. Studies have shown stealthy triggers, e.g., the name of a developer in a comment, to mislead LLMs. Further complexity is introduced by the massiveness of code models trained on huge code datasets, which are often unknown to the end users. In this thesis, we tackle this challenge from two angles: (1) we built benchmarks and frameworks to support research in the area, and (2) we built trojan detection techniques for LLMs of code. Towards (1) we built a diverse repository of clean and trojaned code models, along with a poisoning framework that allows researchers to evaluate poisoning techniques on benchmark datasets for coding tasks. In addition, we provided a unifying taxonomy of triggers for Trojan AI for Code. Towards (2), we built and evaluated two trojan detection techniques – a black-box approach, OSeqL, for detecting input triggers to trojaned code models, and a white-box approach for detecting trojan signatures from such models. Finally, we evaluated the effects of quantization on the trojan behaviour of models. Our results indicate that OSeqL is effective in detecting and identifying certain triggers, detecting trojans from solely weight-based analysis is a hard problem, and that the effects of quantization can vary for different models. As trojans become more complex, mitigation techniques would need to account for impacts of trojans on model internals – towards which we provide insights-based future directions.

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Houston
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Hussain, Aftab
Advisor dc:contributor.advisor
  • Alipour, Mohammad Amin
Committee members dc:contributor.committeemember
  • Huang, Stephen
  • Lin, Sen
  • Gnawali, Omprakash
  • Hellendoorn, Vincent J.
  • Xu, Bowen

Rights

Language dc:language.iso
English

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/10657/18302
OAI identifier oai:identifier
oai:uh-ir.tdl.org:10657/18302

Chain of custody

source
Harvested from
University of Houston
Base URL
uh-ir.tdl.org/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Hussain, Aftab. Trojan Detection in Large Language Models of Code. University of Houston, 2024. https://hdl.handle.net/10657/18302