Technische Universität Berlin
Software defect localization using explainable deep learning
Abstract
dc:description.abstractThe rapid proliferation of software has led to an increase in security threats, causing data breaches that have severe privacy implications and substantial financial consequences. As a result, software developers are under pressure to efficiently identify and mitigate vulnerabilities. One category of tools that has gained prominence in supporting developers in this regard is the field of machine learning-based software vulnerability detection, where models are trained to classify code as either vulnerable or clean. These models offer advantages over traditional static application testing tools, including adaptability to project-specific code and tunable decision boundaries. They have shown promising performance in vulnerability discovery, outperforming traditional static analysis techniques. Despite their potential, machine learning-based vulnerability detectors face challenges. They exhibit low transferability and generalizability, meaning that models trained on one project may not perform well on another. The scarcity of high-quality training data, along with issues related to model interpretability, poses additional hurdles. Many deep learning models are used as black boxes, making it difficult for security practitioners to understand their reasoning. Explainable AI (XAI) methods have been proposed to address the interpretability issue, allowing practitioners to gain insights into the model's decision process. However, these explanations can be noisy, and small changes in the input can lead to different results. Additionally, the choice of the right explanation method remains a challenge. The context-sensitivity of discovery models, or their ability to detect defects that span multiple modules or analyze code interprocedurally, also influences their detection capabilities. In the scope of this thesis, we explore and tackle the challenges associated with machine learning-based vulnerability discovery methods. Our focus encompasses four crucial dimensions: data quality, model interpretability, robustness, and context sensitivity. To address the scarcity of data, we employ novel augmentation techniques specifically tailored to code, which helps to increase model accuracy. Furthermore, we integrate explanation methods with dynamic program analysis to enable more effective comparisons. In terms of enhancing detection robustness, we employ causal learning techniques to effectively reduce confounding effects by up to $50\%$. Finally, we bolster defect detection by leveraging taint analysis, thereby expanding the context without encountering the issue of exponentially increasing feature spaces and, most prominently, increasing the detection rate. In this thesis, we propose solutions for learning-based vulnerability discovery to be more effectively applied in real-world scenarios. Finally, this work also provides a comprehensive overview of the challenges and advancements in the field, offering insights into the future of machine learning-based vulnerability discovery models.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Ganz, Tom
- Advisor dc:contributor.advisor
-
- Rieck, Konrad
Rights
- Licence dc:rights.uri
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Identifier URI
- https://doi.org/10.14279/depositonce-20402
- OAI identifier oai:identifier
- oai:depositonce.tu-berlin.de:11303/21601