Abstract
dc:description.abstractAnomaly detection is the process of identifying unusual/rare patterns that deviate from normal behaviour. While the collected data form extra-high dimensional datasets, challenges on the predictive models’ scalability regarding both data size and dimension arises. However, extensive labeled training data for anomaly detection is enormously expensive and often unavailable in data-sensitive applications due to privacy constraints. Unsupervised approaches are promising due to their ability to learn from unlabeled data. Nonetheless, these methods offer inadequate anomaly explanations compared to supervised methods. This motivates the development of semi-supervised solutions that can provide reasonable detection accuracy and interpretation of the anomalies. This thesis breaks the problem down, explores related work, and presents possible solutions. We first survey existing anomaly detection methods of different levels of supervision, with their advantages and limitations. Afterwards, we review feature selection approaches and provide discussion on their adaptability with semi-supervised anomaly detectors. Noticing most of these methods suffer when very limited labels are provided, we propose solutions for both feature selection and anomaly detection tasks. We propose a histogram-based label-spreading approach that maximizes the use of labeled data, which propagates the small amount of class information to the local neighbourhoods. We show that this mechanism enables information-gain-based calculations, which generally require the full ground truth to be known. We extend this work as a filter-based feature selection method that is computationally light and adaptable with many statistical measurements on distributions. Furthermore, we propose a tree-based anomaly detection method that utilizes this mechanism. We show empirically that this method is effective with very limited labels and provides feature-wise anomaly explanations that showing consistencies to fully supervised models such as Random Forest.
Degree
thesis:*- Name thesis:degree_name
- PhD
- Level thesis:degree_level
- Doctoral
- Discipline thesis:degree_discipline
- Computer Science
- Grantor dc:publisher
- ResearchSpace@Auckland
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Zhang, Jingrui
- Advisors dc:contributor.advisor
-
- Pham, Ninh
- Dobbie, Gillian
Rights
dc:rights- Statement dc:rights
-
- Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated.
- Licence dc:rights.uri
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/2292/74319
- OAI identifier oai:identifier
- oai:researchspace.auckland.ac.nz:2292/74319