Back to search

ResearchSpace@Auckland

Explainable Anomaly Detection with Few Labeled Data

Abstract

dc:description.abstract

Anomaly detection is the process of identifying unusual/rare patterns that deviate from normal behaviour. While the collected data form extra-high dimensional datasets, challenges on the predictive models’ scalability regarding both data size and dimension arises. However, extensive labeled training data for anomaly detection is enormously expensive and often unavailable in data-sensitive applications due to privacy constraints. Unsupervised approaches are promising due to their ability to learn from unlabeled data. Nonetheless, these methods offer inadequate anomaly explanations compared to supervised methods. This motivates the development of semi-supervised solutions that can provide reasonable detection accuracy and interpretation of the anomalies. This thesis breaks the problem down, explores related work, and presents possible solutions. We first survey existing anomaly detection methods of different levels of supervision, with their advantages and limitations. Afterwards, we review feature selection approaches and provide discussion on their adaptability with semi-supervised anomaly detectors. Noticing most of these methods suffer when very limited labels are provided, we propose solutions for both feature selection and anomaly detection tasks. We propose a histogram-based label-spreading approach that maximizes the use of labeled data, which propagates the small amount of class information to the local neighbourhoods. We show that this mechanism enables information-gain-based calculations, which generally require the full ground truth to be known. We extend this work as a filter-based feature selection method that is computationally light and adaptable with many statistical measurements on distributions. Furthermore, we propose a tree-based anomaly detection method that utilizes this mechanism. We show empirically that this method is effective with very limited labels and provides feature-wise anomaly explanations that showing consistencies to fully supervised models such as Random Forest.

Degree

thesis:*
Name thesis:degree_name
PhD
Level thesis:degree_level
Doctoral
Discipline thesis:degree_discipline
Computer Science
Grantor dc:publisher
ResearchSpace@Auckland
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zhang, Jingrui
Advisors dc:contributor.advisor
  • Pham, Ninh
  • Dobbie, Gillian

Rights

dc:rights
Statement dc:rights
  • Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated.

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/2292/74319
OAI identifier oai:identifier
oai:researchspace.auckland.ac.nz:2292/74319

Chain of custody

source
Harvested from
University of Auckland
Base URL
researchspace.auckland.ac.nz/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Zhang, Jingrui. Explainable Anomaly Detection with Few Labeled Data. Doctoral thesis, ResearchSpace@Auckland, 2024. https://hdl.handle.net/2292/74319