{"id":{"repo_id":"auckland-ms","oai_identifier":"oai:researchspace.auckland.ac.nz:2292/74319"},"canonical_url":"https://search.dev.ndltd.org/etd/auckland-ms/oai:researchspace.auckland.ac.nz:2292/74319","repository":{"repo_id":"auckland-ms","name":"University of Auckland","base_url":"https://researchspace.auckland.ac.nz/server/oai/request"},"display":{"title":"Explainable Anomaly Detection with Few Labeled Data","abstract":"Anomaly detection is the process of identifying unusual/rare patterns that deviate from normal behaviour. While the collected data form extra-high dimensional datasets, challenges on the predictive models’ scalability regarding both data size and dimension arises. However, extensive labeled training data for anomaly detection is enormously expensive and often unavailable in data-sensitive applications due to privacy constraints. Unsupervised approaches are promising due to their ability to learn from unlabeled data. Nonetheless, these methods offer inadequate anomaly explanations compared to supervised methods. This motivates the development of semi-supervised solutions that can provide reasonable detection accuracy and interpretation of the anomalies. This thesis breaks the problem down, explores related work, and presents possible solutions. We first survey existing anomaly detection methods of different levels of supervision, with their advantages and limitations. Afterwards, we review feature selection approaches and provide discussion on their adaptability with semi-supervised anomaly detectors. Noticing most of these methods suffer when very limited labels are provided, we propose solutions for both feature selection and anomaly detection tasks. We propose a histogram-based label-spreading approach that maximizes the use of labeled data, which propagates the small amount of class information to the local neighbourhoods. We show that this mechanism enables information-gain-based calculations, which generally require the full ground truth to be known. We extend this work as a filter-based feature selection method that is computationally light and adaptable with many statistical measurements on distributions. Furthermore, we propose a tree-based anomaly detection method that utilizes this mechanism. We show empirically that this method is effective with very limited labels and provides feature-wise anomaly explanations that showing consistencies to fully supervised models such as Random Forest.","abstract_html":"Anomaly detection is the process of identifying unusual/rare patterns that deviate from normal behaviour. While the collected data form extra-high dimensional datasets, challenges on the predictive models’ scalability regarding both data size and dimension arises. However, extensive labeled training data for anomaly detection is enormously expensive and often unavailable in data-sensitive applications due to privacy constraints. Unsupervised approaches are promising due to their ability to learn from unlabeled data. Nonetheless, these methods offer inadequate anomaly explanations compared to supervised methods. This motivates the development of semi-supervised solutions that can provide reasonable detection accuracy and interpretation of the anomalies. This thesis breaks the problem down, explores related work, and presents possible solutions. We first survey existing anomaly detection methods of different levels of supervision, with their advantages and limitations. Afterwards, we review feature selection approaches and provide discussion on their adaptability with semi-supervised anomaly detectors. Noticing most of these methods suffer when very limited labels are provided, we propose solutions for both feature selection and anomaly detection tasks. We propose a histogram-based label-spreading approach that maximizes the use of labeled data, which propagates the small amount of class information to the local neighbourhoods. We show that this mechanism enables information-gain-based calculations, which generally require the full ground truth to be known. We extend this work as a filter-based feature selection method that is computationally light and adaptable with many statistical measurements on distributions. Furthermore, we propose a tree-based anomaly detection method that utilizes this mechanism. We show empirically that this method is effective with very limited labels and provides feature-wise anomaly explanations that showing consistencies to fully supervised models such as Random Forest.","abstract_has_math":false,"creators":["Zhang, Jingrui"],"institution":"ResearchSpace@Auckland","degree_name":"PhD","degree_level":"Doctoral","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Pham, Ninh","Dobbie, Gillian"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024","date_published":"2024","updated_at":"2026-07-24T01:03:56Z","subjects":[],"languages":[],"rights":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."],"rights_urls":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2292/74319","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Pham, Ninh","Dobbie, Gillian"]},{"key":"dc:creator","label":"Author","values":["Zhang, Jingrui"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-01-04T19:07:28Z"]},{"key":"dc:date.issued","label":"Date","values":["2024"]},{"key":"dc:publisher","label":"Institution","values":["ResearchSpace@Auckland"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["PhD"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["The University of Auckland"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/2292/74319"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Anomaly detection is the process of identifying unusual/rare patterns that deviate from normal behaviour. While the collected data form extra-high dimensional datasets, challenges on the predictive models’ scalability regarding both data size and dimension arises. However, extensive labeled training data for anomaly detection is enormously expensive and often unavailable in data-sensitive applications due to privacy constraints. Unsupervised approaches are promising due to their ability to learn from unlabeled data. Nonetheless, these methods offer inadequate anomaly explanations compared to supervised methods. This motivates the development of semi-supervised solutions that can provide reasonable detection accuracy and interpretation of the anomalies. This thesis breaks the problem down, explores related work, and presents possible solutions. We first survey existing anomaly detection methods of different levels of supervision, with their advantages and limitations. Afterwards, we review feature selection approaches and provide discussion on their adaptability with semi-supervised anomaly detectors. Noticing most of these methods suffer when very limited labels are provided, we propose solutions for both feature selection and anomaly detection tasks. We propose a histogram-based label-spreading approach that maximizes the use of labeled data, which propagates the small amount of class information to the local neighbourhoods. We show that this mechanism enables information-gain-based calculations, which generally require the full ground truth to be known. We extend this work as a filter-based feature selection method that is computationally light and adaptable with many statistical measurements on distributions. Furthermore, we propose a tree-based anomaly detection method that utilizes this mechanism. We show empirically that this method is effective with very limited labels and provides feature-wise anomaly explanations that showing consistencies to fully supervised models such as Random Forest."]},{"key":"dc:title","label":"Title","values":["Explainable Anomaly Detection with Few Labeled Data"]}]}],"canonical_facts":{"dc:contributor.advisor":["Pham, Ninh","Dobbie, Gillian"],"dc:creator":["Zhang, Jingrui"],"dc:date.accessioned":["2026-01-04T19:07:28Z"],"dc:date.issued":["2024"],"dc:description.abstract":["Anomaly detection is the process of identifying unusual/rare patterns that deviate from normal behaviour. While the collected data form extra-high dimensional datasets, challenges on the predictive models’ scalability regarding both data size and dimension arises. However, extensive labeled training data for anomaly detection is enormously expensive and often unavailable in data-sensitive applications due to privacy constraints. Unsupervised approaches are promising due to their ability to learn from unlabeled data. Nonetheless, these methods offer inadequate anomaly explanations compared to supervised methods. This motivates the development of semi-supervised solutions that can provide reasonable detection accuracy and interpretation of the anomalies. This thesis breaks the problem down, explores related work, and presents possible solutions. We first survey existing anomaly detection methods of different levels of supervision, with their advantages and limitations. Afterwards, we review feature selection approaches and provide discussion on their adaptability with semi-supervised anomaly detectors. Noticing most of these methods suffer when very limited labels are provided, we propose solutions for both feature selection and anomaly detection tasks. We propose a histogram-based label-spreading approach that maximizes the use of labeled data, which propagates the small amount of class information to the local neighbourhoods. We show that this mechanism enables information-gain-based calculations, which generally require the full ground truth to be known. We extend this work as a filter-based feature selection method that is computationally light and adaptable with many statistical measurements on distributions. Furthermore, we propose a tree-based anomaly detection method that utilizes this mechanism. We show empirically that this method is effective with very limited labels and provides feature-wise anomaly explanations that showing consistencies to fully supervised models such as Random Forest."],"dc:identifier.uri":["https://hdl.handle.net/2292/74319"],"dc:publisher":["ResearchSpace@Auckland"],"dc:rights":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."],"dc:rights.uri":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"],"dc:title":["Explainable Anomaly Detection with Few Labeled Data"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["PhD"],"thesis:institution_name":["The University of Auckland"]},"updated_at":"2026-07-24T01:03:56Z"}