{"id":{"repo_id":"uwo","oai_identifier":"oai:uwo.scholaris.ca:20.500.14721/28653"},"canonical_url":"https://search.dev.ndltd.org/etd/uwo/oai:uwo.scholaris.ca:20.500.14721/28653","repository":{"repo_id":"uwo","name":"Western University","base_url":"https://uwo.scholaris.ca/server/oai/request"},"display":{"title":"Beyond Limits: Detecting Anomalies in Sparse, High-dimensional Data","abstract":"Anomaly detection is a critical aspect of data-driven decision-making, particularly in high-stakes areas such as fraud detection and identifying manufacturing defects. However, the proprietary nature and specialized use cases of such data often result in data that is both high-dimensional and has limited samples. These challenges arise because the data typically involves complex systems with numerous variables, and acquiring sufficient labeled examples is often cost-prohibitive or time-consuming. As a result, the data becomes sparse, and its high-dimensionality complicates the training of accurate models. This thesis addresses these issues by proposing a novel approach SparseDetect designed to detect anomalies in high dimensional and low sample situations. SparseDetect combines semi supervised anomaly detection algorithms with advanced statistical techniques, overcoming the limitations of traditional methods. By strategically selecting and grouping relevant data, the approach reduces the impact of high-dimensionality while ensuring robust model performance. The results demonstrate that SparseDetect achieves recall scores exceeding 92%, outperforming conventional anomaly detection methods, especially in scenarios with limited samples. This research offers valuable insights into anomaly detection for complex datasets, filling a critical gap in the literature and laying the foundation for future advancements in the field.","abstract_html":"Anomaly detection is a critical aspect of data-driven decision-making, particularly in high-stakes areas such as fraud detection and identifying manufacturing defects. However, the proprietary nature and specialized use cases of such data often result in data that is both high-dimensional and has limited samples. These challenges arise because the data typically involves complex systems with numerous variables, and acquiring sufficient labeled examples is often cost-prohibitive or time-consuming. As a result, the data becomes sparse, and its high-dimensionality complicates the training of accurate models. This thesis addresses these issues by proposing a novel approach SparseDetect designed to detect anomalies in high dimensional and low sample situations. SparseDetect combines semi supervised anomaly detection algorithms with advanced statistical techniques, overcoming the limitations of traditional methods. By strategically selecting and grouping relevant data, the approach reduces the impact of high-dimensionality while ensuring robust model performance. The results demonstrate that SparseDetect achieves recall scores exceeding 92%, outperforming conventional anomaly detection methods, especially in scenarios with limited samples. This research offers valuable insights into anomaly detection for complex datasets, filling a critical gap in the literature and laying the foundation for future advancements in the field.","abstract_has_math":false,"creators":["Shah, Ayush"],"institution":"The University of Western Ontario","degree_name":"M Sc","degree_level":null,"degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Narayan, Apurva"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-12-10","date_published":"2024-12-10","updated_at":"2026-07-27T21:56:18Z","subjects":["Anomaly detection","high dimensionality","low sample availability","Semi-supervised","Isolation Forest"],"languages":["en_ca"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/20.500.14721/28653","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Narayan, Apurva"]},{"key":"dc:creator","label":"Author","values":["Shah, Ayush"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-07-10T16:13:39Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-07-10T16:13:39Z"]},{"key":"dc:date.issued","label":"Date","values":["2024-12-10"]},{"key":"dc:publisher","label":"Institution","values":["The University of Western Ontario"]},{"key":"dc:type","label":"Dc Type","values":["thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M Sc"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Anomaly detection","high dimensionality","low sample availability","Semi-supervised","Isolation Forest"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_ca"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/20.500.14721/28653"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The thesis cover page in the PDF document includes references to Western University’s previous institutional repository platform, known as Scholarship@Western, and links to that platform (beginning with ir.lib.uwo.ca). In citing or referring to this thesis, use the DOI or handle from this page instead. Sample citation: Author name, \"Thesis title.\" (Year). Western University Open Repository. https://doi.org/10.71858/123456."]},{"key":"dc:description.abstract","label":"Abstract","values":["Anomaly detection is a critical aspect of data-driven decision-making, particularly in high-stakes areas such as fraud detection and identifying manufacturing defects. However, the proprietary nature and specialized use cases of such data often result in data that is both high-dimensional and has limited samples. These challenges arise because the data typically involves complex systems with numerous variables, and acquiring sufficient labeled examples is often cost-prohibitive or time-consuming. As a result, the data becomes sparse, and its high-dimensionality complicates the training of accurate models. This thesis addresses these issues by proposing a novel approach SparseDetect designed to detect anomalies in high dimensional and low sample situations. SparseDetect combines semi supervised anomaly detection algorithms with advanced statistical techniques, overcoming the limitations of traditional methods. By strategically selecting and grouping relevant data, the approach reduces the impact of high-dimensionality while ensuring robust model performance. The results demonstrate that SparseDetect achieves recall scores exceeding 92%, outperforming conventional anomaly detection methods, especially in scenarios with limited samples. This research offers valuable insights into anomaly detection for complex datasets, filling a critical gap in the literature and laying the foundation for future advancements in the field."]},{"key":"dc:title","label":"Title","values":["Beyond Limits: Detecting Anomalies in Sparse, High-dimensional Data"]}]}],"canonical_facts":{"dc:contributor.advisor":["Narayan, Apurva"],"dc:creator":["Shah, Ayush"],"dc:date.accessioned":["2025-07-10T16:13:39Z"],"dc:date.available":["2025-07-10T16:13:39Z"],"dc:date.issued":["2024-12-10"],"dc:description":["The thesis cover page in the PDF document includes references to Western University’s previous institutional repository platform, known as Scholarship@Western, and links to that platform (beginning with ir.lib.uwo.ca). In citing or referring to this thesis, use the DOI or handle from this page instead. Sample citation: Author name, \"Thesis title.\" (Year). Western University Open Repository. https://doi.org/10.71858/123456."],"dc:description.abstract":["Anomaly detection is a critical aspect of data-driven decision-making, particularly in high-stakes areas such as fraud detection and identifying manufacturing defects. However, the proprietary nature and specialized use cases of such data often result in data that is both high-dimensional and has limited samples. These challenges arise because the data typically involves complex systems with numerous variables, and acquiring sufficient labeled examples is often cost-prohibitive or time-consuming. As a result, the data becomes sparse, and its high-dimensionality complicates the training of accurate models. This thesis addresses these issues by proposing a novel approach SparseDetect designed to detect anomalies in high dimensional and low sample situations. SparseDetect combines semi supervised anomaly detection algorithms with advanced statistical techniques, overcoming the limitations of traditional methods. By strategically selecting and grouping relevant data, the approach reduces the impact of high-dimensionality while ensuring robust model performance. The results demonstrate that SparseDetect achieves recall scores exceeding 92%, outperforming conventional anomaly detection methods, especially in scenarios with limited samples. This research offers valuable insights into anomaly detection for complex datasets, filling a critical gap in the literature and laying the foundation for future advancements in the field."],"dc:identifier.uri":["https://hdl.handle.net/20.500.14721/28653"],"dc:language.iso":["en_ca"],"dc:publisher":["The University of Western Ontario"],"dc:subject":["Anomaly detection","high dimensionality","low sample availability","Semi-supervised","Isolation Forest"],"dc:title":["Beyond Limits: Detecting Anomalies in Sparse, High-dimensional Data"],"dc:type":["thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_name":["M Sc"]},"updated_at":"2026-07-27T21:56:18Z"}