{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/24511"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/24511","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Filtering and refinement: a two-stage approach for efficient and effective anomaly detection","abstract":"Anomaly detection is an important data mining task. Most existing methods treat anomalies as inconsistencies and spend the majority amount of time on modeling normal instances. A recently proposed, sampling-based approach may substantially boost the efficiency in anomaly detection but may lead to weaker accuracy and robustness. In this study, we propose a two-stage approach to find anomalies in complex datasets with high accuracy as well as low time complexity and space cost. Instead of analyzing normal instances, our algorithm first employs an efficient deterministic space partition algorithm to eliminate obvious normal instances and generates a small set of anomaly candidates with a single scan of the dataset. It then checks each candidate with density-based multiple criteria to determine the final results. This two-stage framework also detects anomalies of different notions. Our experiments show that this new approach finds anomalies successfully in different conditions and ensures a good balance of efficiency, accuracy, and robustness.","abstract_html":"Anomaly detection is an important data mining task. Most existing methods treat anomalies as inconsistencies and spend the majority amount of time on modeling normal instances. A recently proposed, sampling-based approach may substantially boost the efficiency in anomaly detection but may lead to weaker accuracy and robustness. In this study, we propose a two-stage approach to find anomalies in complex datasets with high accuracy as well as low time complexity and space cost. Instead of analyzing normal instances, our algorithm first employs an efficient deterministic space partition algorithm to eliminate obvious normal instances and generates a small set of anomaly candidates with a single scan of the dataset. It then checks each candidate with density-based multiple criteria to determine the final results. This two-stage framework also detects anomalies of different notions. Our experiments show that this new approach finds anomalies successfully in different conditions and ensures a good balance of efficiency, accuracy, and robustness.","abstract_has_math":false,"creators":["Yu, Xiao"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-05-25T14:25:58Z","date_published":"2011-05-25T14:25:58Z","updated_at":"2026-07-22T22:25:23Z","subjects":["anomaly detection"],"languages":["en"],"rights":["Copyright 2011 Xiao Yu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/24511","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei"]},{"key":"dc:creator","label":"Author","values":["Yu, Xiao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2011-05-25T14:25:58Z","2013-05-26T10:00:21Z","2011-05"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["anomaly detection"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2011 Xiao Yu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/24511"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Anomaly detection is an important data mining task. Most existing methods treat anomalies as inconsistencies and spend the majority amount of time on modeling normal instances. A recently proposed, sampling-based approach may substantially boost the efficiency in anomaly detection but may lead to weaker accuracy and robustness. In this study, we propose a two-stage approach to find anomalies in complex datasets with high accuracy as well as low time complexity and space cost. Instead of analyzing normal instances, our algorithm first employs an efficient deterministic space partition algorithm to eliminate obvious normal instances and generates a small set of anomaly candidates with a single scan of the dataset. It then checks each candidate with density-based multiple criteria to determine the final results. This two-stage framework also detects anomalies of different notions. Our experiments show that this new approach finds anomalies successfully in different conditions and ensures a good balance of efficiency, accuracy, and robustness.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-04-26T13:15:20Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 thesis_anomaly_xiaoyu1.pdf: 349387 bytes, checksum: df8c74272271ae8db62384d07812dc6a (MD5)","Made available in DSpace on 2011-05-25T14:25:58Z (GMT). No. of bitstreams: 2 Yu_Xiao.pdf: 349387 bytes, checksum: df8c74272271ae8db62384d07812dc6a (MD5) license.txt: 4056 bytes, checksum: 31a40d8ba642017fb42a3d0e5df6685e (MD5)","Item marked as restricted to the 'Administrator' Group (id=1) by William Ingram (wingram2@illinois.edu) on 2011-05-25T14:30:11Z Item is restricted until 2013-05-25T14:29:35Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2013-05-26T10:00:21Z Item was in collections: University of Illinois Dissertations and Theses (ID: 204) Dissertations and Theses - Materials Science and Engineering (ID: 649) Dissertations and Theses - Computer Science (ID: 587) No. of bitstreams: 3 Yu_Xiao.pdf.txt: 44711 bytes, checksum: da101eb54d820ae09fe778992bf1f802 (MD5) Yu_Xiao.pdf: 349387 bytes, checksum: df8c74272271ae8db62384d07812dc6a (MD5) license.txt: 4056 bytes, checksum: 31a40d8ba642017fb42a3d0e5df6685e (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2013-05-26T10:00:21Z"]},{"key":"dc:title","label":"Title","values":["Filtering and refinement: a two-stage approach for efficient and effective anomaly detection"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei"],"dc:creator":["Yu, Xiao"],"dc:date":["2011-05-25T14:25:58Z","2013-05-26T10:00:21Z","2011-05"],"dc:description":["Anomaly detection is an important data mining task. Most existing methods treat anomalies as inconsistencies and spend the majority amount of time on modeling normal instances. A recently proposed, sampling-based approach may substantially boost the efficiency in anomaly detection but may lead to weaker accuracy and robustness. In this study, we propose a two-stage approach to find anomalies in complex datasets with high accuracy as well as low time complexity and space cost. Instead of analyzing normal instances, our algorithm first employs an efficient deterministic space partition algorithm to eliminate obvious normal instances and generates a small set of anomaly candidates with a single scan of the dataset. It then checks each candidate with density-based multiple criteria to determine the final results. This two-stage framework also detects anomalies of different notions. Our experiments show that this new approach finds anomalies successfully in different conditions and ensures a good balance of efficiency, accuracy, and robustness.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-04-26T13:15:20Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 thesis_anomaly_xiaoyu1.pdf: 349387 bytes, checksum: df8c74272271ae8db62384d07812dc6a (MD5)","Made available in DSpace on 2011-05-25T14:25:58Z (GMT). No. of bitstreams: 2 Yu_Xiao.pdf: 349387 bytes, checksum: df8c74272271ae8db62384d07812dc6a (MD5) license.txt: 4056 bytes, checksum: 31a40d8ba642017fb42a3d0e5df6685e (MD5)","Item marked as restricted to the 'Administrator' Group (id=1) by William Ingram (wingram2@illinois.edu) on 2011-05-25T14:30:11Z Item is restricted until 2013-05-25T14:29:35Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2013-05-26T10:00:21Z Item was in collections: University of Illinois Dissertations and Theses (ID: 204) Dissertations and Theses - Materials Science and Engineering (ID: 649) Dissertations and Theses - Computer Science (ID: 587) No. of bitstreams: 3 Yu_Xiao.pdf.txt: 44711 bytes, checksum: da101eb54d820ae09fe778992bf1f802 (MD5) Yu_Xiao.pdf: 349387 bytes, checksum: df8c74272271ae8db62384d07812dc6a (MD5) license.txt: 4056 bytes, checksum: 31a40d8ba642017fb42a3d0e5df6685e (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2013-05-26T10:00:21Z"],"dc:identifier":["http://hdl.handle.net/2142/24511"],"dc:language":["en"],"dc:rights":["Copyright 2011 Xiao Yu"],"dc:subject":["anomaly detection"],"dc:title":["Filtering and refinement: a two-stage approach for efficient and effective anomaly detection"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:23Z"}