{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/97645"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/97645","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"New approaches for outlier detection","abstract":"Outlier detection has relevance in many modern day contexts, including health care, engineering, data processing and analysis, credit card fraud, monitoring computer and internet intrusions and wearable personal health sensors. Outlier detection once represented a single pre-processing step, completed prior to the analysis of data proper. Today it has importance in all stages of the data analysis pipeline, from initial processing to defining data points of interest, such as when a sensor detects an anomaly. Moreover, as data sets have grown to encompass millions and billions of observations and variables, it is imperative to have outlier detection methods capable of effectively and automatically winnowing through large amounts of data with few or no inputs from a data analyst. Many existing outlier detection methods are constrained in certain ways which might limit their utility and efficacy. For instance, it is not uncommon for outlier detection methods to require some knowledge about the data under study or require the analyst to specify information about the number of outliers in the data. Another possible constraint of many outlier detection methods is the use of the raw data. Sometimes outliers can readily be detected in the raw data; but sometimes not, in which case one can achieve greater sensitivity and accuracy from features derived from data. This study uses feature extraction on multivariate time series data and demonstrates the efficacy of a set of features and their potential for aggregation through the use of Voronoi diagrams. Voronoi diagrams are constructed from the data to create tessellations which satisfy certain geometric properties. The covariance based outlier detection is proposed and demonstrated to addresses both of these challenges. It utilizes covariance information in the data and its efficacy lies in its ability to take a set of features constructed from the data and determine which feature is best at detecting outliers. The method is shown to work effectively on time series data; but it is general and can be applied or extended to other types of data objects and data sets.","abstract_html":"Outlier detection has relevance in many modern day contexts, including health care, engineering, data processing and analysis, credit card fraud, monitoring computer and internet intrusions and wearable personal health sensors. Outlier detection once represented a single pre-processing step, completed prior to the analysis of data proper. Today it has importance in all stages of the data analysis pipeline, from initial processing to defining data points of interest, such as when a sensor detects an anomaly. Moreover, as data sets have grown to encompass millions and billions of observations and variables, it is imperative to have outlier detection methods capable of effectively and automatically winnowing through large amounts of data with few or no inputs from a data analyst. Many existing outlier detection methods are constrained in certain ways which might limit their utility and efficacy. For instance, it is not uncommon for outlier detection methods to require some knowledge about the data under study or require the analyst to specify information about the number of outliers in the data. Another possible constraint of many outlier detection methods is the use of the raw data. Sometimes outliers can readily be detected in the raw data; but sometimes not, in which case one can achieve greater sensitivity and accuracy from features derived from data. This study uses feature extraction on multivariate time series data and demonstrates the efficacy of a set of features and their potential for aggregation through the use of Voronoi diagrams. Voronoi diagrams are constructed from the data to create tessellations which satisfy certain geometric properties. The covariance based outlier detection is proposed and demonstrated to addresses both of these challenges. It utilizes covariance information in the data and its efficacy lies in its ability to take a set of features constructed from the data and determine which feature is best at detecting outliers. The method is shown to work effectively on time series data; but it is general and can be applied or extended to other types of data objects and data sets.","abstract_has_math":false,"creators":["Zwilling, Christopher Eric"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Psychology","degree_department":null,"school":null,"contributors":["Wang, Michelle Y.","Anderson, Carolyn","Köhn, Hans-Friedrich","Marden, John","Drasgow, Fritz"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-08-10T20:32:18Z","date_published":"2017-08-10T20:32:18Z","updated_at":"2026-07-22T22:24:34Z","subjects":["Outliers","Covariance","Time series"],"languages":["en"],"rights":["Copyright 2017 Christopher Eric Zwilling"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/97645","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Wang, Michelle Y.","Anderson, Carolyn","Köhn, Hans-Friedrich","Marden, John","Drasgow, Fritz"]},{"key":"dc:creator","label":"Author","values":["Zwilling, Christopher Eric"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-08-10T20:32:18Z","2019-08-11T09:15:21Z","2016-09-15","2017-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Psychology"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Outliers","Covariance","Time series"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Christopher Eric Zwilling"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/97645"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Outlier detection has relevance in many modern day contexts, including health care, engineering, data processing and analysis, credit card fraud, monitoring computer and internet intrusions and wearable personal health sensors. Outlier detection once represented a single pre-processing step, completed prior to the analysis of data proper. Today it has importance in all stages of the data analysis pipeline, from initial processing to defining data points of interest, such as when a sensor detects an anomaly. Moreover, as data sets have grown to encompass millions and billions of observations and variables, it is imperative to have outlier detection methods capable of effectively and automatically winnowing through large amounts of data with few or no inputs from a data analyst. Many existing outlier detection methods are constrained in certain ways which might limit their utility and efficacy. For instance, it is not uncommon for outlier detection methods to require some knowledge about the data under study or require the analyst to specify information about the number of outliers in the data. Another possible constraint of many outlier detection methods is the use of the raw data. Sometimes outliers can readily be detected in the raw data; but sometimes not, in which case one can achieve greater sensitivity and accuracy from features derived from data. This study uses feature extraction on multivariate time series data and demonstrates the efficacy of a set of features and their potential for aggregation through the use of Voronoi diagrams. Voronoi diagrams are constructed from the data to create tessellations which satisfy certain geometric properties. The covariance based outlier detection is proposed and demonstrated to addresses both of these challenges. It utilizes covariance information in the data and its efficacy lies in its ability to take a set of features constructed from the data and determine which feature is best at detecting outliers. The method is shown to work effectively on time series data; but it is general and can be applied or extended to other types of data objects and data sets.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-05-01","The student, Christopher Zwilling, accepted the attached license on 2016-09-12 at 10:14.","The student, Christopher Zwilling, submitted this Dissertation for approval on 2016-09-12 at 10:24.","This Dissertation was approved for publication on 2016-09-15 at 08:22.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10150 on 2017-08-10 at 15:03:59","Made available in DSpace on 2017-08-10T20:32:18Z (GMT). No. of bitstreams: 3 ZWILLING-DISSERTATION-2017.pdf: 1321681 bytes, checksum: f5b507de99520aa2e182d585badd070f (MD5) LICENSE.txt: 4217 bytes, checksum: 8b0086af27ea710b6455fa92a13f9a77 (MD5) PROQUEST_LICENSE.txt: 4563 bytes, checksum: 4d2f5013ed87b72bbb7d8cae7023b2ca (MD5) Previous issue date: 2016-09-15","Embargo set by: Colleen Fallaw for item 102698 Lift date: 2019-08-10T21:27:21Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 102698 on 2019-08-11T09:15:21Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["New approaches for outlier detection"]}]}],"canonical_facts":{"dc:contributor":["Wang, Michelle Y.","Anderson, Carolyn","Köhn, Hans-Friedrich","Marden, John","Drasgow, Fritz"],"dc:creator":["Zwilling, Christopher Eric"],"dc:date":["2017-08-10T20:32:18Z","2019-08-11T09:15:21Z","2016-09-15","2017-05"],"dc:description":["Outlier detection has relevance in many modern day contexts, including health care, engineering, data processing and analysis, credit card fraud, monitoring computer and internet intrusions and wearable personal health sensors. Outlier detection once represented a single pre-processing step, completed prior to the analysis of data proper. Today it has importance in all stages of the data analysis pipeline, from initial processing to defining data points of interest, such as when a sensor detects an anomaly. Moreover, as data sets have grown to encompass millions and billions of observations and variables, it is imperative to have outlier detection methods capable of effectively and automatically winnowing through large amounts of data with few or no inputs from a data analyst. Many existing outlier detection methods are constrained in certain ways which might limit their utility and efficacy. For instance, it is not uncommon for outlier detection methods to require some knowledge about the data under study or require the analyst to specify information about the number of outliers in the data. Another possible constraint of many outlier detection methods is the use of the raw data. Sometimes outliers can readily be detected in the raw data; but sometimes not, in which case one can achieve greater sensitivity and accuracy from features derived from data. This study uses feature extraction on multivariate time series data and demonstrates the efficacy of a set of features and their potential for aggregation through the use of Voronoi diagrams. Voronoi diagrams are constructed from the data to create tessellations which satisfy certain geometric properties. The covariance based outlier detection is proposed and demonstrated to addresses both of these challenges. It utilizes covariance information in the data and its efficacy lies in its ability to take a set of features constructed from the data and determine which feature is best at detecting outliers. The method is shown to work effectively on time series data; but it is general and can be applied or extended to other types of data objects and data sets.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-05-01","The student, Christopher Zwilling, accepted the attached license on 2016-09-12 at 10:14.","The student, Christopher Zwilling, submitted this Dissertation for approval on 2016-09-12 at 10:24.","This Dissertation was approved for publication on 2016-09-15 at 08:22.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10150 on 2017-08-10 at 15:03:59","Made available in DSpace on 2017-08-10T20:32:18Z (GMT). No. of bitstreams: 3 ZWILLING-DISSERTATION-2017.pdf: 1321681 bytes, checksum: f5b507de99520aa2e182d585badd070f (MD5) LICENSE.txt: 4217 bytes, checksum: 8b0086af27ea710b6455fa92a13f9a77 (MD5) PROQUEST_LICENSE.txt: 4563 bytes, checksum: 4d2f5013ed87b72bbb7d8cae7023b2ca (MD5) Previous issue date: 2016-09-15","Embargo set by: Colleen Fallaw for item 102698 Lift date: 2019-08-10T21:27:21Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 102698 on 2019-08-11T09:15:21Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/97645"],"dc:language":["en"],"dc:rights":["Copyright 2017 Christopher Eric Zwilling"],"dc:subject":["Outliers","Covariance","Time series"],"dc:title":["New approaches for outlier detection"],"dc:type":["text"],"thesis:degree_discipline":["Psychology"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:34Z"}