{"id":{"repo_id":"duquesne","oai_identifier":"oai:dsc.duq.edu:etd-2195"},"canonical_url":"https://search.dev.ndltd.org/etd/duquesne/oai:dsc.duq.edu:etd-2195","repository":{"repo_id":"duquesne","name":"Duquesne","base_url":"https://dsc.duq.edu/do/oai/"},"display":{"title":"Multivariate Outlier Mining Using Cluster Analysis: Case Study - National Health Interview Survey","abstract":"Outlier mining is a fundamental issue in many statistical analyses, especially in multivariate cases. Outliers may exert undue influence on outcomes of the analysis. In most cases, it is a big challenge to reveal the pattern of the outliers and the \"outlyingness\". There are several approaches and methods to detect anomalous data points in data. But no single method is perfect for every data set especially when the data dimension and volume is high. In this thesis, I review distance-based clustering methods for multivariate outlier mining and demonstrate the usefulness of it in a medical setting. Specifically, I discuss Hierarchical clustering and the multivariate methods of determining appropriate cluster(s). After mining the multivariate outliers, I examine and describe the characteristics of the variables for those outliers. Finally, I demonstrate the application of these methods using the National Health Interview Survey (NHIS) 2008 database for the purposes of studying adolescent obesity.","abstract_html":"Outlier mining is a fundamental issue in many statistical analyses, especially in multivariate cases. Outliers may exert undue influence on outcomes of the analysis. In most cases, it is a big challenge to reveal the pattern of the outliers and the &quot;outlyingness&quot;. There are several approaches and methods to detect anomalous data points in data. But no single method is perfect for every data set especially when the data dimension and volume is high. In this thesis, I review distance-based clustering methods for multivariate outlier mining and demonstrate the usefulness of it in a medical setting. Specifically, I discuss Hierarchical clustering and the multivariate methods of determining appropriate cluster(s). After mining the multivariate outliers, I examine and describe the characteristics of the variables for those outliers. Finally, I demonstrate the application of these methods using the National Health Interview Survey (NHIS) 2008 database for the purposes of studying adolescent obesity.","abstract_has_math":false,"creators":["Sharker, Md Monir Hossain"],"institution":null,"degree_name":"MS","degree_level":"Immediate Access","degree_discipline":"Computational Mathematics","degree_department":null,"school":null,"contributors":["Frank D'Amico","John Kern","John Fleming"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2010,"date_issued":"2010-01-01T08:00:00Z","date_published":"2010-01-01T08:00:00Z","updated_at":"2026-07-24T02:10:29Z","subjects":["Adolescent obesity","Hierarchical clustering","Multivariate Outlier","NHIS","Outlier Mining","Similarity(Distance) Measure"],"languages":["English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://dsc.duq.edu/etd/1179","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Frank D'Amico","John Kern","John Fleming"]},{"key":"dc:creator","label":"Author","values":["Sharker, Md Monir Hossain"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2018-08-03T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computational Mathematics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Immediate Access"]},{"key":"thesis:degree_name","label":"Degree Name","values":["MS"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Adolescent obesity","Hierarchical clustering","Multivariate Outlier","NHIS","Outlier Mining","Similarity(Distance) Measure"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://dsc.duq.edu/etd/1179"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Outlier mining is a fundamental issue in many statistical analyses, especially in multivariate cases. Outliers may exert undue influence on outcomes of the analysis. In most cases, it is a big challenge to reveal the pattern of the outliers and the \"outlyingness\". There are several approaches and methods to detect anomalous data points in data. But no single method is perfect for every data set especially when the data dimension and volume is high. In this thesis, I review distance-based clustering methods for multivariate outlier mining and demonstrate the usefulness of it in a medical setting. Specifically, I discuss Hierarchical clustering and the multivariate methods of determining appropriate cluster(s). After mining the multivariate outliers, I examine and describe the characteristics of the variables for those outliers. Finally, I demonstrate the application of these methods using the National Health Interview Survey (NHIS) 2008 database for the purposes of studying adolescent obesity."]},{"key":"dc:title","label":"Title","values":["Multivariate Outlier Mining Using Cluster Analysis: Case Study - National Health Interview Survey"]}]}],"canonical_facts":{"dc:contributor":["Frank D'Amico","John Kern","John Fleming"],"dc:creator":["Sharker, Md Monir Hossain"],"dc:date.available":["2018-08-03T07:00:00Z"],"dc:description.abstract":["Outlier mining is a fundamental issue in many statistical analyses, especially in multivariate cases. Outliers may exert undue influence on outcomes of the analysis. In most cases, it is a big challenge to reveal the pattern of the outliers and the \"outlyingness\". There are several approaches and methods to detect anomalous data points in data. But no single method is perfect for every data set especially when the data dimension and volume is high. In this thesis, I review distance-based clustering methods for multivariate outlier mining and demonstrate the usefulness of it in a medical setting. Specifically, I discuss Hierarchical clustering and the multivariate methods of determining appropriate cluster(s). After mining the multivariate outliers, I examine and describe the characteristics of the variables for those outliers. Finally, I demonstrate the application of these methods using the National Health Interview Survey (NHIS) 2008 database for the purposes of studying adolescent obesity."],"dc:identifier":["https://dsc.duq.edu/etd/1179"],"dc:language":["English"],"dc:subject":["Adolescent obesity","Hierarchical clustering","Multivariate Outlier","NHIS","Outlier Mining","Similarity(Distance) Measure"],"dc:title":["Multivariate Outlier Mining Using Cluster Analysis: Case Study - National Health Interview Survey"],"thesis:degree_discipline":["Computational Mathematics"],"thesis:degree_level":["Immediate Access"],"thesis:degree_name":["MS"]},"updated_at":"2026-07-24T02:10:29Z"}