{"id":{"repo_id":"ohiolink","oai_identifier":"oai:etd.ohiolink.edu:ucin1353154433"},"canonical_url":"https://search.dev.ndltd.org/etd/ohiolink/oai:etd.ohiolink.edu:ucin1353154433","repository":{"repo_id":"ohiolink","name":"OhioLINK","base_url":"https://etd.ohiolink.edu/acprod/odb_etd/ws/oai/oai"},"display":{"title":"Feature Selection for High-risk Pattern Discovery in Medical Data","abstract":"Along with the constant improvement in the utilization of electronic information systems in the healthcare industry, substantial amounts of electronic medical data have been accumulated. Mining such data could provide useful knowledge for health care provider to improve the quality of the service been delivered. One application is to identify the patients at high-risk of severe chronic or recurrent illness. Traditionally, statistical data analysis is used in well-designed epidemiology studies to identify risk factors. This approach produced satisfactory results for some illnesses but not for others. Possible causes for the difficulty with certain illnesses include: (1) the number of risk factors considered may not be sufficient, and (2) complex multi-factor interactions may exist that cannot be efficiently discovered using conventional statistical analysis. The use of electronic medical data allows the extraction of many potential risk factors. On the other hand, complex interaction effects can be uncovered using a combination of data mining and data visualization techniques. In order to uncover clinically meaningful risk patterns, a feature selection method based on relative risk measure is proposed to identify important risk factors. To make the result friendly to subject matter experts, a visualization technique has been designed to represent the high-risk patterns. Experiments on simulated data have been conducted to compare the performance of the proposed method with existing solutions. The proposed method has also been successfully applied to the real-world case study of antidepressant utilization as well as further applied to the analysis of recurrence of major depression.","abstract_html":"Along with the constant improvement in the utilization of electronic information systems in the healthcare industry, substantial amounts of electronic medical data have been accumulated. Mining such data could provide useful knowledge for health care provider to improve the quality of the service been delivered. One application is to identify the patients at high-risk of severe chronic or recurrent illness. Traditionally, statistical data analysis is used in well-designed epidemiology studies to identify risk factors. This approach produced satisfactory results for some illnesses but not for others. Possible causes for the difficulty with certain illnesses include: (1) the number of risk factors considered may not be sufficient, and (2) complex multi-factor interactions may exist that cannot be efficiently discovered using conventional statistical analysis. The use of electronic medical data allows the extraction of many potential risk factors. On the other hand, complex interaction effects can be uncovered using a combination of data mining and data visualization techniques. In order to uncover clinically meaningful risk patterns, a feature selection method based on relative risk measure is proposed to identify important risk factors. To make the result friendly to subject matter experts, a visualization technique has been designed to represent the high-risk patterns. Experiments on simulated data have been conducted to compare the performance of the proposed method with existing solutions. The proposed method has also been successfully applied to the real-world case study of antidepressant utilization as well as further applied to the analysis of recurrence of major depression.","abstract_has_math":false,"creators":["Li, Hua"],"institution":"University of Cincinnati","degree_name":"PhD","degree_level":"doctoral","degree_discipline":"Engineering and Applied Science: Industrial Engineering","degree_department":null,"school":null,"contributors":["Huang, Hongdao"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012","date_published":"2012","updated_at":"2026-07-24T03:36:23Z","subjects":["Engineering","feature selection","high risk","healthcare","relative risk","occurrence rate","visualization"],"languages":["English"],"rights":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://rave.ohiolink.edu/etdc/view?acc_num=ucin1353154433","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Huang, Hongdao"]},{"key":"dc:creator","label":"Author","values":["Li, Hua"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012"]},{"key":"dc:publisher","label":"Institution","values":["University of Cincinnati / OhioLINK"]},{"key":"dc:type","label":"Dc Type","values":["Electronic Thesis or Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Engineering and Applied Science: Industrial Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["PhD"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Cincinnati"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Engineering","feature selection","high risk","healthcare","relative risk","occurrence rate","visualization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:rights","label":"Dc Rights","values":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://rave.ohiolink.edu/etdc/view?acc_num=ucin1353154433"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Along with the constant improvement in the utilization of electronic information systems in the healthcare industry, substantial amounts of electronic medical data have been accumulated. Mining such data could provide useful knowledge for health care provider to improve the quality of the service been delivered. One application is to identify the patients at high-risk of severe chronic or recurrent illness. Traditionally, statistical data analysis is used in well-designed epidemiology studies to identify risk factors. This approach produced satisfactory results for some illnesses but not for others. Possible causes for the difficulty with certain illnesses include: (1) the number of risk factors considered may not be sufficient, and (2) complex multi-factor interactions may exist that cannot be efficiently discovered using conventional statistical analysis. The use of electronic medical data allows the extraction of many potential risk factors. On the other hand, complex interaction effects can be uncovered using a combination of data mining and data visualization techniques. In order to uncover clinically meaningful risk patterns, a feature selection method based on relative risk measure is proposed to identify important risk factors. To make the result friendly to subject matter experts, a visualization technique has been designed to represent the high-risk patterns. Experiments on simulated data have been conducted to compare the performance of the proposed method with existing solutions. The proposed method has also been successfully applied to the real-world case study of antidepressant utilization as well as further applied to the analysis of recurrence of major depression."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf","p.78","1.83 MB"]},{"key":"dc:title","label":"Title","values":["Feature Selection for High-risk Pattern Discovery in Medical Data"]}]}],"canonical_facts":{"dc:contributor":["Huang, Hongdao"],"dc:creator":["Li, Hua"],"dc:date":["2012"],"dc:description":["Along with the constant improvement in the utilization of electronic information systems in the healthcare industry, substantial amounts of electronic medical data have been accumulated. Mining such data could provide useful knowledge for health care provider to improve the quality of the service been delivered. One application is to identify the patients at high-risk of severe chronic or recurrent illness. Traditionally, statistical data analysis is used in well-designed epidemiology studies to identify risk factors. This approach produced satisfactory results for some illnesses but not for others. Possible causes for the difficulty with certain illnesses include: (1) the number of risk factors considered may not be sufficient, and (2) complex multi-factor interactions may exist that cannot be efficiently discovered using conventional statistical analysis. The use of electronic medical data allows the extraction of many potential risk factors. On the other hand, complex interaction effects can be uncovered using a combination of data mining and data visualization techniques. In order to uncover clinically meaningful risk patterns, a feature selection method based on relative risk measure is proposed to identify important risk factors. To make the result friendly to subject matter experts, a visualization technique has been designed to represent the high-risk patterns. Experiments on simulated data have been conducted to compare the performance of the proposed method with existing solutions. The proposed method has also been successfully applied to the real-world case study of antidepressant utilization as well as further applied to the analysis of recurrence of major depression."],"dc:format":["application/pdf","p.78","1.83 MB"],"dc:identifier":["http://rave.ohiolink.edu/etdc/view?acc_num=ucin1353154433"],"dc:language":["English"],"dc:publisher":["University of Cincinnati / OhioLINK"],"dc:rights":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."],"dc:subject":["Engineering","feature selection","high risk","healthcare","relative risk","occurrence rate","visualization"],"dc:title":["Feature Selection for High-risk Pattern Discovery in Medical Data"],"dc:type":["Electronic Thesis or Dissertation"],"thesis:degree_discipline":["Engineering and Applied Science: Industrial Engineering"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["PhD"],"thesis:institution_name":["University of Cincinnati"]},"updated_at":"2026-07-24T03:36:23Z"}