{"id":{"repo_id":"regina","oai_identifier":"oai:uregina.scholaris.ca:10294/14356"},"canonical_url":"https://search.dev.ndltd.org/etd/regina/oai:uregina.scholaris.ca:10294/14356","repository":{"repo_id":"regina","name":"University of Regina","base_url":"https://uregina.scholaris.ca/server/oai/request"},"display":{"title":"A Machine Learning Classifiers Approach for Cardiovascular Disease Diagnosis","abstract":"This Thesis study investigates the application of Machine Learning in cardiology and the role that ensemble classifiers can play to help diagnose cardiovascular disease. The computing power and technology available to humans has helped in the development of the application of computers in cardiology. With this development comes a redundancy of some data. When there is a large amount of data as input variables, a Machine Learning model performs poorly instead of helping us to make better decisions. In light of the above, this Thesis study investigated how to select the essential variables from among a set of routine clinical data for Cardiovascular Disease (CVD) diagnosis, to determine if an individual has cardiovascular disease or not. Upon review of the current literature and the consideration of existing methodologies, it was realized that most studies used single classifiers. Moreover. these studies were based on limited datasets which were not sufficient to evaluate the performance of the models; on the other hand, they considered a large number of input variables that had inherent problems of overfitting. In this Thesis the dataset from Kaggle on Cardiovascular Disease (CVD) diagnosis and Python tools on Anaconda platform were used. The data was cleaned, and 5 feature reduction techniques were investigated. Here, in addition a statistical unbiased ensemble feature reduction is proposed by imposing a unitary weight on all intersecting features. This Thesis study showed that by considering only 7 features, the Recurrent Feature Elimination and the proposed unbiased-ensemble feature reduction techniques were effective for reducing variables from routine clinical data. From each feature reduction method, the diverse selected features are then fed into a set of Machine Learning techniques to compose a corresponding classifier. This Machine Learning approach in turn considers 5 independent classifiers and one additional proposed Ensemble Classifier based on those 5 classifiers. This proposed Ensemble Classifier consisted of: Multilayer Perceptron, Random Forest, Support Vector Machine, Logistic Regression and K-Nearest Neighbor classifiers. The output of the Machine Learning Classifiers approach is a classification to determine: an individual with cardiovascular disease; or an individual that is free from cardiovascular disease. In this Thesis, it is found that based on the Recurrent Feature Elimination technique, the proposed Ensemble Classifier with a mean accuracy of 0.73408, ROC_AUC of 0.73164 and standard deviation of 0.0038, had the best performance. By considering the effective Recursive Feature Elimination method and the proposed Ensemble Classifier it was demonstrated that the body weight of an individual, systolic and diastolic blood pressure, cholesterol level, glucose level, level of physical activity, and the age are decisive in diagnosing the CVD condition of an individual. It is relevant to mention that a genetic feature was not available from the considered database; therefore, this potentially important factor was not considered in this Thesis study.","abstract_html":"This Thesis study investigates the application of Machine Learning in cardiology and the role that ensemble classifiers can play to help diagnose cardiovascular disease. The computing power and technology available to humans has helped in the development of the application of computers in cardiology. With this development comes a redundancy of some data. When there is a large amount of data as input variables, a Machine Learning model performs poorly instead of helping us to make better decisions. In light of the above, this Thesis study investigated how to select the essential variables from among a set of routine clinical data for Cardiovascular Disease (CVD) diagnosis, to determine if an individual has cardiovascular disease or not. Upon review of the current literature and the consideration of existing methodologies, it was realized that most studies used single classifiers. Moreover. these studies were based on limited datasets which were not sufficient to evaluate the performance of the models; on the other hand, they considered a large number of input variables that had inherent problems of overfitting. In this Thesis the dataset from Kaggle on Cardiovascular Disease (CVD) diagnosis and Python tools on Anaconda platform were used. The data was cleaned, and 5 feature reduction techniques were investigated. Here, in addition a statistical unbiased ensemble feature reduction is proposed by imposing a unitary weight on all intersecting features. This Thesis study showed that by considering only 7 features, the Recurrent Feature Elimination and the proposed unbiased-ensemble feature reduction techniques were effective for reducing variables from routine clinical data. From each feature reduction method, the diverse selected features are then fed into a set of Machine Learning techniques to compose a corresponding classifier. This Machine Learning approach in turn considers 5 independent classifiers and one additional proposed Ensemble Classifier based on those 5 classifiers. This proposed Ensemble Classifier consisted of: Multilayer Perceptron, Random Forest, Support Vector Machine, Logistic Regression and K-Nearest Neighbor classifiers. The output of the Machine Learning Classifiers approach is a classification to determine: an individual with cardiovascular disease; or an individual that is free from cardiovascular disease. In this Thesis, it is found that based on the Recurrent Feature Elimination technique, the proposed Ensemble Classifier with a mean accuracy of 0.73408, ROC_AUC of 0.73164 and standard deviation of 0.0038, had the best performance. By considering the effective Recursive Feature Elimination method and the proposed Ensemble Classifier it was demonstrated that the body weight of an individual, systolic and diastolic blood pressure, cholesterol level, glucose level, level of physical activity, and the age are decisive in diagnosing the CVD condition of an individual. It is relevant to mention that a genetic feature was not available from the considered database; therefore, this potentially important factor was not considered in this Thesis study.","abstract_has_math":false,"creators":["Oyelude, Oyetunde Philip"],"institution":"Faculty of Graduate Studies and Research, University of Regina","degree_name":"Master of Applied Science (MASc)","degree_level":"Master&apos;s","degree_discipline":"Engineering - Industrial Systems","degree_department":null,"school":null,"contributors":[],"advisors":["Mayorga, Rene"],"committee_chairs":[],"committee_members":["Peng, Wei","Hussein, Esam"],"year":2020,"date_issued":"2020-10","date_published":"2020-10","updated_at":"2026-07-24T04:03:50Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.82465/5070"],"render_values":[{"text":"https://doi.org/10.82465/5070","href":"https://doi.org/10.82465/5070","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10294/14356","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Mayorga, Rene"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Peng, Wei","Hussein, Esam"]},{"key":"dc:creator","label":"Author","values":["Oyelude, Oyetunde Philip"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2021-09-22T22:20:36Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2021-09-22T22:20:36Z"]},{"key":"dc:date.issued","label":"Date","values":["2020-10"]},{"key":"dc:publisher","label":"Institution","values":["Faculty of Graduate Studies and Research, University of Regina"]},{"key":"dc:type","label":"Dc Type","values":["master thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Engineering - Industrial Systems"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Master&apos;s"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Applied Science (MASc)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Faculty of Graduate Studies and Research, University of Regina"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.82465/5070"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10294/14356"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["A Thesis Submitted to the Faculty of Graduate Studies and Research In Partial Fulfillment of the Requirements for the Degree of Master of Applied Science in Industrial Systems Engineering, University of Regina. xv, 230 p."]},{"key":"dc:description.abstract","label":"Abstract","values":["This Thesis study investigates the application of Machine Learning in cardiology and the role that ensemble classifiers can play to help diagnose cardiovascular disease. The computing power and technology available to humans has helped in the development of the application of computers in cardiology. With this development comes a redundancy of some data. When there is a large amount of data as input variables, a Machine Learning model performs poorly instead of helping us to make better decisions. In light of the above, this Thesis study investigated how to select the essential variables from among a set of routine clinical data for Cardiovascular Disease (CVD) diagnosis, to determine if an individual has cardiovascular disease or not. Upon review of the current literature and the consideration of existing methodologies, it was realized that most studies used single classifiers. Moreover. these studies were based on limited datasets which were not sufficient to evaluate the performance of the models; on the other hand, they considered a large number of input variables that had inherent problems of overfitting. In this Thesis the dataset from Kaggle on Cardiovascular Disease (CVD) diagnosis and Python tools on Anaconda platform were used. The data was cleaned, and 5 feature reduction techniques were investigated. Here, in addition a statistical unbiased ensemble feature reduction is proposed by imposing a unitary weight on all intersecting features. This Thesis study showed that by considering only 7 features, the Recurrent Feature Elimination and the proposed unbiased-ensemble feature reduction techniques were effective for reducing variables from routine clinical data. From each feature reduction method, the diverse selected features are then fed into a set of Machine Learning techniques to compose a corresponding classifier. This Machine Learning approach in turn considers 5 independent classifiers and one additional proposed Ensemble Classifier based on those 5 classifiers. This proposed Ensemble Classifier consisted of: Multilayer Perceptron, Random Forest, Support Vector Machine, Logistic Regression and K-Nearest Neighbor classifiers. The output of the Machine Learning Classifiers approach is a classification to determine: an individual with cardiovascular disease; or an individual that is free from cardiovascular disease. In this Thesis, it is found that based on the Recurrent Feature Elimination technique, the proposed Ensemble Classifier with a mean accuracy of 0.73408, ROC_AUC of 0.73164 and standard deviation of 0.0038, had the best performance. By considering the effective Recursive Feature Elimination method and the proposed Ensemble Classifier it was demonstrated that the body weight of an individual, systolic and diastolic blood pressure, cholesterol level, glucose level, level of physical activity, and the age are decisive in diagnosing the CVD condition of an individual. It is relevant to mention that a genetic feature was not available from the considered database; therefore, this potentially important factor was not considered in this Thesis study."]},{"key":"dc:title","label":"Title","values":["A Machine Learning Classifiers Approach for Cardiovascular Disease Diagnosis"]}]}],"canonical_facts":{"dc:contributor.advisor":["Mayorga, Rene"],"dc:contributor.committeemember":["Peng, Wei","Hussein, Esam"],"dc:creator":["Oyelude, Oyetunde Philip"],"dc:date.accessioned":["2021-09-22T22:20:36Z"],"dc:date.available":["2021-09-22T22:20:36Z"],"dc:date.issued":["2020-10"],"dc:description":["A Thesis Submitted to the Faculty of Graduate Studies and Research In Partial Fulfillment of the Requirements for the Degree of Master of Applied Science in Industrial Systems Engineering, University of Regina. xv, 230 p."],"dc:description.abstract":["This Thesis study investigates the application of Machine Learning in cardiology and the role that ensemble classifiers can play to help diagnose cardiovascular disease. The computing power and technology available to humans has helped in the development of the application of computers in cardiology. With this development comes a redundancy of some data. When there is a large amount of data as input variables, a Machine Learning model performs poorly instead of helping us to make better decisions. In light of the above, this Thesis study investigated how to select the essential variables from among a set of routine clinical data for Cardiovascular Disease (CVD) diagnosis, to determine if an individual has cardiovascular disease or not. Upon review of the current literature and the consideration of existing methodologies, it was realized that most studies used single classifiers. Moreover. these studies were based on limited datasets which were not sufficient to evaluate the performance of the models; on the other hand, they considered a large number of input variables that had inherent problems of overfitting. In this Thesis the dataset from Kaggle on Cardiovascular Disease (CVD) diagnosis and Python tools on Anaconda platform were used. The data was cleaned, and 5 feature reduction techniques were investigated. Here, in addition a statistical unbiased ensemble feature reduction is proposed by imposing a unitary weight on all intersecting features. This Thesis study showed that by considering only 7 features, the Recurrent Feature Elimination and the proposed unbiased-ensemble feature reduction techniques were effective for reducing variables from routine clinical data. From each feature reduction method, the diverse selected features are then fed into a set of Machine Learning techniques to compose a corresponding classifier. This Machine Learning approach in turn considers 5 independent classifiers and one additional proposed Ensemble Classifier based on those 5 classifiers. This proposed Ensemble Classifier consisted of: Multilayer Perceptron, Random Forest, Support Vector Machine, Logistic Regression and K-Nearest Neighbor classifiers. The output of the Machine Learning Classifiers approach is a classification to determine: an individual with cardiovascular disease; or an individual that is free from cardiovascular disease. In this Thesis, it is found that based on the Recurrent Feature Elimination technique, the proposed Ensemble Classifier with a mean accuracy of 0.73408, ROC_AUC of 0.73164 and standard deviation of 0.0038, had the best performance. By considering the effective Recursive Feature Elimination method and the proposed Ensemble Classifier it was demonstrated that the body weight of an individual, systolic and diastolic blood pressure, cholesterol level, glucose level, level of physical activity, and the age are decisive in diagnosing the CVD condition of an individual. It is relevant to mention that a genetic feature was not available from the considered database; therefore, this potentially important factor was not considered in this Thesis study."],"dc:identifier.doi":["https://doi.org/10.82465/5070"],"dc:identifier.uri":["https://hdl.handle.net/10294/14356"],"dc:language.iso":["en"],"dc:publisher":["Faculty of Graduate Studies and Research, University of Regina"],"dc:title":["A Machine Learning Classifiers Approach for Cardiovascular Disease Diagnosis"],"dc:type":["master thesis"],"thesis:degree_discipline":["Engineering - Industrial Systems"],"thesis:degree_level":["Master&apos;s"],"thesis:degree_name":["Master of Applied Science (MASc)"],"thesis:institution_name":["Faculty of Graduate Studies and Research, University of Regina"]},"updated_at":"2026-07-24T04:03:50Z"}