{"id":{"repo_id":"usm","oai_identifier":"oai:aquila.usm.edu:masters_theses-2047"},"canonical_url":"https://search.dev.ndltd.org/etd/usm/oai:aquila.usm.edu:masters_theses-2047","repository":{"repo_id":"usm","name":"University of Southern Mississippi","base_url":"https://aquila.usm.edu/do/oai/"},"display":{"title":"Predicting Suicide Risk Among Youths Using Machine Learning Methods","abstract":"<p>Suicide is the second leading cause of death among youths in the USA. Although machine learning approaches have provided great potential for predicting suicide risk using survey data, prediction accuracy may not meet the need for clinical diagnosis due to the intrinsic characteristics of datasets. In this study, I perform a comparative study of six classification algorithms including naïve Bayes (NB), logistic regression (LR), multilayer perceptron (MLP), AdaBoost (Ada), random forest (RF), and bagging using YRBSS dataset and investigate the effectiveness of several data handling techniques to improve the overall performance of suicide risk prediction.</p> <p>The dataset consists of 76 health risk-related questions with 13,437 responses collected from 136 high school students in the USA. Various preprocessing techniques such as missing value imputation, feature selection, and sampling techniques for handling the imbalanced ratio of the class label were applied to the dataset. The data was partitioned into a training dataset (70%) and a test dataset (30%) using a stratified partitioning method. The performance of the classifiers was evaluated using five evaluation metrics including accuracy, precision, recall, F2 score, and area under the receiver operating characteristic curve (AUROC). The result showed that RF classifier with undersampling method achieved the highest recall of 0.84, F2 measure of 0.72, and AUROC of 0.85 followed by LR and Ada classifiers.</p> <p>Therefore, I can conclude that RF, LR, AdaBoost are powerful tools for predicting suicidal tendencies in youth. Feature selection and undersampling methods are crucial preprocessing steps necessary to identify adolescents who are at high suicide risk.</p>","abstract_html":"&lt;p&gt;Suicide is the second leading cause of death among youths in the USA. Although machine learning approaches have provided great potential for predicting suicide risk using survey data, prediction accuracy may not meet the need for clinical diagnosis due to the intrinsic characteristics of datasets. In this study, I perform a comparative study of six classification algorithms including naïve Bayes (NB), logistic regression (LR), multilayer perceptron (MLP), AdaBoost (Ada), random forest (RF), and bagging using YRBSS dataset and investigate the effectiveness of several data handling techniques to improve the overall performance of suicide risk prediction.&lt;/p&gt; &lt;p&gt;The dataset consists of 76 health risk-related questions with 13,437 responses collected from 136 high school students in the USA. Various preprocessing techniques such as missing value imputation, feature selection, and sampling techniques for handling the imbalanced ratio of the class label were applied to the dataset. The data was partitioned into a training dataset (70%) and a test dataset (30%) using a stratified partitioning method. The performance of the classifiers was evaluated using five evaluation metrics including accuracy, precision, recall, F2 score, and area under the receiver operating characteristic curve (AUROC). The result showed that RF classifier with undersampling method achieved the highest recall of 0.84, F2 measure of 0.72, and AUROC of 0.85 followed by LR and Ada classifiers.&lt;/p&gt; &lt;p&gt;Therefore, I can conclude that RF, LR, AdaBoost are powerful tools for predicting suicidal tendencies in youth. Feature selection and undersampling methods are crucial preprocessing steps necessary to identify adolescents who are at high suicide risk.&lt;/p&gt;","abstract_has_math":false,"creators":["Bhattacharjee, Saswati"],"institution":null,"degree_name":"Master of Science (MS)","degree_level":"Masters Thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Dr. Chaoyang Zhang","Dr. Sarah Lee","Dr. Ahmed Sherif"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-05-09T07:00:00Z","date_published":"2023-05-09T07:00:00Z","updated_at":"2026-07-24T05:45:40Z","subjects":["Supervised learning","machine learning","feature selection","sampling","stratified partition","ensembled learning","Other Computer Engineering"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://aquila.usm.edu/masters_theses/973","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dr. Chaoyang Zhang","Dr. Sarah Lee","Dr. Ahmed Sherif"]},{"key":"dc:creator","label":"Author","values":["Bhattacharjee, Saswati"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2024-05-31T07:00:00Z"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Supervised learning","machine learning","feature selection","sampling","stratified partition","ensembled learning","Other Computer Engineering"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://aquila.usm.edu/masters_theses/973"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Suicide is the second leading cause of death among youths in the USA. Although machine learning approaches have provided great potential for predicting suicide risk using survey data, prediction accuracy may not meet the need for clinical diagnosis due to the intrinsic characteristics of datasets. In this study, I perform a comparative study of six classification algorithms including naïve Bayes (NB), logistic regression (LR), multilayer perceptron (MLP), AdaBoost (Ada), random forest (RF), and bagging using YRBSS dataset and investigate the effectiveness of several data handling techniques to improve the overall performance of suicide risk prediction.</p> <p>The dataset consists of 76 health risk-related questions with 13,437 responses collected from 136 high school students in the USA. Various preprocessing techniques such as missing value imputation, feature selection, and sampling techniques for handling the imbalanced ratio of the class label were applied to the dataset. The data was partitioned into a training dataset (70%) and a test dataset (30%) using a stratified partitioning method. The performance of the classifiers was evaluated using five evaluation metrics including accuracy, precision, recall, F2 score, and area under the receiver operating characteristic curve (AUROC). The result showed that RF classifier with undersampling method achieved the highest recall of 0.84, F2 measure of 0.72, and AUROC of 0.85 followed by LR and Ada classifiers.</p> <p>Therefore, I can conclude that RF, LR, AdaBoost are powerful tools for predicting suicidal tendencies in youth. Feature selection and undersampling methods are crucial preprocessing steps necessary to identify adolescents who are at high suicide risk.</p>"]},{"key":"dc:title","label":"Title","values":["Predicting Suicide Risk Among Youths Using Machine Learning Methods"]}]}],"canonical_facts":{"dc:contributor":["Dr. Chaoyang Zhang","Dr. Sarah Lee","Dr. Ahmed Sherif"],"dc:creator":["Bhattacharjee, Saswati"],"dc:date.available":["2024-05-31T07:00:00Z"],"dc:description.abstract":["<p>Suicide is the second leading cause of death among youths in the USA. Although machine learning approaches have provided great potential for predicting suicide risk using survey data, prediction accuracy may not meet the need for clinical diagnosis due to the intrinsic characteristics of datasets. In this study, I perform a comparative study of six classification algorithms including naïve Bayes (NB), logistic regression (LR), multilayer perceptron (MLP), AdaBoost (Ada), random forest (RF), and bagging using YRBSS dataset and investigate the effectiveness of several data handling techniques to improve the overall performance of suicide risk prediction.</p> <p>The dataset consists of 76 health risk-related questions with 13,437 responses collected from 136 high school students in the USA. Various preprocessing techniques such as missing value imputation, feature selection, and sampling techniques for handling the imbalanced ratio of the class label were applied to the dataset. The data was partitioned into a training dataset (70%) and a test dataset (30%) using a stratified partitioning method. The performance of the classifiers was evaluated using five evaluation metrics including accuracy, precision, recall, F2 score, and area under the receiver operating characteristic curve (AUROC). The result showed that RF classifier with undersampling method achieved the highest recall of 0.84, F2 measure of 0.72, and AUROC of 0.85 followed by LR and Ada classifiers.</p> <p>Therefore, I can conclude that RF, LR, AdaBoost are powerful tools for predicting suicidal tendencies in youth. Feature selection and undersampling methods are crucial preprocessing steps necessary to identify adolescents who are at high suicide risk.</p>"],"dc:identifier":["https://aquila.usm.edu/masters_theses/973"],"dc:subject":["Supervised learning","machine learning","feature selection","sampling","stratified partition","ensembled learning","Other Computer Engineering"],"dc:title":["Predicting Suicide Risk Among Youths Using Machine Learning Methods"],"thesis:degree_level":["Masters Thesis"],"thesis:degree_name":["Master of Science (MS)"]},"updated_at":"2026-07-24T05:45:40Z"}