{"id":{"repo_id":"regina","oai_identifier":"oai:uregina.scholaris.ca:10294/16190"},"canonical_url":"https://search.dev.ndltd.org/etd/regina/oai:uregina.scholaris.ca:10294/16190","repository":{"repo_id":"regina","name":"University of Regina","base_url":"https://uregina.scholaris.ca/server/oai/request"},"display":{"title":"Self-training for cyberbully detection: Achieving high accuracy with a balanced multi-class dataset","abstract":"Cyberbullying has become an alarming issue in the digital era, causing significant harm to its victims. The development of automated methods for detecting cyberbullying in social media is of paramount importance to safeguard vulnerable individuals. In this thesis, we propose a robust approach based on Machine Learning (ML) and Deep Learning (DL) techniques for cyberbully detection in social media platforms. Our approach involves the meticulous curation of a balanced dataset specifically designed for training the ML/ DL models. To overcome the challenge of limited labeled data, we employ a semi-supervised self-training algorithm, which effectively expands the size of the labeled dataset. By leveraging real-world social media data, we train and test the model, evaluating its performance using key metrics such as precision, recall, and F1-score. In addition, we present our meticulously annotated dataset comprising 99,991 tweets, which we have made publicly available for future scientific investigations. This dataset serves as a valuable resource for further research in this field, facilitating the development and evaluation of novel techniques for cyberbully detection. Our results underscore the near-perfect performance of the proposed approach in the context of cyberbully detection, reaffirming the efficacy of ML and DL techniques for addressing this pervasive problem. These findings offer crucial insights for future research endeavors in this domain and hold practical implications for the development of automated systems capable of detecting and combating cyberbullying in social media platforms. By continuously advancing our understanding of cyberbullying detection and developing sophisticated ML and DL models, we can foster safer digital environments and protect individuals from the detrimental effects of cyberbullying.","abstract_html":"Cyberbullying has become an alarming issue in the digital era, causing significant harm to its victims. The development of automated methods for detecting cyberbullying in social media is of paramount importance to safeguard vulnerable individuals. In this thesis, we propose a robust approach based on Machine Learning (ML) and Deep Learning (DL) techniques for cyberbully detection in social media platforms. Our approach involves the meticulous curation of a balanced dataset specifically designed for training the ML/ DL models. To overcome the challenge of limited labeled data, we employ a semi-supervised self-training algorithm, which effectively expands the size of the labeled dataset. By leveraging real-world social media data, we train and test the model, evaluating its performance using key metrics such as precision, recall, and F1-score. In addition, we present our meticulously annotated dataset comprising 99,991 tweets, which we have made publicly available for future scientific investigations. This dataset serves as a valuable resource for further research in this field, facilitating the development and evaluation of novel techniques for cyberbully detection. Our results underscore the near-perfect performance of the proposed approach in the context of cyberbully detection, reaffirming the efficacy of ML and DL techniques for addressing this pervasive problem. These findings offer crucial insights for future research endeavors in this domain and hold practical implications for the development of automated systems capable of detecting and combating cyberbullying in social media platforms. By continuously advancing our understanding of cyberbullying detection and developing sophisticated ML and DL models, we can foster safer digital environments and protect individuals from the detrimental effects of cyberbullying.","abstract_has_math":false,"creators":["Ahmadinejad, Mohamad Hosein"],"institution":"Faculty of Graduate Studies and Research, University of Regina","degree_name":"Master of Science (MSc)","degree_level":"Master&apos;s","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Nashid, Shahriar","Lisa, Fan"],"committee_chairs":[],"committee_members":["Samira, Sadaoui"],"year":2023,"date_issued":"2023-08","date_published":"2023-08","updated_at":"2026-07-24T04:03:38Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.82465/4388"],"render_values":[{"text":"https://doi.org/10.82465/4388","href":"https://doi.org/10.82465/4388","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10294/16190","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Nashid, Shahriar","Lisa, Fan"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Samira, Sadaoui"]},{"key":"dc:creator","label":"Author","values":["Ahmadinejad, Mohamad Hosein"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2023-12-18T17:32:49Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2023-12-18T17:32:49Z"]},{"key":"dc:date.issued","label":"Date","values":["2023-08"]},{"key":"dc:publisher","label":"Institution","values":["Faculty of Graduate Studies and Research, University of Regina"]},{"key":"dc:type","label":"Dc Type","values":["master thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Master&apos;s"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MSc)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Faculty of Graduate Studies and Research, University of Regina"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.82465/4388"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10294/16190"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["A Thesis Submitted to the Faculty of Graduate Studies and Research In Partial Fulfillment of the Requirements for the Degree of Master of Science in Computer Science, University of Regina. xiii, 99."]},{"key":"dc:description.abstract","label":"Abstract","values":["Cyberbullying has become an alarming issue in the digital era, causing significant harm to its victims. The development of automated methods for detecting cyberbullying in social media is of paramount importance to safeguard vulnerable individuals. In this thesis, we propose a robust approach based on Machine Learning (ML) and Deep Learning (DL) techniques for cyberbully detection in social media platforms. Our approach involves the meticulous curation of a balanced dataset specifically designed for training the ML/ DL models. To overcome the challenge of limited labeled data, we employ a semi-supervised self-training algorithm, which effectively expands the size of the labeled dataset. By leveraging real-world social media data, we train and test the model, evaluating its performance using key metrics such as precision, recall, and F1-score. In addition, we present our meticulously annotated dataset comprising 99,991 tweets, which we have made publicly available for future scientific investigations. This dataset serves as a valuable resource for further research in this field, facilitating the development and evaluation of novel techniques for cyberbully detection. Our results underscore the near-perfect performance of the proposed approach in the context of cyberbully detection, reaffirming the efficacy of ML and DL techniques for addressing this pervasive problem. These findings offer crucial insights for future research endeavors in this domain and hold practical implications for the development of automated systems capable of detecting and combating cyberbullying in social media platforms. By continuously advancing our understanding of cyberbullying detection and developing sophisticated ML and DL models, we can foster safer digital environments and protect individuals from the detrimental effects of cyberbullying."]},{"key":"dc:title","label":"Title","values":["Self-training for cyberbully detection: Achieving high accuracy with a balanced multi-class dataset"]}]}],"canonical_facts":{"dc:contributor.advisor":["Nashid, Shahriar","Lisa, Fan"],"dc:contributor.committeemember":["Samira, Sadaoui"],"dc:creator":["Ahmadinejad, Mohamad Hosein"],"dc:date.accessioned":["2023-12-18T17:32:49Z"],"dc:date.available":["2023-12-18T17:32:49Z"],"dc:date.issued":["2023-08"],"dc:description":["A Thesis Submitted to the Faculty of Graduate Studies and Research In Partial Fulfillment of the Requirements for the Degree of Master of Science in Computer Science, University of Regina. xiii, 99."],"dc:description.abstract":["Cyberbullying has become an alarming issue in the digital era, causing significant harm to its victims. The development of automated methods for detecting cyberbullying in social media is of paramount importance to safeguard vulnerable individuals. In this thesis, we propose a robust approach based on Machine Learning (ML) and Deep Learning (DL) techniques for cyberbully detection in social media platforms. Our approach involves the meticulous curation of a balanced dataset specifically designed for training the ML/ DL models. To overcome the challenge of limited labeled data, we employ a semi-supervised self-training algorithm, which effectively expands the size of the labeled dataset. By leveraging real-world social media data, we train and test the model, evaluating its performance using key metrics such as precision, recall, and F1-score. In addition, we present our meticulously annotated dataset comprising 99,991 tweets, which we have made publicly available for future scientific investigations. This dataset serves as a valuable resource for further research in this field, facilitating the development and evaluation of novel techniques for cyberbully detection. Our results underscore the near-perfect performance of the proposed approach in the context of cyberbully detection, reaffirming the efficacy of ML and DL techniques for addressing this pervasive problem. These findings offer crucial insights for future research endeavors in this domain and hold practical implications for the development of automated systems capable of detecting and combating cyberbullying in social media platforms. By continuously advancing our understanding of cyberbullying detection and developing sophisticated ML and DL models, we can foster safer digital environments and protect individuals from the detrimental effects of cyberbullying."],"dc:identifier.doi":["https://doi.org/10.82465/4388"],"dc:identifier.uri":["https://hdl.handle.net/10294/16190"],"dc:language.iso":["en"],"dc:publisher":["Faculty of Graduate Studies and Research, University of Regina"],"dc:title":["Self-training for cyberbully detection: Achieving high accuracy with a balanced multi-class dataset"],"dc:type":["master thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Master&apos;s"],"thesis:degree_name":["Master of Science (MSc)"],"thesis:institution_name":["Faculty of Graduate Studies and Research, University of Regina"]},"updated_at":"2026-07-24T04:03:38Z"}