{"id":{"repo_id":"uwtsd","oai_identifier":"oai:repository.uwtsd.ac.uk:4224"},"canonical_url":"https://search.dev.ndltd.org/etd/uwtsd/oai:repository.uwtsd.ac.uk:4224","repository":{"repo_id":"uwtsd","name":"University of Wales Trinity Saint David","base_url":"https://repository.uwtsd.ac.uk/cgi/oai2"},"display":{"title":"Optimizing Deep Learning and Machine Learning Models for Real-Time Intrusion Detection in IoT Networks","abstract":"The rapid growth of Internet of Things (IoT) deployments has expanded the network attack surface and increased the operational need for intrusion detection systems (IDS) that remain accurate under multi-class class imbalance while also being computationally feasible for near-real-time use. Using the CICIoT2023 benchmark dataset, this dissertation develops a reproducible end-to-end experimental pipeline and evaluates representative families of intrusion detection models, including the traditional machine learning, deep learning, ensemble learning, and proposed hybrid approaches. Five CICIoT2023 CSV partitions were ingested with a robust Drive-to-local caching strategy, then controlled via per-class capping (cap=3,225) to manage scale. After rare-class filtering (minimum 200 samples/class), the final experimental dataset comprised 73,211 samples, 50 features, and 27 classes, split stratified into 51,539 train, 7,029 validation, and 14,643 test instances. Across 27 evaluated models, the Proposed Voting Ensemble achieved the strongest overall detection quality (Accuracy=0.9528, Macro-F1=0.9441, MCC=0.9508) with a training time of 224.0 s, marginally exceeding strong ensemble baselines such as Bagging (DT) (MacroF1=0.9417, 95.8 s) and XGBoost (Macro-F1=0.9416, 52.9 s). Notably, runtime–performance trade-offs were material: Gradient Boosting delivered competitive Macro-F1 (0.9391) but incurred substantially higher training time (2835.6 s), while LightGBM offered a favorable efficiency profile (Macro-F1=0.9360, 40.6 s). Deep learning baselines underperformed leading ensembles (best DL: Wide MLP Macro-F1=0.7666) and one attention-based model collapsed (MacroF1=0.0031), indicating sensitivity of generic neural architectures to tabular IoT traffic representations under imbalance. Knowledge distillation did not improve detection quality versus ensembles in this setting (KD student Macro-F1=0.7247), but remains relevant for future edge focused compression studies. The dissertation’s contributions are (i) a transparent, imbalance aware benchmark emphasizing Macro-F1 and MCC, (ii) validated ensemble and tuning strategies (Optuna-guided HistGB, soft voting), and (iii) a reproducible experimental workflow with saved artefacts and figures.","abstract_html":"The rapid growth of Internet of Things (IoT) deployments has expanded the network attack surface and increased the operational need for intrusion detection systems (IDS) that remain accurate under multi-class class imbalance while also being computationally feasible for near-real-time use. Using the CICIoT2023 benchmark dataset, this dissertation develops a reproducible end-to-end experimental pipeline and evaluates representative families of intrusion detection models, including the traditional machine learning, deep learning, ensemble learning, and proposed hybrid approaches. Five CICIoT2023 CSV partitions were ingested with a robust Drive-to-local caching strategy, then controlled via per-class capping (cap=3,225) to manage scale. After rare-class filtering (minimum 200 samples/class), the final experimental dataset comprised 73,211 samples, 50 features, and 27 classes, split stratified into 51,539 train, 7,029 validation, and 14,643 test instances. Across 27 evaluated models, the Proposed Voting Ensemble achieved the strongest overall detection quality (Accuracy=0.9528, Macro-F1=0.9441, MCC=0.9508) with a training time of 224.0 s, marginally exceeding strong ensemble baselines such as Bagging (DT) (MacroF1=0.9417, 95.8 s) and XGBoost (Macro-F1=0.9416, 52.9 s). Notably, runtime–performance trade-offs were material: Gradient Boosting delivered competitive Macro-F1 (0.9391) but incurred substantially higher training time (2835.6 s), while LightGBM offered a favorable efficiency profile (Macro-F1=0.9360, 40.6 s). Deep learning baselines underperformed leading ensembles (best DL: Wide MLP Macro-F1=0.7666) and one attention-based model collapsed (MacroF1=0.0031), indicating sensitivity of generic neural architectures to tabular IoT traffic representations under imbalance. Knowledge distillation did not improve detection quality versus ensembles in this setting (KD student Macro-F1=0.7247), but remains relevant for future edge focused compression studies. The dissertation’s contributions are (i) a transparent, imbalance aware benchmark emphasizing Macro-F1 and MCC, (ii) validated ensemble and tuning strategies (Optuna-guided HistGB, soft voting), and (iii) a reproducible experimental workflow with saved artefacts and figures.","abstract_has_math":false,"creators":["Islam, Rezuan"],"institution":"University of Wales Trinity Saint David","degree_name":"msc","degree_level":"masters","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-03","date_published":"2026-03","updated_at":"2026-07-24T05:53:11Z","subjects":["QA75 Cyfrifiaduron electronig"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.grantnumber","label":"Dc Identifier Grantnumber","values":["UWTSD"],"render_values":[{"text":"UWTSD","href":null,"code":true}]}]},"links":{"outbound_url":"https://doi.org/10.82227/repository.uwtsd.ac.uk.00004224","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.sponsor","label":"Sponsor","values":["University of Wales Trinity Saint David"]},{"key":"dc:creator","label":"Author","values":["Islam, Rezuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2026-03-25"]},{"key":"dc:date.issued","label":"Date","values":["2026-03"]},{"key":"dc:publisher.commercial","label":"Dc Publisher Commercial","values":["University of Wales Trinity Saint David"]},{"key":"dc:publisher.department","label":"Dc Publisher Department","values":["Traethodau Meistr","Institute of Inner City Learning"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Wales Trinity Saint David"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://repository.uwtsd.ac.uk/id/eprint/4224/"]},{"key":"dc:type","label":"Dc Type","values":["Gosodiad"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["masters"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["msc"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["QA75 Cyfrifiaduron electronig"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["10.82227/repository.uwtsd.ac.uk.00004224"]},{"key":"dc:identifier.grantnumber","label":"Dc Identifier Grantnumber","values":["UWTSD"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://repository.uwtsd.ac.uk/id/eprint/4224/1/Islam_R_MSc_Thesis.pdf"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The rapid growth of Internet of Things (IoT) deployments has expanded the network attack surface and increased the operational need for intrusion detection systems (IDS) that remain accurate under multi-class class imbalance while also being computationally feasible for near-real-time use. Using the CICIoT2023 benchmark dataset, this dissertation develops a reproducible end-to-end experimental pipeline and evaluates representative families of intrusion detection models, including the traditional machine learning, deep learning, ensemble learning, and proposed hybrid approaches. Five CICIoT2023 CSV partitions were ingested with a robust Drive-to-local caching strategy, then controlled via per-class capping (cap=3,225) to manage scale. After rare-class filtering (minimum 200 samples/class), the final experimental dataset comprised 73,211 samples, 50 features, and 27 classes, split stratified into 51,539 train, 7,029 validation, and 14,643 test instances. Across 27 evaluated models, the Proposed Voting Ensemble achieved the strongest overall detection quality (Accuracy=0.9528, Macro-F1=0.9441, MCC=0.9508) with a training time of 224.0 s, marginally exceeding strong ensemble baselines such as Bagging (DT) (MacroF1=0.9417, 95.8 s) and XGBoost (Macro-F1=0.9416, 52.9 s). Notably, runtime–performance trade-offs were material: Gradient Boosting delivered competitive Macro-F1 (0.9391) but incurred substantially higher training time (2835.6 s), while LightGBM offered a favorable efficiency profile (Macro-F1=0.9360, 40.6 s). Deep learning baselines underperformed leading ensembles (best DL: Wide MLP Macro-F1=0.7666) and one attention-based model collapsed (MacroF1=0.0031), indicating sensitivity of generic neural architectures to tabular IoT traffic representations under imbalance. Knowledge distillation did not improve detection quality versus ensembles in this setting (KD student Macro-F1=0.7247), but remains relevant for future edge focused compression studies. The dissertation’s contributions are (i) a transparent, imbalance aware benchmark emphasizing Macro-F1 and MCC, (ii) validated ensemble and tuning strategies (Optuna-guided HistGB, soft voting), and (iii) a reproducible experimental workflow with saved artefacts and figures."]},{"key":"dc:format","label":"Dc Format","values":["text"]},{"key":"dc:title","label":"Title","values":["Optimizing Deep Learning and Machine Learning Models for Real-Time Intrusion Detection in IoT Networks"]}]}],"canonical_facts":{"dc:contributor.sponsor":["University of Wales Trinity Saint David"],"dc:creator":["Islam, Rezuan"],"dc:date":["2026-03-25"],"dc:date.issued":["2026-03"],"dc:description.abstract":["The rapid growth of Internet of Things (IoT) deployments has expanded the network attack surface and increased the operational need for intrusion detection systems (IDS) that remain accurate under multi-class class imbalance while also being computationally feasible for near-real-time use. Using the CICIoT2023 benchmark dataset, this dissertation develops a reproducible end-to-end experimental pipeline and evaluates representative families of intrusion detection models, including the traditional machine learning, deep learning, ensemble learning, and proposed hybrid approaches. Five CICIoT2023 CSV partitions were ingested with a robust Drive-to-local caching strategy, then controlled via per-class capping (cap=3,225) to manage scale. After rare-class filtering (minimum 200 samples/class), the final experimental dataset comprised 73,211 samples, 50 features, and 27 classes, split stratified into 51,539 train, 7,029 validation, and 14,643 test instances. Across 27 evaluated models, the Proposed Voting Ensemble achieved the strongest overall detection quality (Accuracy=0.9528, Macro-F1=0.9441, MCC=0.9508) with a training time of 224.0 s, marginally exceeding strong ensemble baselines such as Bagging (DT) (MacroF1=0.9417, 95.8 s) and XGBoost (Macro-F1=0.9416, 52.9 s). Notably, runtime–performance trade-offs were material: Gradient Boosting delivered competitive Macro-F1 (0.9391) but incurred substantially higher training time (2835.6 s), while LightGBM offered a favorable efficiency profile (Macro-F1=0.9360, 40.6 s). Deep learning baselines underperformed leading ensembles (best DL: Wide MLP Macro-F1=0.7666) and one attention-based model collapsed (MacroF1=0.0031), indicating sensitivity of generic neural architectures to tabular IoT traffic representations under imbalance. Knowledge distillation did not improve detection quality versus ensembles in this setting (KD student Macro-F1=0.7247), but remains relevant for future edge focused compression studies. The dissertation’s contributions are (i) a transparent, imbalance aware benchmark emphasizing Macro-F1 and MCC, (ii) validated ensemble and tuning strategies (Optuna-guided HistGB, soft voting), and (iii) a reproducible experimental workflow with saved artefacts and figures."],"dc:format":["text"],"dc:identifier.doi":["10.82227/repository.uwtsd.ac.uk.00004224"],"dc:identifier.grantnumber":["UWTSD"],"dc:identifier.uri":["https://repository.uwtsd.ac.uk/id/eprint/4224/1/Islam_R_MSc_Thesis.pdf"],"dc:publisher.commercial":["University of Wales Trinity Saint David"],"dc:publisher.department":["Traethodau Meistr","Institute of Inner City Learning"],"dc:publisher.institution":["University of Wales Trinity Saint David"],"dc:relation.isreferencedby":["https://repository.uwtsd.ac.uk/id/eprint/4224/"],"dc:subject":["QA75 Cyfrifiaduron electronig"],"dc:title":["Optimizing Deep Learning and Machine Learning Models for Real-Time Intrusion Detection in IoT Networks"],"dc:type":["Gosodiad"],"dc:type.qualificationlevel":["masters"],"dc:type.qualificationname":["msc"]},"updated_at":"2026-07-24T05:53:11Z"}