{"id":{"repo_id":"nps","oai_identifier":"oai:calhoun.nps.edu:10945/73378"},"canonical_url":"https://search.dev.ndltd.org/etd/nps/oai:calhoun.nps.edu:10945/73378","repository":{"repo_id":"nps","name":"Naval Postgraduate School","base_url":"https://calhoun.nps.edu/server/oai/request"},"display":{"title":"USING DCGAN TO GENERATE SYNTHETIC PACKET FLOWS FOR THREAT DETECTION","abstract":"Labeled network traffic data is needed in the training of machine learning models used in threat detection. Such traffic is scarce and often imbalanced as the labeling is intensive and requires domain expertise. Deep Convolutional Generative Adversarial Networks (DCGAN) are known for their image recognition and generation capabilities by learning inherent features within the image. In this thesis, we looked at how DCGAN could be trained to learn and generate synthetic packet flow data to supplement an imbalanced dataset to improve the performance of Random Forest and Convolutional Neural Network (CNN) models. The study assessed the correctness and quality of the generated data and investigated its effects on classifier models at various scarcity levels of malicious data available for training. We found that although DCGAN had problems generating valid data, post-processing them improved their validity. The resulting generated data improved the performance of both classifier model types but fell behind improvements made using up-sampling on Random Forest models. The study later showed that synthetic data could defeat up-sampling improvements at certain scarcity levels by significantly reducing the number of false negatives at the cost of higher false positives. This provides not only a solution for practitioners to overcome data scarcity problems, but also insights to enhancing threat detection using a combination of DCGAN synthetic data and CNN-based threat detection models","abstract_html":"Labeled network traffic data is needed in the training of machine learning models used in threat detection. Such traffic is scarce and often imbalanced as the labeling is intensive and requires domain expertise. Deep Convolutional Generative Adversarial Networks (DCGAN) are known for their image recognition and generation capabilities by learning inherent features within the image. In this thesis, we looked at how DCGAN could be trained to learn and generate synthetic packet flow data to supplement an imbalanced dataset to improve the performance of Random Forest and Convolutional Neural Network (CNN) models. The study assessed the correctness and quality of the generated data and investigated its effects on classifier models at various scarcity levels of malicious data available for training. We found that although DCGAN had problems generating valid data, post-processing them improved their validity. The resulting generated data improved the performance of both classifier model types but fell behind improvements made using up-sampling on Random Forest models. The study later showed that synthetic data could defeat up-sampling improvements at certain scarcity levels by significantly reducing the number of false negatives at the cost of higher false positives. This provides not only a solution for practitioners to overcome data scarcity problems, but also insights to enhancing threat detection using a combination of DCGAN synthetic data and CNN-based threat detection models","abstract_has_math":false,"creators":["Tan, Swee Khoon"],"institution":"Monterey, CA; Naval Postgraduate School","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Computer Science (CS)","school":null,"contributors":[],"advisors":["Barton, Armon C."],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-09","date_published":"2024-09","updated_at":"2026-07-27T20:24:23Z","subjects":[],"languages":[],"rights":["�Copyright�is reserved by the copyright owner."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10945/73378","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Barton, Armon C."]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science (CS)"]},{"key":"dc:creator","label":"Author","values":["Tan, Swee Khoon"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2024-11-01T19:16:10Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2024-11-01T19:16:10Z"]},{"key":"dc:date.issued","label":"Date","values":["2024-09"]},{"key":"dc:publisher","label":"Institution","values":["Monterey, CA; Naval Postgraduate School"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["�Copyright�is reserved by the copyright owner."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10945/73378"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Labeled network traffic data is needed in the training of machine learning models used in threat detection. Such traffic is scarce and often imbalanced as the labeling is intensive and requires domain expertise. Deep Convolutional Generative Adversarial Networks (DCGAN) are known for their image recognition and generation capabilities by learning inherent features within the image. In this thesis, we looked at how DCGAN could be trained to learn and generate synthetic packet flow data to supplement an imbalanced dataset to improve the performance of Random Forest and Convolutional Neural Network (CNN) models. The study assessed the correctness and quality of the generated data and investigated its effects on classifier models at various scarcity levels of malicious data available for training. We found that although DCGAN had problems generating valid data, post-processing them improved their validity. The resulting generated data improved the performance of both classifier model types but fell behind improvements made using up-sampling on Random Forest models. The study later showed that synthetic data could defeat up-sampling improvements at certain scarcity levels by significantly reducing the number of false negatives at the cost of higher false positives. This provides not only a solution for practitioners to overcome data scarcity problems, but also insights to enhancing threat detection using a combination of DCGAN synthetic data and CNN-based threat detection models"]},{"key":"dc:title","label":"Title","values":["USING DCGAN TO GENERATE SYNTHETIC PACKET FLOWS FOR THREAT DETECTION"]}]}],"canonical_facts":{"dc:contributor.advisor":["Barton, Armon C."],"dc:contributor.department":["Computer Science (CS)"],"dc:creator":["Tan, Swee Khoon"],"dc:date.accessioned":["2024-11-01T19:16:10Z"],"dc:date.available":["2024-11-01T19:16:10Z"],"dc:date.issued":["2024-09"],"dc:description.abstract":["Labeled network traffic data is needed in the training of machine learning models used in threat detection. Such traffic is scarce and often imbalanced as the labeling is intensive and requires domain expertise. Deep Convolutional Generative Adversarial Networks (DCGAN) are known for their image recognition and generation capabilities by learning inherent features within the image. In this thesis, we looked at how DCGAN could be trained to learn and generate synthetic packet flow data to supplement an imbalanced dataset to improve the performance of Random Forest and Convolutional Neural Network (CNN) models. The study assessed the correctness and quality of the generated data and investigated its effects on classifier models at various scarcity levels of malicious data available for training. We found that although DCGAN had problems generating valid data, post-processing them improved their validity. The resulting generated data improved the performance of both classifier model types but fell behind improvements made using up-sampling on Random Forest models. The study later showed that synthetic data could defeat up-sampling improvements at certain scarcity levels by significantly reducing the number of false negatives at the cost of higher false positives. This provides not only a solution for practitioners to overcome data scarcity problems, but also insights to enhancing threat detection using a combination of DCGAN synthetic data and CNN-based threat detection models"],"dc:identifier.uri":["https://hdl.handle.net/10945/73378"],"dc:publisher":["Monterey, CA; Naval Postgraduate School"],"dc:rights":["�Copyright�is reserved by the copyright owner."],"dc:title":["USING DCGAN TO GENERATE SYNTHETIC PACKET FLOWS FOR THREAT DETECTION"],"dc:type":["Thesis"]},"updated_at":"2026-07-27T20:24:23Z"}