{"id":{"repo_id":"utc","oai_identifier":"oai:scholar.utc.edu:theses-1854"},"canonical_url":"https://search.dev.ndltd.org/etd/utc/oai:scholar.utc.edu:theses-1854","repository":{"repo_id":"utc","name":"University of Tennessee - Chattanooga","base_url":"https://scholar.utc.edu/do/oai/"},"display":{"title":"How negative sampling provides class balance to rare event case data using a vehicular accident prediction project as a use case scenario","abstract":"Rare event case data occur at such an infrequent rate that even having high amounts of it can leave researchers starving for more information. There has always existed a tug and pull relationship among rare event case data, where a higher count of entries often leads to a lack of explanatory variables, and vice versa. In the research spectrum of rare event case probability prediction, several methods of data sampling exist to remedy the main issue of rare event case data: a lack of data to collect and learn from. The most effective methods often involve altering the distribution of the training samples in a data set. The least utilized of these methods is negative sampling, where positive entries in a data set are used to generate negative entries. To outline the utility of negative sampling, this work discusses the application of five types of negative sampling on a vehicular accident prediction project, where non-accident records are generated through manipulating the temporal and spatial attributes of existing accident records. Moreover, different methods of data manipulation, including feature selection and different negative to positive data ratios, are used to explore what types of explanatory variables are most important when predicting vehicular accidents. Additionally, two types of predictive models, a Multilayer Perceptron and a Logistic Regression model, are created and directly compared in terms of predictive capability. Ultimately, the best model for predictive performance is heavily dependent on the specific implementation and desired results.","abstract_html":"Rare event case data occur at such an infrequent rate that even having high amounts of it can leave researchers starving for more information. There has always existed a tug and pull relationship among rare event case data, where a higher count of entries often leads to a lack of explanatory variables, and vice versa. In the research spectrum of rare event case probability prediction, several methods of data sampling exist to remedy the main issue of rare event case data: a lack of data to collect and learn from. The most effective methods often involve altering the distribution of the training samples in a data set. The least utilized of these methods is negative sampling, where positive entries in a data set are used to generate negative entries. To outline the utility of negative sampling, this work discusses the application of five types of negative sampling on a vehicular accident prediction project, where non-accident records are generated through manipulating the temporal and spatial attributes of existing accident records. Moreover, different methods of data manipulation, including feature selection and different negative to positive data ratios, are used to explore what types of explanatory variables are most important when predicting vehicular accidents. Additionally, two types of predictive models, a Multilayer Perceptron and a Logistic Regression model, are created and directly compared in terms of predictive capability. Ultimately, the best model for predictive performance is heavily dependent on the specific implementation and desired results.","abstract_has_math":false,"creators":["Roland, Jeremy"],"institution":"University of Tennessee at Chattanooga","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Sartipi, Mina","Osman, Osama A.; Wu, Dalei","College of Engineering and Computer Science"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":null,"date_issued":"","date_published":null,"updated_at":"2026-07-24T05:46:59Z","subjects":["Machine learning","Sampling (Statistics)","Traffic accidents--Mathematical models"],"languages":["English","eng"],"rights":[],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://scholar.utc.edu/theses/681","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Sartipi, Mina","Osman, Osama A.; Wu, Dalei","College of Engineering and Computer Science"]},{"key":"dc:creator","label":"Author","values":["Roland, Jeremy"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-12-01T08:00:00Z"]},{"key":"dc:publisher","label":"Institution","values":["University of Tennessee at Chattanooga","Chattanooga (Tenn.)"]},{"key":"dc:relation","label":"Dc Relation","values":["Masters Theses and Doctoral Dissertations"]},{"key":"dc:type","label":"Dc Type","values":["Masters theses","Text"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine learning","Sampling (Statistics)","Traffic accidents--Mathematical models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://scholar.utc.edu/theses/681"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Dept. of Computer Science and Engineering","M. S.; A thesis submitted to the faculty of the University of Tennessee at Chattanooga in partial fulfillment of the requirements of the degree of Master of Science."]},{"key":"dc:description.abstract","label":"Abstract","values":["Rare event case data occur at such an infrequent rate that even having high amounts of it can leave researchers starving for more information. There has always existed a tug and pull relationship among rare event case data, where a higher count of entries often leads to a lack of explanatory variables, and vice versa. In the research spectrum of rare event case probability prediction, several methods of data sampling exist to remedy the main issue of rare event case data: a lack of data to collect and learn from. The most effective methods often involve altering the distribution of the training samples in a data set. The least utilized of these methods is negative sampling, where positive entries in a data set are used to generate negative entries. To outline the utility of negative sampling, this work discusses the application of five types of negative sampling on a vehicular accident prediction project, where non-accident records are generated through manipulating the temporal and spatial attributes of existing accident records. Moreover, different methods of data manipulation, including feature selection and different negative to positive data ratios, are used to explore what types of explanatory variables are most important when predicting vehicular accidents. Additionally, two types of predictive models, a Multilayer Perceptron and a Logistic Regression model, are created and directly compared in terms of predictive capability. Ultimately, the best model for predictive performance is heavily dependent on the specific implementation and desired results."]},{"key":"dc:title","label":"Title","values":["How negative sampling provides class balance to rare event case data using a vehicular accident prediction project as a use case scenario"]}]}],"canonical_facts":{"dc:contributor":["Sartipi, Mina","Osman, Osama A.; Wu, Dalei","College of Engineering and Computer Science"],"dc:creator":["Roland, Jeremy"],"dc:date":["2020-12-01T08:00:00Z"],"dc:description":["Dept. of Computer Science and Engineering","M. S.; A thesis submitted to the faculty of the University of Tennessee at Chattanooga in partial fulfillment of the requirements of the degree of Master of Science."],"dc:description.abstract":["Rare event case data occur at such an infrequent rate that even having high amounts of it can leave researchers starving for more information. There has always existed a tug and pull relationship among rare event case data, where a higher count of entries often leads to a lack of explanatory variables, and vice versa. In the research spectrum of rare event case probability prediction, several methods of data sampling exist to remedy the main issue of rare event case data: a lack of data to collect and learn from. The most effective methods often involve altering the distribution of the training samples in a data set. The least utilized of these methods is negative sampling, where positive entries in a data set are used to generate negative entries. To outline the utility of negative sampling, this work discusses the application of five types of negative sampling on a vehicular accident prediction project, where non-accident records are generated through manipulating the temporal and spatial attributes of existing accident records. Moreover, different methods of data manipulation, including feature selection and different negative to positive data ratios, are used to explore what types of explanatory variables are most important when predicting vehicular accidents. Additionally, two types of predictive models, a Multilayer Perceptron and a Logistic Regression model, are created and directly compared in terms of predictive capability. Ultimately, the best model for predictive performance is heavily dependent on the specific implementation and desired results."],"dc:identifier":["https://scholar.utc.edu/theses/681"],"dc:language":["English","eng"],"dc:publisher":["University of Tennessee at Chattanooga","Chattanooga (Tenn.)"],"dc:relation":["Masters Theses and Doctoral Dissertations"],"dc:rights":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Machine learning","Sampling (Statistics)","Traffic accidents--Mathematical models"],"dc:title":["How negative sampling provides class balance to rare event case data using a vehicular accident prediction project as a use case scenario"],"dc:type":["Masters theses","Text"]},"updated_at":"2026-07-24T05:46:59Z"}