{"id":{"repo_id":"sask","oai_identifier":"oai:harvest.usask.ca:10388/13864"},"canonical_url":"https://search.dev.ndltd.org/etd/sask/oai:harvest.usask.ca:10388/13864","repository":{"repo_id":"sask","name":"University of Saskatchewan","base_url":"https://harvest.usask.ca/server/oai/request"},"display":{"title":"Influence of training dataset selection on the performance of a machine learning model","abstract":"To observe the growth dynamics of the canola flowers during the blooming season and estimate the harvest forecast of the Canola crops, an application called ‘Flower Counter’ has been developed by the researchers of P2IRC located at the University of Saskatchewan. The model has been developed using Deep Learning (DL) based Multi-column Convolutional Neural Network (MCNN) algorithm and TensorFlow framework. This is an object counting model, that counts the Canola flowers from the images based on the learning from a given set of training images, called ‘ground-truths’. This work proposes to compose a good training dataset that would give good accuracy with a robust object detection model by using different training and testing combinations. Various evaluation techniques have been used in this work to check the impact of the training dataset, on the testing results of the model and generalizability. The primary goal of this research work is to define a good training dataset composition having diversity. A good composition also consists of different characteristics present in the dataset, that can impact the testing results and can help in creating a robust object counting model. Different characteristics of the training datasets and testing datasets are used to evaluate the most prominent characteristics and features that impact the test results. The objective is also to evaluate the impact of training dataset selection on testing results produced by the ML model in terms of accuracy. This work would help the researchers and plant scientists gain knowledge about the diversity of characteristics for the composition of a training dataset. This can give insights to reduce the manual effort which is required to create ground truth for training models by identifying the characteristics that impact testing results. Since the entire training of the model depends on the datasets collected during diverse weather conditions, there could be factors that could impact some of the experimental results. The research area for training dataset selection has not been explored much, and this research work will give good insights about model generalization capability and scopes for manual work utilization for getting a robust object counting model.","abstract_html":"To observe the growth dynamics of the canola flowers during the blooming season and estimate the harvest forecast of the Canola crops, an application called ‘Flower Counter’ has been developed by the researchers of P2IRC located at the University of Saskatchewan. The model has been developed using Deep Learning (DL) based Multi-column Convolutional Neural Network (MCNN) algorithm and TensorFlow framework. This is an object counting model, that counts the Canola flowers from the images based on the learning from a given set of training images, called ‘ground-truths’. This work proposes to compose a good training dataset that would give good accuracy with a robust object detection model by using different training and testing combinations. Various evaluation techniques have been used in this work to check the impact of the training dataset, on the testing results of the model and generalizability. The primary goal of this research work is to define a good training dataset composition having diversity. A good composition also consists of different characteristics present in the dataset, that can impact the testing results and can help in creating a robust object counting model. Different characteristics of the training datasets and testing datasets are used to evaluate the most prominent characteristics and features that impact the test results. The objective is also to evaluate the impact of training dataset selection on testing results produced by the ML model in terms of accuracy. This work would help the researchers and plant scientists gain knowledge about the diversity of characteristics for the composition of a training dataset. This can give insights to reduce the manual effort which is required to create ground truth for training models by identifying the characteristics that impact testing results. Since the entire training of the model depends on the datasets collected during diverse weather conditions, there could be factors that could impact some of the experimental results. The research area for training dataset selection has not been explored much, and this research work will give good insights about model generalization capability and scopes for manual work utilization for getting a robust object counting model.","abstract_has_math":false,"creators":["Mouli, Srishti"],"institution":"University of Saskatchewan","degree_name":"Master of Science (M.Sc.)","degree_level":"Masters","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Makaroff, Dwight","Eager, Derek"],"committee_chairs":[],"committee_members":["Stavness, Ian","Keil, Mark","Nguyen, Ha"],"year":2022,"date_issued":"2022-04-05","date_published":"2022-04-05","updated_at":"2026-07-24T04:27:06Z","subjects":["Machine Learning, Training Dataset selection"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10388/13864","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Makaroff, Dwight","Eager, Derek"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Stavness, Ian","Keil, Mark","Nguyen, Ha"]},{"key":"dc:creator","label":"Author","values":["Mouli, Srishti"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2022-04-05T15:01:23Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2022-04-05T15:01:23Z"]},{"key":"dc:date.issued","label":"Date","values":["2022-04-05"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (M.Sc.)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Saskatchewan"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning, Training Dataset selection"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10388/13864"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["To observe the growth dynamics of the canola flowers during the blooming season and estimate the harvest forecast of the Canola crops, an application called ‘Flower Counter’ has been developed by the researchers of P2IRC located at the University of Saskatchewan. The model has been developed using Deep Learning (DL) based Multi-column Convolutional Neural Network (MCNN) algorithm and TensorFlow framework. This is an object counting model, that counts the Canola flowers from the images based on the learning from a given set of training images, called ‘ground-truths’. This work proposes to compose a good training dataset that would give good accuracy with a robust object detection model by using different training and testing combinations. Various evaluation techniques have been used in this work to check the impact of the training dataset, on the testing results of the model and generalizability. The primary goal of this research work is to define a good training dataset composition having diversity. A good composition also consists of different characteristics present in the dataset, that can impact the testing results and can help in creating a robust object counting model. Different characteristics of the training datasets and testing datasets are used to evaluate the most prominent characteristics and features that impact the test results. The objective is also to evaluate the impact of training dataset selection on testing results produced by the ML model in terms of accuracy. This work would help the researchers and plant scientists gain knowledge about the diversity of characteristics for the composition of a training dataset. This can give insights to reduce the manual effort which is required to create ground truth for training models by identifying the characteristics that impact testing results. Since the entire training of the model depends on the datasets collected during diverse weather conditions, there could be factors that could impact some of the experimental results. The research area for training dataset selection has not been explored much, and this research work will give good insights about model generalization capability and scopes for manual work utilization for getting a robust object counting model."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Influence of training dataset selection on the performance of a machine learning model"]}]}],"canonical_facts":{"dc:contributor.advisor":["Makaroff, Dwight","Eager, Derek"],"dc:contributor.committeemember":["Stavness, Ian","Keil, Mark","Nguyen, Ha"],"dc:creator":["Mouli, Srishti"],"dc:date.accessioned":["2022-04-05T15:01:23Z"],"dc:date.available":["2022-04-05T15:01:23Z"],"dc:date.issued":["2022-04-05"],"dc:description.abstract":["To observe the growth dynamics of the canola flowers during the blooming season and estimate the harvest forecast of the Canola crops, an application called ‘Flower Counter’ has been developed by the researchers of P2IRC located at the University of Saskatchewan. The model has been developed using Deep Learning (DL) based Multi-column Convolutional Neural Network (MCNN) algorithm and TensorFlow framework. This is an object counting model, that counts the Canola flowers from the images based on the learning from a given set of training images, called ‘ground-truths’. This work proposes to compose a good training dataset that would give good accuracy with a robust object detection model by using different training and testing combinations. Various evaluation techniques have been used in this work to check the impact of the training dataset, on the testing results of the model and generalizability. The primary goal of this research work is to define a good training dataset composition having diversity. A good composition also consists of different characteristics present in the dataset, that can impact the testing results and can help in creating a robust object counting model. Different characteristics of the training datasets and testing datasets are used to evaluate the most prominent characteristics and features that impact the test results. The objective is also to evaluate the impact of training dataset selection on testing results produced by the ML model in terms of accuracy. This work would help the researchers and plant scientists gain knowledge about the diversity of characteristics for the composition of a training dataset. This can give insights to reduce the manual effort which is required to create ground truth for training models by identifying the characteristics that impact testing results. Since the entire training of the model depends on the datasets collected during diverse weather conditions, there could be factors that could impact some of the experimental results. The research area for training dataset selection has not been explored much, and this research work will give good insights about model generalization capability and scopes for manual work utilization for getting a robust object counting model."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/10388/13864"],"dc:subject":["Machine Learning, Training Dataset selection"],"dc:title":["Influence of training dataset selection on the performance of a machine learning model"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Masters"],"thesis:degree_name":["Master of Science (M.Sc.)"],"thesis:institution_name":["University of Saskatchewan"]},"updated_at":"2026-07-24T04:27:06Z"}