{"id":{"repo_id":"umkc","oai_identifier":"oai:mospace.umsystem.edu:10355/74350"},"canonical_url":"https://search.dev.ndltd.org/etd/umkc/oai:mospace.umsystem.edu:10355/74350","repository":{"repo_id":"umkc","name":"University of Missouri - Kansas City","base_url":"https://mospace.umsystem.edu/oai/request"},"display":{"title":"DeepSampling: Image Sampling Technique for Cost-Effective Deep Learning","abstract":"Deep learning is beneficial from big data while facing computationally expensive, with an increase in data size. Some severe data issues, such as the presence of highly skewed, sparse, and imbalanced data, would substantially influence the findings of machine learning. Due to the complexity of such data, the ability to assess and evaluate the data is central to cost-effective deep learning. More specifically, in Deep Learning, choosing the right validation method is vital to ensure the accuracy and biases of the validation process. Current validation techniques, including k-fold cross-validation or random split of training and testing datasets, are hampered by the lack of systematic sampling with a comprehensive understanding of the data. In this thesis, we proposed a sampling technique called DeepSampling that aims at achieving cost-effective deep learning for a given application. For the proposed DeepSampling framework, two sampling schemes are designed [1] to resolve the imbalanced data issues using Generative Adversarial Networks (GANs), [2] to develop an effective sampling technique based on clustering. The clustering techniques are based on Mahalanobis distance metric and use t-SNE (T-distributed Stochastic Neighbor Embedding), to overcome the data skewness and sparseness issues. The proposed DeepSampling technique for cost-effective deep learning has been evaluated with three Deep Learning models and four benchmark datasets, including MNIST, Breast Histology, Malaria cell images, and Stanford dog. The results confirm that the accuracies obtained by DeepSampling are improved by approximately 2-3% for image classification, compared to traditional evaluation techniques on the same dataset.","abstract_html":"Deep learning is beneficial from big data while facing computationally expensive, with an increase in data size. Some severe data issues, such as the presence of highly skewed, sparse, and imbalanced data, would substantially influence the findings of machine learning. Due to the complexity of such data, the ability to assess and evaluate the data is central to cost-effective deep learning. More specifically, in Deep Learning, choosing the right validation method is vital to ensure the accuracy and biases of the validation process. Current validation techniques, including k-fold cross-validation or random split of training and testing datasets, are hampered by the lack of systematic sampling with a comprehensive understanding of the data. In this thesis, we proposed a sampling technique called DeepSampling that aims at achieving cost-effective deep learning for a given application. For the proposed DeepSampling framework, two sampling schemes are designed [1] to resolve the imbalanced data issues using Generative Adversarial Networks (GANs), [2] to develop an effective sampling technique based on clustering. The clustering techniques are based on Mahalanobis distance metric and use t-SNE (T-distributed Stochastic Neighbor Embedding), to overcome the data skewness and sparseness issues. The proposed DeepSampling technique for cost-effective deep learning has been evaluated with three Deep Learning models and four benchmark datasets, including MNIST, Breast Histology, Malaria cell images, and Stanford dog. The results confirm that the accuracies obtained by DeepSampling are improved by approximately 2-3% for image classification, compared to traditional evaluation techniques on the same dataset.","abstract_has_math":false,"creators":["Gaikwad, Priyanka V."],"institution":"University of Missouri--Kansas City","degree_name":"M.S. (Master of Science)","degree_level":"Masters","degree_discipline":"Computer Science (UMKC)","degree_department":null,"school":null,"contributors":[],"advisors":["Lee, Yugyung, 1960-"],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020","date_published":"2020","updated_at":"2026-07-24T05:16:20Z","subjects":[],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10355/74350","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Lee, Yugyung, 1960-"]},{"key":"dc:creator","label":"Author","values":["Gaikwad, Priyanka V."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2020-06-22T22:13:17Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2020-06-22T22:13:17Z"]},{"key":"dc:date.issued","label":"Date","values":["2020"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science (UMKC)"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S. (Master of Science)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Missouri--Kansas City"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10355/74350"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Title from PDF of title page viewed June 24, 2020","Thesis advisor: Yugyung Lee","Vita","Includes bibliographical references (pages 44-46)","Thesis (M.S.)--School of Computing and Engineering. University of Missouri--Kansas City, 2020"]},{"key":"dc:description.abstract","label":"Abstract","values":["Deep learning is beneficial from big data while facing computationally expensive, with an increase in data size. Some severe data issues, such as the presence of highly skewed, sparse, and imbalanced data, would substantially influence the findings of machine learning. Due to the complexity of such data, the ability to assess and evaluate the data is central to cost-effective deep learning. More specifically, in Deep Learning, choosing the right validation method is vital to ensure the accuracy and biases of the validation process. Current validation techniques, including k-fold cross-validation or random split of training and testing datasets, are hampered by the lack of systematic sampling with a comprehensive understanding of the data. In this thesis, we proposed a sampling technique called DeepSampling that aims at achieving cost-effective deep learning for a given application. For the proposed DeepSampling framework, two sampling schemes are designed [1] to resolve the imbalanced data issues using Generative Adversarial Networks (GANs), [2] to develop an effective sampling technique based on clustering. The clustering techniques are based on Mahalanobis distance metric and use t-SNE (T-distributed Stochastic Neighbor Embedding), to overcome the data skewness and sparseness issues. The proposed DeepSampling technique for cost-effective deep learning has been evaluated with three Deep Learning models and four benchmark datasets, including MNIST, Breast Histology, Malaria cell images, and Stanford dog. The results confirm that the accuracies obtained by DeepSampling are improved by approximately 2-3% for image classification, compared to traditional evaluation techniques on the same dataset."]},{"key":"dc:title","label":"Title","values":["DeepSampling: Image Sampling Technique for Cost-Effective Deep Learning"]}]}],"canonical_facts":{"dc:contributor.advisor":["Lee, Yugyung, 1960-"],"dc:creator":["Gaikwad, Priyanka V."],"dc:date.accessioned":["2020-06-22T22:13:17Z"],"dc:date.available":["2020-06-22T22:13:17Z"],"dc:date.issued":["2020"],"dc:description":["Title from PDF of title page viewed June 24, 2020","Thesis advisor: Yugyung Lee","Vita","Includes bibliographical references (pages 44-46)","Thesis (M.S.)--School of Computing and Engineering. University of Missouri--Kansas City, 2020"],"dc:description.abstract":["Deep learning is beneficial from big data while facing computationally expensive, with an increase in data size. Some severe data issues, such as the presence of highly skewed, sparse, and imbalanced data, would substantially influence the findings of machine learning. Due to the complexity of such data, the ability to assess and evaluate the data is central to cost-effective deep learning. More specifically, in Deep Learning, choosing the right validation method is vital to ensure the accuracy and biases of the validation process. Current validation techniques, including k-fold cross-validation or random split of training and testing datasets, are hampered by the lack of systematic sampling with a comprehensive understanding of the data. In this thesis, we proposed a sampling technique called DeepSampling that aims at achieving cost-effective deep learning for a given application. For the proposed DeepSampling framework, two sampling schemes are designed [1] to resolve the imbalanced data issues using Generative Adversarial Networks (GANs), [2] to develop an effective sampling technique based on clustering. The clustering techniques are based on Mahalanobis distance metric and use t-SNE (T-distributed Stochastic Neighbor Embedding), to overcome the data skewness and sparseness issues. The proposed DeepSampling technique for cost-effective deep learning has been evaluated with three Deep Learning models and four benchmark datasets, including MNIST, Breast Histology, Malaria cell images, and Stanford dog. The results confirm that the accuracies obtained by DeepSampling are improved by approximately 2-3% for image classification, compared to traditional evaluation techniques on the same dataset."],"dc:identifier.uri":["https://hdl.handle.net/10355/74350"],"dc:title":["DeepSampling: Image Sampling Technique for Cost-Effective Deep Learning"],"thesis:degree_discipline":["Computer Science (UMKC)"],"thesis:degree_level":["Masters"],"thesis:degree_name":["M.S. (Master of Science)"],"thesis:institution_name":["University of Missouri--Kansas City"]},"updated_at":"2026-07-24T05:16:20Z"}