{"id":{"repo_id":"claremont","oai_identifier":"oai:scholarship.claremont.edu:cgu_etd-1836"},"canonical_url":"https://search.dev.ndltd.org/etd/claremont/oai:scholarship.claremont.edu:cgu_etd-1836","repository":{"repo_id":"claremont","name":"Claremont Graduate University","base_url":"https://scholarship.claremont.edu/do/oai/"},"display":{"title":"Improveing F-beta Score in Classifying Shark Data into Shark Behaviors","abstract":"<p>One metric used to measure classification performance in machine learning is F-beta score. The objective in this thesis is to improve the average F-b score computed in classifying shark data into shark behaviors, namely; Resting, Swimming, Feeding, and Non-Directed Motion (NDM). Synthetic Minority Oversampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN) are utilized to balance the data, from which pre-processed Fast Fourier Transform (FFT), Walsh-Hadamard Transform (WHT), and Autocorrelation (AC) features are extracted then classified using Convolutional Neural Network (CNN) and K-Nearest Neighbors (K-NN). All the combinations of the two balancing techniques, the three feature types, and the two machine learning algorithms are applied then compared to examine the average F-beta score improvement. Other signal processing techniques are also applied, to reduce the noise level of the recorded raw shark data and enhance its Signal-to-Noise Ratio (SNR).</p> <p>The average F-beta scores showed that K-NN performed at its best when using FFT-only features while CNN performed at its best when using WHT-FFT features. In the K-NN case, FFT performed better when it was used alone than when it was combined with any other feature type. On the other hand, WHT performed better when it was combined with any other feature type than when it was used alone. In the CNN case, WHT and FFT performed better together than they did separately. In other words, Combining FFT and WHT features in CNN resulted in considerably improved average F-beta score, while combining them in K-NN averaged their scores. Also, whether alone or combined with other feature types, AC did not work well in CNN as it resulted in poor average F-beta scores. In K-NN, combining AC with other feature types did not improve the average F-beta score from when it is used alone.</p> <p>The average F-beta scores also showed that reducing the data imbalance nature during the pre-processing phase is more effective than mitigating the misleading classification during the machine learning phase. Prior balancing was performed using SMOTE and ADASYN, while later mitigation was performed using weight-sensitive learning. SMOTE, more so ADASYN, reduced the difference between precision and recall scores, and produced higher F-beta scores.</p> <p>Besides the mentioned two balancing techniques, the three feature types, and the two machine learning algorithms, other pre-processing techniques that were applied to the raw data contributed to the improvement of the average F-beta score. These pre-processing techniques included framing, detrending, normalization, Ensemble Average (EA) based low-pass filtering, filter delay compensation, overlap windowing, and <em>k</em>-fold cross validation. For example, the average F-beta scores showed that applying EA-based low-pass filters (LPF) on the data, prior to machine learning and classification, improves Signal Power to Noise Power Ratio (SNR), and sequentially improves average F-beat scores significantly.</p> <p>As an end result, for the shark data used in this thesis, CNN was found to be a better choice than K-NN, and it was a better choice when using WHT-FFT as features and ADASYN as balancing technique.</p>","abstract_html":"&lt;p&gt;One metric used to measure classification performance in machine learning is F-beta score. The objective in this thesis is to improve the average F-b score computed in classifying shark data into shark behaviors, namely; Resting, Swimming, Feeding, and Non-Directed Motion (NDM). Synthetic Minority Oversampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN) are utilized to balance the data, from which pre-processed Fast Fourier Transform (FFT), Walsh-Hadamard Transform (WHT), and Autocorrelation (AC) features are extracted then classified using Convolutional Neural Network (CNN) and K-Nearest Neighbors (K-NN). All the combinations of the two balancing techniques, the three feature types, and the two machine learning algorithms are applied then compared to examine the average F-beta score improvement. Other signal processing techniques are also applied, to reduce the noise level of the recorded raw shark data and enhance its Signal-to-Noise Ratio (SNR).&lt;/p&gt; &lt;p&gt;The average F-beta scores showed that K-NN performed at its best when using FFT-only features while CNN performed at its best when using WHT-FFT features. In the K-NN case, FFT performed better when it was used alone than when it was combined with any other feature type. On the other hand, WHT performed better when it was combined with any other feature type than when it was used alone. In the CNN case, WHT and FFT performed better together than they did separately. In other words, Combining FFT and WHT features in CNN resulted in considerably improved average F-beta score, while combining them in K-NN averaged their scores. Also, whether alone or combined with other feature types, AC did not work well in CNN as it resulted in poor average F-beta scores. In K-NN, combining AC with other feature types did not improve the average F-beta score from when it is used alone.&lt;/p&gt; &lt;p&gt;The average F-beta scores also showed that reducing the data imbalance nature during the pre-processing phase is more effective than mitigating the misleading classification during the machine learning phase. Prior balancing was performed using SMOTE and ADASYN, while later mitigation was performed using weight-sensitive learning. SMOTE, more so ADASYN, reduced the difference between precision and recall scores, and produced higher F-beta scores.&lt;/p&gt; &lt;p&gt;Besides the mentioned two balancing techniques, the three feature types, and the two machine learning algorithms, other pre-processing techniques that were applied to the raw data contributed to the improvement of the average F-beta score. These pre-processing techniques included framing, detrending, normalization, Ensemble Average (EA) based low-pass filtering, filter delay compensation, overlap windowing, and &lt;em&gt;k&lt;/em&gt;-fold cross validation. For example, the average F-beta scores showed that applying EA-based low-pass filters (LPF) on the data, prior to machine learning and classification, improves Signal Power to Noise Power Ratio (SNR), and sequentially improves average F-beat scores significantly.&lt;/p&gt; &lt;p&gt;As an end result, for the shark data used in this thesis, CNN was found to be a better choice than K-NN, and it was a better choice when using WHT-FFT as features and ADASYN as balancing technique.&lt;/p&gt;","abstract_has_math":false,"creators":["Ali, Ibrahim M"],"institution":null,"degree_name":"Computational Science Joint PhD with San Diego State University, PhD","degree_level":"Open Access Dissertation","degree_discipline":"Institute of Mathematical Sciences","degree_department":null,"school":null,"contributors":["Marina Chugunova","Yu Yang","Qidi Peng"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-01-01T08:00:00Z","date_published":"2024-01-01T08:00:00Z","updated_at":"2026-07-24T01:40:36Z","subjects":["Activity Recognition","CNN K-NN","Digital Signal Processing","Imbalanced data","Machine Learning","SMOTE ADASYN","Engineering","Mathematics"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://scholarship.claremont.edu/cgu_etd/814","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Marina Chugunova","Yu Yang","Qidi Peng"]},{"key":"dc:creator","label":"Author","values":["Ali, Ibrahim M"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2024-06-28T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Institute of Mathematical Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Open Access Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Computational Science Joint PhD with San Diego State University, PhD"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Activity Recognition","CNN K-NN","Digital Signal Processing","Imbalanced data","Machine Learning","SMOTE ADASYN","Engineering","Mathematics"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://scholarship.claremont.edu/cgu_etd/814"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>One metric used to measure classification performance in machine learning is F-beta score. The objective in this thesis is to improve the average F-b score computed in classifying shark data into shark behaviors, namely; Resting, Swimming, Feeding, and Non-Directed Motion (NDM). Synthetic Minority Oversampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN) are utilized to balance the data, from which pre-processed Fast Fourier Transform (FFT), Walsh-Hadamard Transform (WHT), and Autocorrelation (AC) features are extracted then classified using Convolutional Neural Network (CNN) and K-Nearest Neighbors (K-NN). All the combinations of the two balancing techniques, the three feature types, and the two machine learning algorithms are applied then compared to examine the average F-beta score improvement. Other signal processing techniques are also applied, to reduce the noise level of the recorded raw shark data and enhance its Signal-to-Noise Ratio (SNR).</p> <p>The average F-beta scores showed that K-NN performed at its best when using FFT-only features while CNN performed at its best when using WHT-FFT features. In the K-NN case, FFT performed better when it was used alone than when it was combined with any other feature type. On the other hand, WHT performed better when it was combined with any other feature type than when it was used alone. In the CNN case, WHT and FFT performed better together than they did separately. In other words, Combining FFT and WHT features in CNN resulted in considerably improved average F-beta score, while combining them in K-NN averaged their scores. Also, whether alone or combined with other feature types, AC did not work well in CNN as it resulted in poor average F-beta scores. In K-NN, combining AC with other feature types did not improve the average F-beta score from when it is used alone.</p> <p>The average F-beta scores also showed that reducing the data imbalance nature during the pre-processing phase is more effective than mitigating the misleading classification during the machine learning phase. Prior balancing was performed using SMOTE and ADASYN, while later mitigation was performed using weight-sensitive learning. SMOTE, more so ADASYN, reduced the difference between precision and recall scores, and produced higher F-beta scores.</p> <p>Besides the mentioned two balancing techniques, the three feature types, and the two machine learning algorithms, other pre-processing techniques that were applied to the raw data contributed to the improvement of the average F-beta score. These pre-processing techniques included framing, detrending, normalization, Ensemble Average (EA) based low-pass filtering, filter delay compensation, overlap windowing, and <em>k</em>-fold cross validation. For example, the average F-beta scores showed that applying EA-based low-pass filters (LPF) on the data, prior to machine learning and classification, improves Signal Power to Noise Power Ratio (SNR), and sequentially improves average F-beat scores significantly.</p> <p>As an end result, for the shark data used in this thesis, CNN was found to be a better choice than K-NN, and it was a better choice when using WHT-FFT as features and ADASYN as balancing technique.</p>"]},{"key":"dc:title","label":"Title","values":["Improveing F-beta Score in Classifying Shark Data into Shark Behaviors"]}]}],"canonical_facts":{"dc:contributor":["Marina Chugunova","Yu Yang","Qidi Peng"],"dc:creator":["Ali, Ibrahim M"],"dc:date.available":["2024-06-28T07:00:00Z"],"dc:description.abstract":["<p>One metric used to measure classification performance in machine learning is F-beta score. The objective in this thesis is to improve the average F-b score computed in classifying shark data into shark behaviors, namely; Resting, Swimming, Feeding, and Non-Directed Motion (NDM). Synthetic Minority Oversampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN) are utilized to balance the data, from which pre-processed Fast Fourier Transform (FFT), Walsh-Hadamard Transform (WHT), and Autocorrelation (AC) features are extracted then classified using Convolutional Neural Network (CNN) and K-Nearest Neighbors (K-NN). All the combinations of the two balancing techniques, the three feature types, and the two machine learning algorithms are applied then compared to examine the average F-beta score improvement. Other signal processing techniques are also applied, to reduce the noise level of the recorded raw shark data and enhance its Signal-to-Noise Ratio (SNR).</p> <p>The average F-beta scores showed that K-NN performed at its best when using FFT-only features while CNN performed at its best when using WHT-FFT features. In the K-NN case, FFT performed better when it was used alone than when it was combined with any other feature type. On the other hand, WHT performed better when it was combined with any other feature type than when it was used alone. In the CNN case, WHT and FFT performed better together than they did separately. In other words, Combining FFT and WHT features in CNN resulted in considerably improved average F-beta score, while combining them in K-NN averaged their scores. Also, whether alone or combined with other feature types, AC did not work well in CNN as it resulted in poor average F-beta scores. In K-NN, combining AC with other feature types did not improve the average F-beta score from when it is used alone.</p> <p>The average F-beta scores also showed that reducing the data imbalance nature during the pre-processing phase is more effective than mitigating the misleading classification during the machine learning phase. Prior balancing was performed using SMOTE and ADASYN, while later mitigation was performed using weight-sensitive learning. SMOTE, more so ADASYN, reduced the difference between precision and recall scores, and produced higher F-beta scores.</p> <p>Besides the mentioned two balancing techniques, the three feature types, and the two machine learning algorithms, other pre-processing techniques that were applied to the raw data contributed to the improvement of the average F-beta score. These pre-processing techniques included framing, detrending, normalization, Ensemble Average (EA) based low-pass filtering, filter delay compensation, overlap windowing, and <em>k</em>-fold cross validation. For example, the average F-beta scores showed that applying EA-based low-pass filters (LPF) on the data, prior to machine learning and classification, improves Signal Power to Noise Power Ratio (SNR), and sequentially improves average F-beat scores significantly.</p> <p>As an end result, for the shark data used in this thesis, CNN was found to be a better choice than K-NN, and it was a better choice when using WHT-FFT as features and ADASYN as balancing technique.</p>"],"dc:identifier":["https://scholarship.claremont.edu/cgu_etd/814"],"dc:subject":["Activity Recognition","CNN K-NN","Digital Signal Processing","Imbalanced data","Machine Learning","SMOTE ADASYN","Engineering","Mathematics"],"dc:title":["Improveing F-beta Score in Classifying Shark Data into Shark Behaviors"],"thesis:degree_discipline":["Institute of Mathematical Sciences"],"thesis:degree_level":["Open Access Dissertation"],"thesis:degree_name":["Computational Science Joint PhD with San Diego State University, PhD"]},"updated_at":"2026-07-24T01:40:36Z"}