{"id":{"repo_id":"calpoly","oai_identifier":"oai:digitalcommons.calpoly.edu:theses-4342"},"canonical_url":"https://search.dev.ndltd.org/etd/calpoly/oai:digitalcommons.calpoly.edu:theses-4342","repository":{"repo_id":"calpoly","name":"Cal Poly","base_url":"https://digitalcommons.calpoly.edu/do/oai/"},"display":{"title":"Enhancing Telecom Churn Prediction: Adaboost with Oversampling and Recursive Feature Elimination Approach","abstract":"<p>Churn prediction is a critical task for businesses to retain their valuable customers. This paper presents a comprehensive study of churn prediction in the telecom sector using 15 approaches, including popular algorithms such as Logistic Regression, Support Vector Machine, Decision Tree, Random Forest, and AdaBoost.</p> <p>The study is segmented into three sets of experiments, each focusing on a different approach to building the churn prediction model. The model is constructed using the original training set in the first set of experiments. The second set involves oversampling the training set to address the issue of imbalanced data. Lastly, the third set combines oversampling with recursive feature selection to enhance the model's performance further.</p> <p>The results demonstrate that the Adaptive Boost classifier, implemented with oversampling and recursive feature selection, outperforms the other 14 techniques. It achieves the highest rank in all three evaluation metrics: recall (0.841), f1-score (0.655), and roc_auc (0.793), further indicating that the proposed approach effectively predicts churn and provides valuable insights into customer behavior.</p>","abstract_html":"&lt;p&gt;Churn prediction is a critical task for businesses to retain their valuable customers. This paper presents a comprehensive study of churn prediction in the telecom sector using 15 approaches, including popular algorithms such as Logistic Regression, Support Vector Machine, Decision Tree, Random Forest, and AdaBoost.&lt;/p&gt; &lt;p&gt;The study is segmented into three sets of experiments, each focusing on a different approach to building the churn prediction model. The model is constructed using the original training set in the first set of experiments. The second set involves oversampling the training set to address the issue of imbalanced data. Lastly, the third set combines oversampling with recursive feature selection to enhance the model&#x27;s performance further.&lt;/p&gt; &lt;p&gt;The results demonstrate that the Adaptive Boost classifier, implemented with oversampling and recursive feature selection, outperforms the other 14 techniques. It achieves the highest rank in all three evaluation metrics: recall (0.841), f1-score (0.655), and roc_auc (0.793), further indicating that the proposed approach effectively predicts churn and provides valuable insights into customer behavior.&lt;/p&gt;","abstract_has_math":false,"creators":["Tran, Long Dinh"],"institution":null,"degree_name":"MS in Electrical Engineering","degree_level":null,"degree_discipline":"Electrical Engineering","degree_department":null,"school":null,"contributors":["Xiaozheng (Jane) Zhang","Electrical Engineering","College of Engineering"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-06-01T07:00:00Z","date_published":"2023-06-01T07:00:00Z","updated_at":"2026-07-24T01:32:29Z","subjects":["Churn Prediction","Unbalanced Datasets","Oversampling","SMOTE","Recursive Feature Selection","RFE","Machine Learning","Other Electrical and Computer Engineering"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.calpoly.edu/theses/2658","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Xiaozheng (Jane) Zhang","Electrical Engineering","College of Engineering"]},{"key":"dc:creator","label":"Author","values":["Tran, Long Dinh"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2023-07-05T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical Engineering"]},{"key":"thesis:degree_name","label":"Degree Name","values":["MS in Electrical Engineering"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Churn Prediction","Unbalanced Datasets","Oversampling","SMOTE","Recursive Feature Selection","RFE","Machine Learning","Other Electrical and Computer Engineering"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.calpoly.edu/theses/2658"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Churn prediction is a critical task for businesses to retain their valuable customers. This paper presents a comprehensive study of churn prediction in the telecom sector using 15 approaches, including popular algorithms such as Logistic Regression, Support Vector Machine, Decision Tree, Random Forest, and AdaBoost.</p> <p>The study is segmented into three sets of experiments, each focusing on a different approach to building the churn prediction model. The model is constructed using the original training set in the first set of experiments. The second set involves oversampling the training set to address the issue of imbalanced data. Lastly, the third set combines oversampling with recursive feature selection to enhance the model's performance further.</p> <p>The results demonstrate that the Adaptive Boost classifier, implemented with oversampling and recursive feature selection, outperforms the other 14 techniques. It achieves the highest rank in all three evaluation metrics: recall (0.841), f1-score (0.655), and roc_auc (0.793), further indicating that the proposed approach effectively predicts churn and provides valuable insights into customer behavior.</p>"]},{"key":"dc:title","label":"Title","values":["Enhancing Telecom Churn Prediction: Adaboost with Oversampling and Recursive Feature Elimination Approach"]}]}],"canonical_facts":{"dc:contributor":["Xiaozheng (Jane) Zhang","Electrical Engineering","College of Engineering"],"dc:creator":["Tran, Long Dinh"],"dc:date.available":["2023-07-05T07:00:00Z"],"dc:description.abstract":["<p>Churn prediction is a critical task for businesses to retain their valuable customers. This paper presents a comprehensive study of churn prediction in the telecom sector using 15 approaches, including popular algorithms such as Logistic Regression, Support Vector Machine, Decision Tree, Random Forest, and AdaBoost.</p> <p>The study is segmented into three sets of experiments, each focusing on a different approach to building the churn prediction model. The model is constructed using the original training set in the first set of experiments. The second set involves oversampling the training set to address the issue of imbalanced data. Lastly, the third set combines oversampling with recursive feature selection to enhance the model's performance further.</p> <p>The results demonstrate that the Adaptive Boost classifier, implemented with oversampling and recursive feature selection, outperforms the other 14 techniques. It achieves the highest rank in all three evaluation metrics: recall (0.841), f1-score (0.655), and roc_auc (0.793), further indicating that the proposed approach effectively predicts churn and provides valuable insights into customer behavior.</p>"],"dc:identifier":["https://digitalcommons.calpoly.edu/theses/2658"],"dc:subject":["Churn Prediction","Unbalanced Datasets","Oversampling","SMOTE","Recursive Feature Selection","RFE","Machine Learning","Other Electrical and Computer Engineering"],"dc:title":["Enhancing Telecom Churn Prediction: Adaboost with Oversampling and Recursive Feature Elimination Approach"],"thesis:degree_discipline":["Electrical Engineering"],"thesis:degree_name":["MS in Electrical Engineering"]},"updated_at":"2026-07-24T01:32:29Z"}