{"id":{"repo_id":"kennesaw","oai_identifier":"oai:digitalcommons.kennesaw.edu:dataphd_etd-1004"},"canonical_url":"https://search.dev.ndltd.org/etd/kennesaw/oai:digitalcommons.kennesaw.edu:dataphd_etd-1004","repository":{"repo_id":"kennesaw","name":"Kennesaw State University","base_url":"https://digitalcommons.kennesaw.edu/do/oai/"},"display":{"title":"A Novel Penalized Log-likelihood Function for Class Imbalance Problem","abstract":"<p>The log-likelihood function is the optimization objective in the maximum likelihood method for estimating models (e.g., logistic regression, neural network). However, its formulation is based on assumptions that the target classes are equally distributed and the overall accuracy is maximized, which do not apply to class imbalance problems (e.g., fraud detection, rare disease diagnoses, customer conversion prediction, cybersecurity, predictive maintenance). When trained on imbalanced data, the resulting models tend to be biased towards the majority class (i.e. non-event), which can bring great loss in practice. One strategy for mitigating such bias is to penalize the misclassification costs of observations differently in the log-likelihood objective function in the learning process. Existing penalized log-likelihood functions require either hard hyperparameter estimation or high computational complexity. In the present work, we propose a novel penalized log-likelihood function by including penalty weights as decision variables for observations in the minority class (i.e. event) and learning them from data along with model coefficients/parameters. The proposed log-likelihood function is applied to train logistic regression and neural network models, which are compared with models trained by existing penalized log-likelihood functions on 10 public imbalanced datasets. The model performance is measured by the statistics of Area under ROC Curve (i.e. AUROC or AUC) over repeated runs of 10-fold stratified cross validation, including 95\\% confidence interval, mean and standard deviation, as well as the training time. A more detailed analysis is conducted to examine the estimated probability distributions and additional performance measurements (i.e. Type I error, Type II error, accuracy) under the chosen probability cutoff. The results demonstrate that the discrimination ability of the models is improved by using the proposed log-likelihood function as the learning objective while reducing or maintaining the computational complexity compared with existing ones.</p>","abstract_html":"&lt;p&gt;The log-likelihood function is the optimization objective in the maximum likelihood method for estimating models (e.g., logistic regression, neural network). However, its formulation is based on assumptions that the target classes are equally distributed and the overall accuracy is maximized, which do not apply to class imbalance problems (e.g., fraud detection, rare disease diagnoses, customer conversion prediction, cybersecurity, predictive maintenance). When trained on imbalanced data, the resulting models tend to be biased towards the majority class (i.e. non-event), which can bring great loss in practice. One strategy for mitigating such bias is to penalize the misclassification costs of observations differently in the log-likelihood objective function in the learning process. Existing penalized log-likelihood functions require either hard hyperparameter estimation or high computational complexity. In the present work, we propose a novel penalized log-likelihood function by including penalty weights as decision variables for observations in the minority class (i.e. event) and learning them from data along with model coefficients/parameters. The proposed log-likelihood function is applied to train logistic regression and neural network models, which are compared with models trained by existing penalized log-likelihood functions on 10 public imbalanced datasets. The model performance is measured by the statistics of Area under ROC Curve (i.e. AUROC or AUC) over repeated runs of 10-fold stratified cross validation, including 95\\% confidence interval, mean and standard deviation, as well as the training time. A more detailed analysis is conducted to examine the estimated probability distributions and additional performance measurements (i.e. Type I error, Type II error, accuracy) under the chosen probability cutoff. The results demonstrate that the discrimination ability of the models is improved by using the proposed log-likelihood function as the learning objective while reducing or maintaining the computational complexity compared with existing ones.&lt;/p&gt;","abstract_has_math":false,"creators":["Zhang, Lili"],"institution":null,"degree_name":"Doctor of Philosophy in Analytic and Data Science","degree_level":"Dissertation","degree_discipline":"Statistics and Analytical Sciences","degree_department":null,"school":null,"contributors":["Herman Ray","Joseph DeMaio","Lin Li","Sherry Ni","Ying Xie"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-03-24T07:00:00Z","date_published":"2020-03-24T07:00:00Z","updated_at":"2026-07-24T02:43:33Z","subjects":["Penalized Log-likelihood Function","Class Imbalance Problem","Logistic Regression","Neural Network","Binary Classification","Business Analytics","Statistics and Probability"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.kennesaw.edu/dataphd_etd/5","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Herman Ray","Joseph DeMaio","Lin Li","Sherry Ni","Ying Xie"]},{"key":"dc:creator","label":"Author","values":["Zhang, Lili"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2021-12-31T08:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics and Analytical Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy in Analytic and Data Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Penalized Log-likelihood Function","Class Imbalance Problem","Logistic Regression","Neural Network","Binary Classification","Business Analytics","Statistics and Probability"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.kennesaw.edu/dataphd_etd/5"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>The log-likelihood function is the optimization objective in the maximum likelihood method for estimating models (e.g., logistic regression, neural network). However, its formulation is based on assumptions that the target classes are equally distributed and the overall accuracy is maximized, which do not apply to class imbalance problems (e.g., fraud detection, rare disease diagnoses, customer conversion prediction, cybersecurity, predictive maintenance). When trained on imbalanced data, the resulting models tend to be biased towards the majority class (i.e. non-event), which can bring great loss in practice. One strategy for mitigating such bias is to penalize the misclassification costs of observations differently in the log-likelihood objective function in the learning process. Existing penalized log-likelihood functions require either hard hyperparameter estimation or high computational complexity. In the present work, we propose a novel penalized log-likelihood function by including penalty weights as decision variables for observations in the minority class (i.e. event) and learning them from data along with model coefficients/parameters. The proposed log-likelihood function is applied to train logistic regression and neural network models, which are compared with models trained by existing penalized log-likelihood functions on 10 public imbalanced datasets. The model performance is measured by the statistics of Area under ROC Curve (i.e. AUROC or AUC) over repeated runs of 10-fold stratified cross validation, including 95\\% confidence interval, mean and standard deviation, as well as the training time. A more detailed analysis is conducted to examine the estimated probability distributions and additional performance measurements (i.e. Type I error, Type II error, accuracy) under the chosen probability cutoff. The results demonstrate that the discrimination ability of the models is improved by using the proposed log-likelihood function as the learning objective while reducing or maintaining the computational complexity compared with existing ones.</p>"]},{"key":"dc:title","label":"Title","values":["A Novel Penalized Log-likelihood Function for Class Imbalance Problem"]}]}],"canonical_facts":{"dc:contributor":["Herman Ray","Joseph DeMaio","Lin Li","Sherry Ni","Ying Xie"],"dc:creator":["Zhang, Lili"],"dc:date.available":["2021-12-31T08:00:00Z"],"dc:description.abstract":["<p>The log-likelihood function is the optimization objective in the maximum likelihood method for estimating models (e.g., logistic regression, neural network). However, its formulation is based on assumptions that the target classes are equally distributed and the overall accuracy is maximized, which do not apply to class imbalance problems (e.g., fraud detection, rare disease diagnoses, customer conversion prediction, cybersecurity, predictive maintenance). When trained on imbalanced data, the resulting models tend to be biased towards the majority class (i.e. non-event), which can bring great loss in practice. One strategy for mitigating such bias is to penalize the misclassification costs of observations differently in the log-likelihood objective function in the learning process. Existing penalized log-likelihood functions require either hard hyperparameter estimation or high computational complexity. In the present work, we propose a novel penalized log-likelihood function by including penalty weights as decision variables for observations in the minority class (i.e. event) and learning them from data along with model coefficients/parameters. The proposed log-likelihood function is applied to train logistic regression and neural network models, which are compared with models trained by existing penalized log-likelihood functions on 10 public imbalanced datasets. The model performance is measured by the statistics of Area under ROC Curve (i.e. AUROC or AUC) over repeated runs of 10-fold stratified cross validation, including 95\\% confidence interval, mean and standard deviation, as well as the training time. A more detailed analysis is conducted to examine the estimated probability distributions and additional performance measurements (i.e. Type I error, Type II error, accuracy) under the chosen probability cutoff. The results demonstrate that the discrimination ability of the models is improved by using the proposed log-likelihood function as the learning objective while reducing or maintaining the computational complexity compared with existing ones.</p>"],"dc:identifier":["https://digitalcommons.kennesaw.edu/dataphd_etd/5"],"dc:subject":["Penalized Log-likelihood Function","Class Imbalance Problem","Logistic Regression","Neural Network","Binary Classification","Business Analytics","Statistics and Probability"],"dc:title":["A Novel Penalized Log-likelihood Function for Class Imbalance Problem"],"thesis:degree_discipline":["Statistics and Analytical Sciences"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Doctor of Philosophy in Analytic and Data Science"]},"updated_at":"2026-07-24T02:43:33Z"}