{"id":{"repo_id":"denver","oai_identifier":"oai:digitalcommons.du.edu:etd-3160"},"canonical_url":"https://search.dev.ndltd.org/etd/denver/oai:digitalcommons.du.edu:etd-3160","repository":{"repo_id":"denver","name":"University of Denver","base_url":"https://digitalcommons.du.edu/do/oai/"},"display":{"title":"A Comparison of Logistic, RIDGE, and LASSO Regression with Heart Failure Risk Data: Effects of Sample Size, Predictor Correlation, and Predictor Weight on Outcome Accuracy","abstract":"<p>Logistic Regression (LR), LASSO regression, and RIDGE regression are standard classification techniques for predicting a dichotomous output. Since these methods are applied for similar purposes and have different features, it is crucial to evaluate the performance of these methods under different controlled conditions. With this information, researchers can apply the optimal method for specific conditions.</p> <p>Following previous research, which reported the effects of conditions such as sample size and multicollinearity on the performance of the classification methods, this research focused on the effects of when sample size, level of predictor collinearity, and predictor variable weight are controlled on the performance of LR, LASSO, and RIDGE regressions. Data were simulated with 100 iterations that generated a total of n = 2,400 observations in R statistical software. A factorial ANOVA with follow-ups was employed to evaluate the effect of conditions on the performance of each technique as measured by accuracy and F-measure.</p> <p>In most conditions for the two outcome performance measures (accuracy and F-measure), the highest effect on performances was observed from the predictor variable weight. However, when the weight was low, all three regression methods were found to have an overall better performance under high correlation and a large sample size. Moreover, the models with high-weight conditions suppressed the effects of every other controlled condition on accuracy and F-measure output values. Therefore, when the study data conditions include a high-weighted variable, regardless of which method was used or which level of correlation or sample size was selected, there were no marked differences between the methods.</p> <p>Based on these results, researchers are encouraged first to consider the problem they are trying to solve. Data nature and feature understanding can lead to more accurate and efficient methods implementation while making it easier to pivot to new analytic problems, adapt when model accuracy drifts, and save data scientists and business users considerable time and effort.</p>","abstract_html":"&lt;p&gt;Logistic Regression (LR), LASSO regression, and RIDGE regression are standard classification techniques for predicting a dichotomous output. Since these methods are applied for similar purposes and have different features, it is crucial to evaluate the performance of these methods under different controlled conditions. With this information, researchers can apply the optimal method for specific conditions.&lt;/p&gt; &lt;p&gt;Following previous research, which reported the effects of conditions such as sample size and multicollinearity on the performance of the classification methods, this research focused on the effects of when sample size, level of predictor collinearity, and predictor variable weight are controlled on the performance of LR, LASSO, and RIDGE regressions. Data were simulated with 100 iterations that generated a total of n = 2,400 observations in R statistical software. A factorial ANOVA with follow-ups was employed to evaluate the effect of conditions on the performance of each technique as measured by accuracy and F-measure.&lt;/p&gt; &lt;p&gt;In most conditions for the two outcome performance measures (accuracy and F-measure), the highest effect on performances was observed from the predictor variable weight. However, when the weight was low, all three regression methods were found to have an overall better performance under high correlation and a large sample size. Moreover, the models with high-weight conditions suppressed the effects of every other controlled condition on accuracy and F-measure output values. Therefore, when the study data conditions include a high-weighted variable, regardless of which method was used or which level of correlation or sample size was selected, there were no marked differences between the methods.&lt;/p&gt; &lt;p&gt;Based on these results, researchers are encouraged first to consider the problem they are trying to solve. Data nature and feature understanding can lead to more accurate and efficient methods implementation while making it easier to pivot to new analytic problems, adapt when model accuracy drifts, and save data scientists and business users considerable time and effort.&lt;/p&gt;","abstract_has_math":false,"creators":["AlJuhani, Mahmoud M."],"institution":null,"degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Nicholas Cutforth","Frederique Chevillot","Kathy Green","Antonio Olmos"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-12-01T08:00:00Z","date_published":"2022-12-01T08:00:00Z","updated_at":"2026-07-24T02:03:12Z","subjects":["Collinearity","LASSO","Logistic regression","RIDGE","Sample size","Weight","Multivariate Analysis","Physical Sciences and Mathematics","Statistical Methodology","Statistics and Probability"],"languages":["en"],"rights":["<p>Copyright is held by the author. User is responsible for all copyright compliance.</p>"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.du.edu/etd/2169","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Nicholas Cutforth","Frederique Chevillot","Kathy Green","Antonio Olmos"]},{"key":"dc:creator","label":"Author","values":["AlJuhani, Mahmoud M."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2025-04-11T07:00:00Z"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Collinearity","LASSO","Logistic regression","RIDGE","Sample size","Weight","Multivariate Analysis","Physical Sciences and Mathematics","Statistical Methodology","Statistics and Probability"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["<p>Copyright is held by the author. User is responsible for all copyright compliance.</p>"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.du.edu/etd/2169"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Logistic Regression (LR), LASSO regression, and RIDGE regression are standard classification techniques for predicting a dichotomous output. Since these methods are applied for similar purposes and have different features, it is crucial to evaluate the performance of these methods under different controlled conditions. With this information, researchers can apply the optimal method for specific conditions.</p> <p>Following previous research, which reported the effects of conditions such as sample size and multicollinearity on the performance of the classification methods, this research focused on the effects of when sample size, level of predictor collinearity, and predictor variable weight are controlled on the performance of LR, LASSO, and RIDGE regressions. Data were simulated with 100 iterations that generated a total of n = 2,400 observations in R statistical software. A factorial ANOVA with follow-ups was employed to evaluate the effect of conditions on the performance of each technique as measured by accuracy and F-measure.</p> <p>In most conditions for the two outcome performance measures (accuracy and F-measure), the highest effect on performances was observed from the predictor variable weight. However, when the weight was low, all three regression methods were found to have an overall better performance under high correlation and a large sample size. Moreover, the models with high-weight conditions suppressed the effects of every other controlled condition on accuracy and F-measure output values. Therefore, when the study data conditions include a high-weighted variable, regardless of which method was used or which level of correlation or sample size was selected, there were no marked differences between the methods.</p> <p>Based on these results, researchers are encouraged first to consider the problem they are trying to solve. Data nature and feature understanding can lead to more accurate and efficient methods implementation while making it easier to pivot to new analytic problems, adapt when model accuracy drifts, and save data scientists and business users considerable time and effort.</p>"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["A Comparison of Logistic, RIDGE, and LASSO Regression with Heart Failure Risk Data: Effects of Sample Size, Predictor Correlation, and Predictor Weight on Outcome Accuracy"]}]}],"canonical_facts":{"dc:contributor":["Nicholas Cutforth","Frederique Chevillot","Kathy Green","Antonio Olmos"],"dc:creator":["AlJuhani, Mahmoud M."],"dc:date.available":["2025-04-11T07:00:00Z"],"dc:description.abstract":["<p>Logistic Regression (LR), LASSO regression, and RIDGE regression are standard classification techniques for predicting a dichotomous output. Since these methods are applied for similar purposes and have different features, it is crucial to evaluate the performance of these methods under different controlled conditions. With this information, researchers can apply the optimal method for specific conditions.</p> <p>Following previous research, which reported the effects of conditions such as sample size and multicollinearity on the performance of the classification methods, this research focused on the effects of when sample size, level of predictor collinearity, and predictor variable weight are controlled on the performance of LR, LASSO, and RIDGE regressions. Data were simulated with 100 iterations that generated a total of n = 2,400 observations in R statistical software. A factorial ANOVA with follow-ups was employed to evaluate the effect of conditions on the performance of each technique as measured by accuracy and F-measure.</p> <p>In most conditions for the two outcome performance measures (accuracy and F-measure), the highest effect on performances was observed from the predictor variable weight. However, when the weight was low, all three regression methods were found to have an overall better performance under high correlation and a large sample size. Moreover, the models with high-weight conditions suppressed the effects of every other controlled condition on accuracy and F-measure output values. Therefore, when the study data conditions include a high-weighted variable, regardless of which method was used or which level of correlation or sample size was selected, there were no marked differences between the methods.</p> <p>Based on these results, researchers are encouraged first to consider the problem they are trying to solve. Data nature and feature understanding can lead to more accurate and efficient methods implementation while making it easier to pivot to new analytic problems, adapt when model accuracy drifts, and save data scientists and business users considerable time and effort.</p>"],"dc:format":["application/pdf"],"dc:identifier":["https://digitalcommons.du.edu/etd/2169"],"dc:language":["en"],"dc:rights":["<p>Copyright is held by the author. User is responsible for all copyright compliance.</p>"],"dc:subject":["Collinearity","LASSO","Logistic regression","RIDGE","Sample size","Weight","Multivariate Analysis","Physical Sciences and Mathematics","Statistical Methodology","Statistics and Probability"],"dc:title":["A Comparison of Logistic, RIDGE, and LASSO Regression with Heart Failure Risk Data: Effects of Sample Size, Predictor Correlation, and Predictor Weight on Outcome Accuracy"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."]},"updated_at":"2026-07-24T02:03:12Z"}