{"id":{"repo_id":"claremont","oai_identifier":"oai:scholarship.claremont.edu:cgu_etd-1913"},"canonical_url":"https://search.dev.ndltd.org/etd/claremont/oai:scholarship.claremont.edu:cgu_etd-1913","repository":{"repo_id":"claremont","name":"Claremont Graduate University","base_url":"https://scholarship.claremont.edu/do/oai/"},"display":{"title":"What Text Information Helps to Reduce Default Risk","abstract":"<p>This dissertation investigates the predictive factors influencing loan default in the context of peer-to-peer (P2P) lending, with a particular focus on the integration of voluntarily provided text data alongside traditional financial, demographic, and loan information. Using a dataset of over 296,000 borrowers from the Lending Club platform, this research employs logistic regression and ensemble machine learning algorithms, including forward and backward stepwise selection and random forests, to rank the importance of various factors in predicting loan default.</p> <p>By including text information, this paper improves the accuracy of predicting the relationship between a borrower's ability to repay a loan and the decision to grant a loan by 5%, compared to 60% accuracy without the inclusion of text information.</p> <p>The analysis reveals that traditional financial variables such as interest rate, loan term, and debt-to-income ratio are the most significant predictors of default risk. However, text variables, especially those reflecting sentiment and psychological states—such as discrepancy, positive emotion, and affective processes—also play a critical role. Borrowers who express optimism or reference moral or emotional factors tend to have lower default rates, while those exhibiting financial discrepancies or negative emotions are more likely to default.</p> <p>This research contributes to the literature by integrating natural language processing (NLP) techniques, specifically the Linguistic Inquiry and Word Count (LIWC2015) tool, to quantify and analyze borrowers’ textual descriptions. The findings suggest that lenders can improve their risk assessment models by combining financial and non-financial data, particularly voluntary text information. The study also highlights the growing potential of machine learning and NLP in enhancing predictive models for credit default. Practical implications include more informed lending decisions and better resource allocation to minimize default risk.</p>","abstract_html":"&lt;p&gt;This dissertation investigates the predictive factors influencing loan default in the context of peer-to-peer (P2P) lending, with a particular focus on the integration of voluntarily provided text data alongside traditional financial, demographic, and loan information. Using a dataset of over 296,000 borrowers from the Lending Club platform, this research employs logistic regression and ensemble machine learning algorithms, including forward and backward stepwise selection and random forests, to rank the importance of various factors in predicting loan default.&lt;/p&gt; &lt;p&gt;By including text information, this paper improves the accuracy of predicting the relationship between a borrower&#x27;s ability to repay a loan and the decision to grant a loan by 5%, compared to 60% accuracy without the inclusion of text information.&lt;/p&gt; &lt;p&gt;The analysis reveals that traditional financial variables such as interest rate, loan term, and debt-to-income ratio are the most significant predictors of default risk. However, text variables, especially those reflecting sentiment and psychological states—such as discrepancy, positive emotion, and affective processes—also play a critical role. Borrowers who express optimism or reference moral or emotional factors tend to have lower default rates, while those exhibiting financial discrepancies or negative emotions are more likely to default.&lt;/p&gt; &lt;p&gt;This research contributes to the literature by integrating natural language processing (NLP) techniques, specifically the Linguistic Inquiry and Word Count (LIWC2015) tool, to quantify and analyze borrowers’ textual descriptions. The findings suggest that lenders can improve their risk assessment models by combining financial and non-financial data, particularly voluntary text information. The study also highlights the growing potential of machine learning and NLP in enhancing predictive models for credit default. Practical implications include more informed lending decisions and better resource allocation to minimize default risk.&lt;/p&gt;","abstract_has_math":false,"creators":["Wang, Guan"],"institution":null,"degree_name":"Economics, PhD","degree_level":"Restricted to Claremont Colleges Dissertation","degree_discipline":"School of Social Science, Politics, and Evaluation","degree_department":null,"school":null,"contributors":["Graham Bird","Levan Efremidze"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-01-01T08:00:00Z","date_published":"2024-01-01T08:00:00Z","updated_at":"2026-07-24T01:41:01Z","subjects":["Economics"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://scholarship.claremont.edu/cgu_etd/891","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Graham Bird","Levan Efremidze"]},{"key":"dc:creator","label":"Author","values":["Wang, Guan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2025-01-07T08:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["School of Social Science, Politics, and Evaluation"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Restricted to Claremont Colleges Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Economics, PhD"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Economics"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://scholarship.claremont.edu/cgu_etd/891"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>This dissertation investigates the predictive factors influencing loan default in the context of peer-to-peer (P2P) lending, with a particular focus on the integration of voluntarily provided text data alongside traditional financial, demographic, and loan information. Using a dataset of over 296,000 borrowers from the Lending Club platform, this research employs logistic regression and ensemble machine learning algorithms, including forward and backward stepwise selection and random forests, to rank the importance of various factors in predicting loan default.</p> <p>By including text information, this paper improves the accuracy of predicting the relationship between a borrower's ability to repay a loan and the decision to grant a loan by 5%, compared to 60% accuracy without the inclusion of text information.</p> <p>The analysis reveals that traditional financial variables such as interest rate, loan term, and debt-to-income ratio are the most significant predictors of default risk. However, text variables, especially those reflecting sentiment and psychological states—such as discrepancy, positive emotion, and affective processes—also play a critical role. Borrowers who express optimism or reference moral or emotional factors tend to have lower default rates, while those exhibiting financial discrepancies or negative emotions are more likely to default.</p> <p>This research contributes to the literature by integrating natural language processing (NLP) techniques, specifically the Linguistic Inquiry and Word Count (LIWC2015) tool, to quantify and analyze borrowers’ textual descriptions. The findings suggest that lenders can improve their risk assessment models by combining financial and non-financial data, particularly voluntary text information. The study also highlights the growing potential of machine learning and NLP in enhancing predictive models for credit default. Practical implications include more informed lending decisions and better resource allocation to minimize default risk.</p>"]},{"key":"dc:title","label":"Title","values":["What Text Information Helps to Reduce Default Risk"]}]}],"canonical_facts":{"dc:contributor":["Graham Bird","Levan Efremidze"],"dc:creator":["Wang, Guan"],"dc:date.available":["2025-01-07T08:00:00Z"],"dc:description.abstract":["<p>This dissertation investigates the predictive factors influencing loan default in the context of peer-to-peer (P2P) lending, with a particular focus on the integration of voluntarily provided text data alongside traditional financial, demographic, and loan information. Using a dataset of over 296,000 borrowers from the Lending Club platform, this research employs logistic regression and ensemble machine learning algorithms, including forward and backward stepwise selection and random forests, to rank the importance of various factors in predicting loan default.</p> <p>By including text information, this paper improves the accuracy of predicting the relationship between a borrower's ability to repay a loan and the decision to grant a loan by 5%, compared to 60% accuracy without the inclusion of text information.</p> <p>The analysis reveals that traditional financial variables such as interest rate, loan term, and debt-to-income ratio are the most significant predictors of default risk. However, text variables, especially those reflecting sentiment and psychological states—such as discrepancy, positive emotion, and affective processes—also play a critical role. Borrowers who express optimism or reference moral or emotional factors tend to have lower default rates, while those exhibiting financial discrepancies or negative emotions are more likely to default.</p> <p>This research contributes to the literature by integrating natural language processing (NLP) techniques, specifically the Linguistic Inquiry and Word Count (LIWC2015) tool, to quantify and analyze borrowers’ textual descriptions. The findings suggest that lenders can improve their risk assessment models by combining financial and non-financial data, particularly voluntary text information. The study also highlights the growing potential of machine learning and NLP in enhancing predictive models for credit default. Practical implications include more informed lending decisions and better resource allocation to minimize default risk.</p>"],"dc:identifier":["https://scholarship.claremont.edu/cgu_etd/891"],"dc:subject":["Economics"],"dc:title":["What Text Information Helps to Reduce Default Risk"],"thesis:degree_discipline":["School of Social Science, Politics, and Evaluation"],"thesis:degree_level":["Restricted to Claremont Colleges Dissertation"],"thesis:degree_name":["Economics, PhD"]},"updated_at":"2026-07-24T01:41:01Z"}