{"id":{"repo_id":"chapman","oai_identifier":"oai:digitalcommons.chapman.edu:cads_theses-1001"},"canonical_url":"https://search.dev.ndltd.org/etd/chapman/oai:digitalcommons.chapman.edu:cads_theses-1001","repository":{"repo_id":"chapman","name":"Chapman University","base_url":"https://digitalcommons.chapman.edu/do/oai/"},"display":{"title":"A Machine Learning Approach to Predicting Alcohol Consumption in Adolescents From Historical Text Messaging Data","abstract":"<p>Techniques based on artificial neural networks represent the current state-of-the-art in machine learning due to the availability of improved hardware and large data sets. Here we employ doc2vec, an unsupervised neural network, to capture the semantic content of text messages sent by adolescents during high school, and encode this semantic content as numeric vectors. These vectors effectively condense the text message data into highly leverageable inputs to a logistic regression classifier in a matter of hours, as compared to the tedious and often quite lengthy task of manually coding data. Using our machine learning approach, we are able to train a logistic regression model to predict adolescents' engagement in substance abuse during distinct life phases with accuracy ranging from 76.5% to 88.1%. We show the effects of grade level and text message aggregation strategy on the efficacy of document embedding generation with doc2vec. Additional examination of the vectorizations for specific terms extracted from the text message data adds quantitative depth to this analysis. We demonstrate the ability of the method used herein to overcome traditional natural language processing concerns related to unconventional orthography. These results suggest that the approach described in this thesis is a competitive and efficient alternative to existing methodologies for predicting substance abuse behaviors. This work reveals the potential for the application of machine learning-based manipulation of text messaging data to development of automatic intervention strategies against substance abuse and other adolescent challenges.</p>","abstract_html":"&lt;p&gt;Techniques based on artificial neural networks represent the current state-of-the-art in machine learning due to the availability of improved hardware and large data sets. Here we employ doc2vec, an unsupervised neural network, to capture the semantic content of text messages sent by adolescents during high school, and encode this semantic content as numeric vectors. These vectors effectively condense the text message data into highly leverageable inputs to a logistic regression classifier in a matter of hours, as compared to the tedious and often quite lengthy task of manually coding data. Using our machine learning approach, we are able to train a logistic regression model to predict adolescents&#x27; engagement in substance abuse during distinct life phases with accuracy ranging from 76.5% to 88.1%. We show the effects of grade level and text message aggregation strategy on the efficacy of document embedding generation with doc2vec. Additional examination of the vectorizations for specific terms extracted from the text message data adds quantitative depth to this analysis. We demonstrate the ability of the method used herein to overcome traditional natural language processing concerns related to unconventional orthography. These results suggest that the approach described in this thesis is a competitive and efficient alternative to existing methodologies for predicting substance abuse behaviors. This work reveals the potential for the application of machine learning-based manipulation of text messaging data to development of automatic intervention strategies against substance abuse and other adolescent challenges.&lt;/p&gt;","abstract_has_math":false,"creators":["Bergh, Adrienne"],"institution":null,"degree_name":"Master of Science (MS)","degree_level":"Thesis","degree_discipline":"Computational and Data Sciences","degree_department":null,"school":null,"contributors":["Erik Linstead, Ph.D.","Elizabeth Stevens, Ph.D.","Samuel E. Ehrenreich, Ph.D."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-05-28T07:00:00Z","date_published":"2019-05-28T07:00:00Z","updated_at":"2026-07-24T01:38:00Z","subjects":["Machine Learning","Natural Language Processing","Neural Embeddings","Adolescence","doc2vec","Computational Engineering","Developmental Psychology"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.chapman.edu/cads_theses/2","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Erik Linstead, Ph.D.","Elizabeth Stevens, Ph.D.","Samuel E. Ehrenreich, Ph.D."]},{"key":"dc:creator","label":"Author","values":["Bergh, Adrienne"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["Computational and Data Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","Natural Language Processing","Neural Embeddings","Adolescence","doc2vec","Computational Engineering","Developmental Psychology"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.chapman.edu/cads_theses/2"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Techniques based on artificial neural networks represent the current state-of-the-art in machine learning due to the availability of improved hardware and large data sets. Here we employ doc2vec, an unsupervised neural network, to capture the semantic content of text messages sent by adolescents during high school, and encode this semantic content as numeric vectors. These vectors effectively condense the text message data into highly leverageable inputs to a logistic regression classifier in a matter of hours, as compared to the tedious and often quite lengthy task of manually coding data. Using our machine learning approach, we are able to train a logistic regression model to predict adolescents' engagement in substance abuse during distinct life phases with accuracy ranging from 76.5% to 88.1%. We show the effects of grade level and text message aggregation strategy on the efficacy of document embedding generation with doc2vec. Additional examination of the vectorizations for specific terms extracted from the text message data adds quantitative depth to this analysis. We demonstrate the ability of the method used herein to overcome traditional natural language processing concerns related to unconventional orthography. These results suggest that the approach described in this thesis is a competitive and efficient alternative to existing methodologies for predicting substance abuse behaviors. This work reveals the potential for the application of machine learning-based manipulation of text messaging data to development of automatic intervention strategies against substance abuse and other adolescent challenges.</p>"]},{"key":"dc:source","label":"Dc Source","values":["A. Bergh, \"A machine learning approach to predicting alcohol consumption in adolescents from historical text messaging data,\" M. S. thesis, Chapman University, Orange, CA, 2019. <a href=\"https://doi.org/10.36837/chapman.000072\">https://doi.org/10.36837/chapman.000072</a>"]},{"key":"dc:title","label":"Title","values":["A Machine Learning Approach to Predicting Alcohol Consumption in Adolescents From Historical Text Messaging Data"]}]}],"canonical_facts":{"dc:contributor":["Erik Linstead, Ph.D.","Elizabeth Stevens, Ph.D.","Samuel E. Ehrenreich, Ph.D."],"dc:creator":["Bergh, Adrienne"],"dc:description.abstract":["<p>Techniques based on artificial neural networks represent the current state-of-the-art in machine learning due to the availability of improved hardware and large data sets. Here we employ doc2vec, an unsupervised neural network, to capture the semantic content of text messages sent by adolescents during high school, and encode this semantic content as numeric vectors. These vectors effectively condense the text message data into highly leverageable inputs to a logistic regression classifier in a matter of hours, as compared to the tedious and often quite lengthy task of manually coding data. Using our machine learning approach, we are able to train a logistic regression model to predict adolescents' engagement in substance abuse during distinct life phases with accuracy ranging from 76.5% to 88.1%. We show the effects of grade level and text message aggregation strategy on the efficacy of document embedding generation with doc2vec. Additional examination of the vectorizations for specific terms extracted from the text message data adds quantitative depth to this analysis. We demonstrate the ability of the method used herein to overcome traditional natural language processing concerns related to unconventional orthography. These results suggest that the approach described in this thesis is a competitive and efficient alternative to existing methodologies for predicting substance abuse behaviors. This work reveals the potential for the application of machine learning-based manipulation of text messaging data to development of automatic intervention strategies against substance abuse and other adolescent challenges.</p>"],"dc:identifier":["https://digitalcommons.chapman.edu/cads_theses/2"],"dc:source":["A. Bergh, \"A machine learning approach to predicting alcohol consumption in adolescents from historical text messaging data,\" M. S. thesis, Chapman University, Orange, CA, 2019. <a href=\"https://doi.org/10.36837/chapman.000072\">https://doi.org/10.36837/chapman.000072</a>"],"dc:subject":["Machine Learning","Natural Language Processing","Neural Embeddings","Adolescence","doc2vec","Computational Engineering","Developmental Psychology"],"dc:title":["A Machine Learning Approach to Predicting Alcohol Consumption in Adolescents From Historical Text Messaging Data"],"thesis:degree_discipline":["Computational and Data Sciences"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Master of Science (MS)"]},"updated_at":"2026-07-24T01:38:00Z"}