{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/143320"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/143320","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Large-Scale Online Conversations About Public Health: Predicting Real-World Outcomes","abstract":"Drug overdose remains one of the most severe public health challenges in the United States. Although online communities contain vast amounts of firsthand accounts, personal experiences, and peer-to-peer discussions about drug use, it is still unclear how these conversations can be analyzed to generate insights that support public health research. This dissertation addresses three core research questions: (1) How can we leverage Large Language Models to extract \"gists\" (causal language patterns) from decade-long online discussions? (2) What kinds of gists characterize how and why people discuss drugs, and how do these gists evolve over time? (3) Do these discussion themes align with, or predict, changes in real-world health outcomes (specifically overdose mortality)? To address these questions, Study 1 develops and validates an instruction-tuned large language model pipeline for extracting causal gists. Study 2 constructs a thematic taxonomy for these gists and uses multiple NLP models to classify and analyze how major discussion themes evolve over the ten years. Study 3 links these online themes to real-world health outcomes by applying time-series models, including autoregressive distributed lag (ARDL) analyses, to test whether changes in topic prevalence correspond with or precede trends in national, state-level, and drug-specific overdose mortality. Together, these studies demonstrate that large-scale online conversations contain structured, meaningful signals that reflect and anticipate real-world patterns in the overdose crisis.","abstract_html":"Drug overdose remains one of the most severe public health challenges in the United States. Although online communities contain vast amounts of firsthand accounts, personal experiences, and peer-to-peer discussions about drug use, it is still unclear how these conversations can be analyzed to generate insights that support public health research. This dissertation addresses three core research questions: (1) How can we leverage Large Language Models to extract &quot;gists&quot; (causal language patterns) from decade-long online discussions? (2) What kinds of gists characterize how and why people discuss drugs, and how do these gists evolve over time? (3) Do these discussion themes align with, or predict, changes in real-world health outcomes (specifically overdose mortality)? To address these questions, Study 1 develops and validates an instruction-tuned large language model pipeline for extracting causal gists. Study 2 constructs a thematic taxonomy for these gists and uses multiple NLP models to classify and analyze how major discussion themes evolve over the ten years. Study 3 links these online themes to real-world health outcomes by applying time-series models, including autoregressive distributed lag (ARDL) analyses, to test whether changes in topic prevalence correspond with or precede trends in national, state-level, and drug-specific overdose mortality. Together, these studies demonstrate that large-scale online conversations contain structured, meaningful signals that reflect and anticipate real-world patterns in the overdose crisis.","abstract_has_math":false,"creators":["Ding, Xiaohan"],"institution":"Virginia Tech","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"Computer Science & Applications","degree_department":"Computer Science and Applications","school":null,"contributors":[],"advisors":[],"committee_chairs":["Rho, Ha Rim"],"committee_members":["Ramakrishnan, Narendran","Lee, Sang Won","Huang, Lifu","North, Christopher L."],"year":2026,"date_issued":"2026-06-08","date_published":"2026-06-08","updated_at":"2026-07-24T05:56:51Z","subjects":["Social Media Discourse","Causal Language Extraction","Large Language Models","Public Health Surveillance"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:46825"],"render_values":[{"text":"vt_gsexam:46825","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/143320","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Rho, Ha Rim"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Ramakrishnan, Narendran","Lee, Sang Won","Huang, Lifu","North, Christopher L."]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science and Applications"]},{"key":"dc:creator","label":"Author","values":["Ding, Xiaohan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-06-09T08:07:03Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-06-09T08:07:03Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-06-08"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science & Applications"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Social Media Discourse","Causal Language Extraction","Large Language Models","Public Health Surveillance"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:46825"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/143320"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Drug overdose remains one of the most severe public health challenges in the United States. Although online communities contain vast amounts of firsthand accounts, personal experiences, and peer-to-peer discussions about drug use, it is still unclear how these conversations can be analyzed to generate insights that support public health research. This dissertation addresses three core research questions: (1) How can we leverage Large Language Models to extract \"gists\" (causal language patterns) from decade-long online discussions? (2) What kinds of gists characterize how and why people discuss drugs, and how do these gists evolve over time? (3) Do these discussion themes align with, or predict, changes in real-world health outcomes (specifically overdose mortality)? To address these questions, Study 1 develops and validates an instruction-tuned large language model pipeline for extracting causal gists. Study 2 constructs a thematic taxonomy for these gists and uses multiple NLP models to classify and analyze how major discussion themes evolve over the ten years. Study 3 links these online themes to real-world health outcomes by applying time-series models, including autoregressive distributed lag (ARDL) analyses, to test whether changes in topic prevalence correspond with or precede trends in national, state-level, and drug-specific overdose mortality. Together, these studies demonstrate that large-scale online conversations contain structured, meaningful signals that reflect and anticipate real-world patterns in the overdose crisis."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Every day in the United States, about 300 people die from drug overdose, making it one of the most serious health problems in the country. Government agencies collect death records to track this crisis, but these records often take six to twelve months to become available. Meanwhile, millions of people share personal stories about drug use on online platforms such as Reddit. This dissertation develops a method that uses advanced computer programs called large language models to read millions of online posts and extract the key \"cause and effect\" ideas that people express. We call these simplified ideas \"gists,\" a term from psychology that refers to the core meaning people rely on when making decisions. Using this method, we analyzed nearly 700,000 Reddit posts written over ten years (2015 to 2024) and identified seven major discussion topics, including reasons for starting drug use, health effects, treatment and recovery, and access to health services. We then compared how often these topics appeared each month with official government records of overdose deaths. Our results show that when discussions about drug use methods and health problems increased, overdose death rates grew faster, while when more people talked about treatment and recovery, death rates grew more slowly. These patterns held at the national level, across individual states, and for specific drugs such as fentanyl and heroin. This research suggests that online conversations can serve as an early signal for public health agencies, helping officials detect warning signs sooner and respond more quickly."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Large-Scale Online Conversations About Public Health: Predicting Real-World Outcomes"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Rho, Ha Rim"],"dc:contributor.committeemember":["Ramakrishnan, Narendran","Lee, Sang Won","Huang, Lifu","North, Christopher L."],"dc:contributor.department":["Computer Science and Applications"],"dc:creator":["Ding, Xiaohan"],"dc:date.accessioned":["2026-06-09T08:07:03Z"],"dc:date.available":["2026-06-09T08:07:03Z"],"dc:date.issued":["2026-06-08"],"dc:description.abstract":["Drug overdose remains one of the most severe public health challenges in the United States. Although online communities contain vast amounts of firsthand accounts, personal experiences, and peer-to-peer discussions about drug use, it is still unclear how these conversations can be analyzed to generate insights that support public health research. This dissertation addresses three core research questions: (1) How can we leverage Large Language Models to extract \"gists\" (causal language patterns) from decade-long online discussions? (2) What kinds of gists characterize how and why people discuss drugs, and how do these gists evolve over time? (3) Do these discussion themes align with, or predict, changes in real-world health outcomes (specifically overdose mortality)? To address these questions, Study 1 develops and validates an instruction-tuned large language model pipeline for extracting causal gists. Study 2 constructs a thematic taxonomy for these gists and uses multiple NLP models to classify and analyze how major discussion themes evolve over the ten years. Study 3 links these online themes to real-world health outcomes by applying time-series models, including autoregressive distributed lag (ARDL) analyses, to test whether changes in topic prevalence correspond with or precede trends in national, state-level, and drug-specific overdose mortality. Together, these studies demonstrate that large-scale online conversations contain structured, meaningful signals that reflect and anticipate real-world patterns in the overdose crisis."],"dc:description.abstractgeneral":["Every day in the United States, about 300 people die from drug overdose, making it one of the most serious health problems in the country. Government agencies collect death records to track this crisis, but these records often take six to twelve months to become available. Meanwhile, millions of people share personal stories about drug use on online platforms such as Reddit. This dissertation develops a method that uses advanced computer programs called large language models to read millions of online posts and extract the key \"cause and effect\" ideas that people express. We call these simplified ideas \"gists,\" a term from psychology that refers to the core meaning people rely on when making decisions. Using this method, we analyzed nearly 700,000 Reddit posts written over ten years (2015 to 2024) and identified seven major discussion topics, including reasons for starting drug use, health effects, treatment and recovery, and access to health services. We then compared how often these topics appeared each month with official government records of overdose deaths. Our results show that when discussions about drug use methods and health problems increased, overdose death rates grew faster, while when more people talked about treatment and recovery, death rates grew more slowly. These patterns held at the national level, across individual states, and for specific drugs such as fentanyl and heroin. This research suggests that online conversations can serve as an early signal for public health agencies, helping officials detect warning signs sooner and respond more quickly."],"dc:description.degree":["Doctor of Philosophy"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:46825"],"dc:identifier.uri":["https://hdl.handle.net/10919/143320"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Social Media Discourse","Causal Language Extraction","Large Language Models","Public Health Surveillance"],"dc:title":["Large-Scale Online Conversations About Public Health: Predicting Real-World Outcomes"],"dc:type":["Dissertation"],"thesis:degree_discipline":["Computer Science & Applications"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-24T05:56:51Z"}