{"id":{"repo_id":"washington","oai_identifier":"oai:digital.lib.washington.edu:1773/49313"},"canonical_url":"https://search.dev.ndltd.org/etd/washington/oai:digital.lib.washington.edu:1773/49313","repository":{"repo_id":"washington","name":"University of Washington","base_url":"https://digital.lib.washington.edu/server/oai/request"},"display":{"title":"Investigating Natural Language Interactions in Communities","abstract":"In this work, we investigate language use in communities. First, we study how authors of scientific textsexplain how their paper relates to another. We propose a new task, relationship explanation generation for scientific texts, by using in-line citation text as a source of evidence. We introduce a new model, SCIGEN, empowered by pretrained language models. We perform extensive human evaluation and discuss potential shortcomings of our system’s generations. Second, we describe work on investigating how NLP models degrade over time due to the dynamic nature of most communities. We find evidence that static models, like SCIGEN, do not generalize temporally. Based on this finding, we investigate how model performance deterioration over time differs across several tasks and domains. We discover that models with temporally misaligned train and testing sets can suffer from large amounts of performance degradation. Finally, we show how elevating temporal characteristics in communities allows us to study particular social phenomena. Namely, we explore the problem of quantifying persuasive skill over time. Using data from an online debate forum, we construct a model of debater skill based on on the Elo ranking model and incorporate historical lingusitic data. In order to estimate skill, we frame our prediction task to forecast the outcome of a debate. Though this work considers a wide range of NLP applications, it is unified by the idea that language emerges from ever-changing communities of people engaging with each other. Our findings demonstrate how the temporal dynamics and social underpinnings of language can inform NLP research and practice.","abstract_html":"In this work, we investigate language use in communities. First, we study how authors of scientific textsexplain how their paper relates to another. We propose a new task, relationship explanation generation for scientific texts, by using in-line citation text as a source of evidence. We introduce a new model, SCIGEN, empowered by pretrained language models. We perform extensive human evaluation and discuss potential shortcomings of our system’s generations. Second, we describe work on investigating how NLP models degrade over time due to the dynamic nature of most communities. We find evidence that static models, like SCIGEN, do not generalize temporally. Based on this finding, we investigate how model performance deterioration over time differs across several tasks and domains. We discover that models with temporally misaligned train and testing sets can suffer from large amounts of performance degradation. Finally, we show how elevating temporal characteristics in communities allows us to study particular social phenomena. Namely, we explore the problem of quantifying persuasive skill over time. Using data from an online debate forum, we construct a model of debater skill based on on the Elo ranking model and incorporate historical lingusitic data. In order to estimate skill, we frame our prediction task to forecast the outcome of a debate. Though this work considers a wide range of NLP applications, it is unified by the idea that language emerges from ever-changing communities of people engaging with each other. Our findings demonstrate how the temporal dynamics and social underpinnings of language can inform NLP research and practice.","abstract_has_math":false,"creators":["Luu, Kelvin"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Smith, Noah A."],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-09-23","date_published":"2022-09-23","updated_at":"2026-07-24T05:58:10Z","subjects":["Computer science"],"languages":["en_US"],"rights":["CC BY-SA"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1773/49313","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Smith, Noah A."]},{"key":"dc:creator","label":"Author","values":["Luu, Kelvin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2022-09-23T20:44:28Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2022-09-23T20:44:28Z"]},{"key":"dc:date.issued","label":"Date","values":["2022-09-23"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]},{"key":"dc:rights","label":"Dc Rights","values":["CC BY-SA"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["Luu_washington_0250E_24843.pdf"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1773/49313"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (Ph.D.)--University of Washington, 2022"]},{"key":"dc:description.abstract","label":"Abstract","values":["In this work, we investigate language use in communities. First, we study how authors of scientific textsexplain how their paper relates to another. We propose a new task, relationship explanation generation for scientific texts, by using in-line citation text as a source of evidence. We introduce a new model, SCIGEN, empowered by pretrained language models. We perform extensive human evaluation and discuss potential shortcomings of our system’s generations. Second, we describe work on investigating how NLP models degrade over time due to the dynamic nature of most communities. We find evidence that static models, like SCIGEN, do not generalize temporally. Based on this finding, we investigate how model performance deterioration over time differs across several tasks and domains. We discover that models with temporally misaligned train and testing sets can suffer from large amounts of performance degradation. Finally, we show how elevating temporal characteristics in communities allows us to study particular social phenomena. Namely, we explore the problem of quantifying persuasive skill over time. Using data from an online debate forum, we construct a model of debater skill based on on the Elo ranking model and incorporate historical lingusitic data. In order to estimate skill, we frame our prediction task to forecast the outcome of a debate. Though this work considers a wide range of NLP applications, it is unified by the idea that language emerges from ever-changing communities of people engaging with each other. Our findings demonstrate how the temporal dynamics and social underpinnings of language can inform NLP research and practice."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Investigating Natural Language Interactions in Communities"]}]}],"canonical_facts":{"dc:contributor.advisor":["Smith, Noah A."],"dc:creator":["Luu, Kelvin"],"dc:date.accessioned":["2022-09-23T20:44:28Z"],"dc:date.available":["2022-09-23T20:44:28Z"],"dc:date.issued":["2022-09-23"],"dc:description":["Thesis (Ph.D.)--University of Washington, 2022"],"dc:description.abstract":["In this work, we investigate language use in communities. First, we study how authors of scientific textsexplain how their paper relates to another. We propose a new task, relationship explanation generation for scientific texts, by using in-line citation text as a source of evidence. We introduce a new model, SCIGEN, empowered by pretrained language models. We perform extensive human evaluation and discuss potential shortcomings of our system’s generations. Second, we describe work on investigating how NLP models degrade over time due to the dynamic nature of most communities. We find evidence that static models, like SCIGEN, do not generalize temporally. Based on this finding, we investigate how model performance deterioration over time differs across several tasks and domains. We discover that models with temporally misaligned train and testing sets can suffer from large amounts of performance degradation. Finally, we show how elevating temporal characteristics in communities allows us to study particular social phenomena. Namely, we explore the problem of quantifying persuasive skill over time. Using data from an online debate forum, we construct a model of debater skill based on on the Elo ranking model and incorporate historical lingusitic data. In order to estimate skill, we frame our prediction task to forecast the outcome of a debate. Though this work considers a wide range of NLP applications, it is unified by the idea that language emerges from ever-changing communities of people engaging with each other. Our findings demonstrate how the temporal dynamics and social underpinnings of language can inform NLP research and practice."],"dc:format.mimetype":["application/pdf"],"dc:identifier.other":["Luu_washington_0250E_24843.pdf"],"dc:identifier.uri":["http://hdl.handle.net/1773/49313"],"dc:language.iso":["en_US"],"dc:rights":["CC BY-SA"],"dc:subject":["Computer science"],"dc:title":["Investigating Natural Language Interactions in Communities"],"dc:type":["Thesis"]},"updated_at":"2026-07-24T05:58:10Z"}