{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/120088"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/120088","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Time series modeling of text data","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-09-01 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2023-09-01 without embargo terms","abstract_has_math":false,"creators":["Dey, Priyanka"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, ChengXiang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-05","date_published":"2023-05","updated_at":"2026-07-22T22:24:56Z","subjects":["Time Series Modeling","Text Data","Word Representations","Interpretability"],"languages":["en","eng"],"rights":["Copyright 2023 Priyanka Dey"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/120088","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, ChengXiang"]},{"key":"dc:creator","label":"Author","values":["Dey, Priyanka"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-05","2023-04-24"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Time Series Modeling","Text Data","Word Representations","Interpretability"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Priyanka Dey"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/120088"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-09-01 without embargo terms","The student, Priyanka Dey, accepted the attached license on 2023-04-24 at 09:12.","The student, Priyanka Dey, submitted this Thesis for approval on 2023-04-24 at 09:19.","This Thesis was approved for publication on 2023-04-24 at 12:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19128 on 2023-09-01 at 16:55:10","The success of machine learning algorithms in recent years has further accelerated the race to develop language models that can accurately and precisely represent text for a wide range of downstream tasks. Unfortunately, the development of these large and powerful models has also led to an increase in model architectures and complexities, thus sometimes making it extremely difficult to understand and interpret these model results. In this research, we present a new methodology for representing text using time series models, namely ARIMA to represent word embeddings as a set of regression equations. Through our experiments and analysis, we show that these representations are often successful in learning language across various domains e.g. sports, politics, and science. We further show that our representations can successfully be used to foster development of downstream applications such as next word prediction, salient dimension lattice generation, and article title generation. With the surge of the field of natural language processing, building models that can accurately represent and generate language has been a major focus for research. Models such as BERT, GPT-3, Chat-GPT are all examples of these large language models that can be used for a wide range of applications with a remarkable amount of precision and accuracy. However, many critique that such models are essentially black boxes. Our work is motivated by the challenging nature of these models to develop a simple yet still effective means of representing text through time series models."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Time series modeling of text data"]}]}],"canonical_facts":{"dc:contributor":["Zhai, ChengXiang"],"dc:creator":["Dey, Priyanka"],"dc:date":["2023-05","2023-04-24"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-09-01 without embargo terms","The student, Priyanka Dey, accepted the attached license on 2023-04-24 at 09:12.","The student, Priyanka Dey, submitted this Thesis for approval on 2023-04-24 at 09:19.","This Thesis was approved for publication on 2023-04-24 at 12:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19128 on 2023-09-01 at 16:55:10","The success of machine learning algorithms in recent years has further accelerated the race to develop language models that can accurately and precisely represent text for a wide range of downstream tasks. Unfortunately, the development of these large and powerful models has also led to an increase in model architectures and complexities, thus sometimes making it extremely difficult to understand and interpret these model results. In this research, we present a new methodology for representing text using time series models, namely ARIMA to represent word embeddings as a set of regression equations. Through our experiments and analysis, we show that these representations are often successful in learning language across various domains e.g. sports, politics, and science. We further show that our representations can successfully be used to foster development of downstream applications such as next word prediction, salient dimension lattice generation, and article title generation. With the surge of the field of natural language processing, building models that can accurately represent and generate language has been a major focus for research. Models such as BERT, GPT-3, Chat-GPT are all examples of these large language models that can be used for a wide range of applications with a remarkable amount of precision and accuracy. However, many critique that such models are essentially black boxes. Our work is motivated by the challenging nature of these models to develop a simple yet still effective means of representing text through time series models."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/120088"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Priyanka Dey"],"dc:subject":["Time Series Modeling","Text Data","Word Representations","Interpretability"],"dc:title":["Time series modeling of text data"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:56Z"}