Back to results

University of Illinois at Urbana-Champaign

Time series modeling of text data

Abstract

dc:description

The success of machine learning algorithms in recent years has further accelerated the race to develop language models that can accurately and precisely represent text for a wide range of downstream tasks. Unfortunately, the development of these large and powerful models has also led to an increase in model architectures and complexities, thus sometimes making it extremely difficult to understand and interpret these model results. In this research, we present a new methodology for representing text using time series models, namely ARIMA to represent word embeddings as a set of regression equations. Through our experiments and analysis, we show that these representations are often successful in learning language across various domains e.g. sports, politics, and science. We further show that our representations can successfully be used to foster development of downstream applications such as next word prediction, salient dimension lattice generation, and article title generation. With the surge of the field of natural language processing, building models that can accurately represent and generate language has been a major focus for research. Models such as BERT, GPT-3, Chat-GPT are all examples of these large language models that can be used for a wide range of applications with a remarkable amount of precision and accuracy. However, many critique that such models are essentially black boxes. Our work is motivated by the challenging nature of these models to develop a simple yet still effective means of representing text through time series models.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Dey, Priyanka
Contributors dc:contributor
  • Zhai, ChengXiang

Subjects

dc:subject × 4

Rights

dc:rights
Statement dc:rights
  • Copyright 2023 Priyanka Dey
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/120088

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Dey, Priyanka. Time series modeling of text data. Thesis thesis, University of Illinois at Urbana-Champaign, 2023. https://hdl.handle.net/2142/120088