Abstract
Written language is one of the most information rich and abundant data sources available. Rigorous statisti- cal study of linguistic phenomena can elevate the digital humanities, linguistic applications to computational social science, and statistics itself. As yet, statistical applications to language, whether in linguistics, compu- tation, or otherwise, are largely ad hoc. Performance gains in modeling have largely been due to two factors: fine tuning transfer learning and increasing the size and complexity of models. Due to the statistical nature of human language—governed by several power law phenomena—fine tuning transfer learning may be ad- vantageous for corpus analyses, not just artificial intelligence applications. Yet state of the art “transformer” models are expensive and opaque. I propose we revisit Latent Dirichlet Allocation (LDA). As a parametric statistical model of a data generating process, it has the potential to be used in statistically rigorous ways to study written language. In this dissertation, I state my case for what I refer to as “natural language statistics”, a more statistically rigorous take on analyzing text data. I link the data generating process for LDA to Zipf’s law and use that relationship to engage in simulation studies for LDA. I also develop a coefficient of deter- mination for topic models, extend LDA to enable pre-train/fine tuning transfer learning, and implement this research and more in an R package, tidylda.
Author and committee
dc:creator, dc:contributor.*- Author
-
- Jones, Thomas
Subjects
dc:subject × 6Identifiers
dc:identifier.*- Identifier
- hdl:1920/13945
- OAI identifier oai:identifier
- oai:MARS:1920/13945