Abstract
dc:description.abstractPersian has two distinct variants, the literary formal form that has been traditionally used in writing and the conversational informal form typically used in oral communications. With the advent of social media, the restrictions against the use of the conversational variant in writing have been relaxed in certain contexts, and speakers tend to write in the spoken informal form. The Use of conversational Persian on the Web creates a new challenge for the analysis of texts, as current grammars, academic textbooks, corpora, and computational models of this language focus mainly on the formal variant.In this dissertation, I present the phonological, morphological, and syntactic distinctions between formal and informal Persian, showing that these two variants have fundamental differences that cannot be attributed solely to pronunciation discrepancies. I then present the open-source Informal Persian Universal Dependency Treebank (iPerUDT), a new treebank created in this effort, and provide empirical evidence for the need of this treebank that focuses explicitly on informal language. Experimenting with two state-of-the-art parsers, I show that the distinctions between the two varieties raise a domain adaptation problem: parsers, primarily trained on formal data, face unknown tokens and structures when evaluated on informal data, and fail to generalize the basic patterns, thus making ill predictions in such conditions. Specifically, it is shown through examples that the majority of erroneous tokens as well as the dependency relations whose performance deteriorates the most represent the unique characteristics of informal grammar (Kabiri et al., 2022). Furthermore, I investigate different transfer learning methods to transfer knowledge from the source domain (i.e. formal) to the target domain (i.e. informal) with no training data or little training data and compare the outcomes to those of a supervised learning model. The results show that when at least a small amount of training data is available for the informal domain, transfer learning techniques can significantly increase the accuracy of the parsers. This emphasizes the significance of linguistic data annotation and indicates that knowledge transfer techniques combined with small corpora can alleviate the time and cost required to develop large datasets. Despite the fact that Persian has a large number of native speakers, it remains a low-resourced language with few annotated datasets and computational models. This dissertation contributes to the availability of more annotated datasets for this language with the goal of improving the language technology infrastructure. In comparison to languages that have received more attention, particularly English, Persian has unique linguistic and orthographic characteristics, posing a variety of challenges on various levels. The methods proposed in this dissertation may facilitate developing language resources and tools for languages with similar properties. In addition, the more accurate syntactic output of parsing models will be of use for future work in a wide variety of Persian natural language processing downstream tasks including but not limited to machine translation, sentiment analysis, summarization, dialogue systems, and speech recognition. The ultimate contribution of this dissertation that demonstrates a broader impact is to provide a stepping-stone to reveal the significance of informal variants of languages, which have been widely overlooked in natural language processing tools across languages.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- doctoral
- Discipline thesis:degree_discipline
- Graduate College
- Grantor dc:publisher
- The University of Arizona.
- Year dc:date.issued
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Kabiri, Roya
- Advisors dc:contributor.advisor
-
- Karimi, Simin
- Surdeanu, Mihai
- Committee member dc:contributor.committeemember
-
- Harley, Heidi B.
Subjects
dc:subject × 6Rights
dc:rights- Statement dc:rights
-
- Copyright © is held by the author. Digital access to this material is made possible by the University Libraries, University of Arizona. Further transmission, reproduction, presentation (such as public display or performance) of protected items is prohibited except with permission of the author.
- Licence dc:rights.uri
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/10150/666154
- OAI identifier oai:identifier
- oai:repository.arizona.edu:10150/666154