The Graduate School and University Center of The City University of New York
Controlling Emotional Text to Speech Using Complex Adverbial Phrases
Abstract
dc:description.abstract<p>This study investigates the usage of adverbial modifiers from audiobook data as a resource for training speech synthesizers with a greater range of speech descriptions. The Tacotron2 text-to-speech (TTS) model was used for the purposes of this study. Utilizing the LibriTTS dataset, the Tacotron2 model is trained under two experimental conditions: one incorporating adverbial modifiers into the input text and the other without. The dataset preprocessing involves embedding descriptions using word embeddings and encoding speaker IDs with machine learning techniques. Additionally, the model architecture includes a prosody encoder inspired by prior research. Evaluation of the trained models involves subjective assessments by human listeners. Results from the evaluation provide insights into the efficacy of adverbial modifiers in controlling speech synthesis with a wider range of speech descriptions and show weakly positive but promising results.</p>
Degree
thesis:*- Name thesis:degree_name
- Master of Arts
- Level thesis:degree_level
- Master
- Discipline thesis:degree_discipline
- Linguistics
- Grantor
- The Graduate School and University Center of The City University of New York
- Year dc:date.available
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Akande, Zainab T
- Advisor dc:contributor.advisor
-
- Rivka Levitan
Subjects
dc:subject × 8Identifiers
dc:identifier.*- Repository record dc:identifier
- https://academicworks.cuny.edu/gc_etds/6027
- OAI identifier oai:identifier
- oai:academicworks.cuny.edu:gc_etds-7148