Abstract
dc:descriptionTextual Inversion (TI) is one of the most recent discoveries in generative audio models that allows for prompt recaptioning of a pretrained model’s concept understanding. In this thesis, we explore a rudimentary method and TI for audio sample outputs, specifically focusing on a custom-trained TI model that embeds new concepts into pretrained image models. The TI recaptioned samples are compared with baseline samples, and we find a significant difference in variation between baseline samples. We believe that this comparison will provide a strong foundation and inspiration to future works related to prompt recaptions and analysis on generative audio models.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois Urbana-Champaign
- Year dc:date
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Matthews, Evan Michael
- Contributors dc:contributor
-
- Smaragdis, Paris
Subjects
dc:subject × 8Rights
dc:rights- Statement dc:rights
-
- Copyright 2025 Evan M. Matthews
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/129612