Back to results

University of Illinois Urbana-Champaign

Text recaptioning for audio diffusion models

Abstract

dc:description

Textual Inversion (TI) is one of the most recent discoveries in generative audio models that allows for prompt recaptioning of a pretrained model’s concept understanding. In this thesis, we explore a rudimentary method and TI for audio sample outputs, specifically focusing on a custom-trained TI model that embeds new concepts into pretrained image models. The TI recaptioned samples are compared with baseline samples, and we find a significant difference in variation between baseline samples. We believe that this comparison will provide a strong foundation and inspiration to future works related to prompt recaptions and analysis on generative audio models.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois Urbana-Champaign
Year dc:date
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Matthews, Evan Michael
Contributors dc:contributor
  • Smaragdis, Paris

Subjects

dc:subject × 8

Rights

dc:rights
Statement dc:rights
  • Copyright 2025 Evan M. Matthews
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/129612

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Matthews, Evan Michael. Text recaptioning for audio diffusion models. Thesis thesis, University of Illinois Urbana-Champaign, 2025. https://hdl.handle.net/2142/129612