Virginia Tech
A Submodular Approach to Find Interpretable Directions in Text-to-Image Models
Abstract
dc:description.abstractText-to-image models have significantly improved the field of image editing. However, finding attributes that the model can actually edit is still a remaining challenge. This thesis proposes a solution to this problem by leveraging a multimodal vision-language model (MMVLM) to find a list of potential attributes that can be used to edit an image, using Flux and ControlNet to generate edits using those keywords, and then applying a submodular ranking method to find which edits actually work. The experiments in this paper demonstrate the robustness of this approach and its ability to produce high-quality edits across various domains, such as dresses and living rooms.
Degree
thesis:*- Name thesis:degree_name
- Master of Science
- Level thesis:degree_level
- masters
- Discipline thesis:degree_discipline
- Computer Science & Applications
- Department dc:contributor.department
- Computer Science and#38; Applications
- Grantor dc:publisher
- Virginia Tech
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Allada, Ritika
- Chair dc:contributor.committeechair
-
- Yanardag Delul, Pinar
- Committee members dc:contributor.committeemember
-
- Eldardiry, Hoda Mohamed
- Thomas, Christopher Lee
- North, Christopher L.
Subjects
dc:subject × 5Rights
dc:rights- Statement dc:rights
-
- Creative Commons Attribution 4.0 International
- Licence dc:rights.uri
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Dc Identifier Other
- vt_gsexam:44224
- OAI identifier oai:identifier
- oai:vtechworks.lib.vt.edu:10919/135475