{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/135475"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/135475","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"A Submodular Approach to Find Interpretable Directions in Text-to-Image Models","abstract":"Text-to-image models have significantly improved the field of image editing. However, finding attributes that the model can actually edit is still a remaining challenge. This thesis proposes a solution to this problem by leveraging a multimodal vision-language model (MMVLM) to find a list of potential attributes that can be used to edit an image, using Flux and ControlNet to generate edits using those keywords, and then applying a submodular ranking method to find which edits actually work. The experiments in this paper demonstrate the robustness of this approach and its ability to produce high-quality edits across various domains, such as dresses and living rooms.","abstract_html":"Text-to-image models have significantly improved the field of image editing. However, finding attributes that the model can actually edit is still a remaining challenge. This thesis proposes a solution to this problem by leveraging a multimodal vision-language model (MMVLM) to find a list of potential attributes that can be used to edit an image, using Flux and ControlNet to generate edits using those keywords, and then applying a submodular ranking method to find which edits actually work. The experiments in this paper demonstrate the robustness of this approach and its ability to produce high-quality edits across various domains, such as dresses and living rooms.","abstract_has_math":false,"creators":["Allada, Ritika"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Computer Science & Applications","degree_department":"Computer Science and#38; Applications","school":null,"contributors":[],"advisors":[],"committee_chairs":["Yanardag Delul, Pinar"],"committee_members":["Eldardiry, Hoda Mohamed","Thomas, Christopher Lee","North, Christopher L."],"year":2025,"date_issued":"2025-06-10","date_published":"2025-06-10","updated_at":"2026-07-22T22:19:01Z","subjects":["Diffusion Models","Interpretability","Image Editing","Explainable AI","Recommendation Systems"],"languages":["en"],"rights":["Creative Commons Attribution 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44224"],"render_values":[{"text":"vt_gsexam:44224","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/135475","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Yanardag Delul, Pinar"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Eldardiry, Hoda Mohamed","Thomas, Christopher Lee","North, Christopher L."]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science and#38; Applications"]},{"key":"dc:creator","label":"Author","values":["Allada, Ritika"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-06-11T08:04:26Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-06-11T08:04:26Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-06-10"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science & Applications"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Diffusion Models","Interpretability","Image Editing","Explainable AI","Recommendation Systems"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Creative Commons Attribution 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44224"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/135475"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Text-to-image models have significantly improved the field of image editing. However, finding attributes that the model can actually edit is still a remaining challenge. This thesis proposes a solution to this problem by leveraging a multimodal vision-language model (MMVLM) to find a list of potential attributes that can be used to edit an image, using Flux and ControlNet to generate edits using those keywords, and then applying a submodular ranking method to find which edits actually work. The experiments in this paper demonstrate the robustness of this approach and its ability to produce high-quality edits across various domains, such as dresses and living rooms."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["In today's world, generative AI models are capable of editing images based on user-specified text prompts. However, finding attributes that the model can actually edit is a time-consuming process. This thesis proposes a solution to this problem by proposing a submodular ranking function that provides users with a list of top attributes that a model can actually edit a particular image with. Compared to existing editing methods, this method is able to find more meaningful attributes and produce high-quality edits across various domains, including fashion and interior design."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["A Submodular Approach to Find Interpretable Directions in Text-to-Image Models"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Yanardag Delul, Pinar"],"dc:contributor.committeemember":["Eldardiry, Hoda Mohamed","Thomas, Christopher Lee","North, Christopher L."],"dc:contributor.department":["Computer Science and#38; Applications"],"dc:creator":["Allada, Ritika"],"dc:date.accessioned":["2025-06-11T08:04:26Z"],"dc:date.available":["2025-06-11T08:04:26Z"],"dc:date.issued":["2025-06-10"],"dc:description.abstract":["Text-to-image models have significantly improved the field of image editing. However, finding attributes that the model can actually edit is still a remaining challenge. This thesis proposes a solution to this problem by leveraging a multimodal vision-language model (MMVLM) to find a list of potential attributes that can be used to edit an image, using Flux and ControlNet to generate edits using those keywords, and then applying a submodular ranking method to find which edits actually work. The experiments in this paper demonstrate the robustness of this approach and its ability to produce high-quality edits across various domains, such as dresses and living rooms."],"dc:description.abstractgeneral":["In today's world, generative AI models are capable of editing images based on user-specified text prompts. However, finding attributes that the model can actually edit is a time-consuming process. This thesis proposes a solution to this problem by proposing a submodular ranking function that provides users with a list of top attributes that a model can actually edit a particular image with. Compared to existing editing methods, this method is able to find more meaningful attributes and produce high-quality edits across various domains, including fashion and interior design."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:44224"],"dc:identifier.uri":["https://hdl.handle.net/10919/135475"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["Creative Commons Attribution 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by/4.0/"],"dc:subject":["Diffusion Models","Interpretability","Image Editing","Explainable AI","Recommendation Systems"],"dc:title":["A Submodular Approach to Find Interpretable Directions in Text-to-Image Models"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science & Applications"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:01Z"}