{"id":{"repo_id":"trento","oai_identifier":"oai:iris.unitn.it:11572/448870"},"canonical_url":"https://search.dev.ndltd.org/etd/trento/oai:iris.unitn.it:11572/448870","repository":{"repo_id":"trento","name":"Università degli Studi di Trento","base_url":"https://iris.unitn.it/oai/request"},"display":{"title":"Interactive and Controlled Visual Content Generation","abstract":"The rapid expansion of the creative economy highlights the need for generative tools that empower users to produce high-quality, engaging digital content efficiently. Although advances in deep learning have significantly reduced the barriers to content creation, generative models in the visual domain often lack intuitive control and customization mechanisms, limiting their applicability in artistic and professional workflows. This thesis explores novel methodologies to enhance user-driven interaction with generative models, with a focus on improving customization, control, and creative freedom across various modalities. We propose a series of contributions addressing these challenges. First, we introduce Interactive Neural Painting, a framework that explores a brushstroke decomposition of images, enabling granular control and iterative collaboration between users and generative systems. Second, we present PAIR Diffusion, a comprehensive image editing framework that integrates diverse tasks such as inpainting, object addition, and shape modification within a single model, emphasizing seamless and interactive customization. Lastly, we extend these principles to video editing with VASE and RAGME, a system that incorporates external motion patterns to enhance the realism and control of generated video dynamics. Through these contributions, we advocate for a shift from the traditional input-output paradigm, positioning users as central participants in the creative process. By integrating retrieval-based mechanisms, interactive interfaces, and object-level control, this thesis bridges technical innovation with user-centric design, fostering a more inclusive and empowering creative economy.","abstract_html":"The rapid expansion of the creative economy highlights the need for generative tools that empower users to produce high-quality, engaging digital content efficiently. Although advances in deep learning have significantly reduced the barriers to content creation, generative models in the visual domain often lack intuitive control and customization mechanisms, limiting their applicability in artistic and professional workflows. This thesis explores novel methodologies to enhance user-driven interaction with generative models, with a focus on improving customization, control, and creative freedom across various modalities. We propose a series of contributions addressing these challenges. First, we introduce Interactive Neural Painting, a framework that explores a brushstroke decomposition of images, enabling granular control and iterative collaboration between users and generative systems. Second, we present PAIR Diffusion, a comprehensive image editing framework that integrates diverse tasks such as inpainting, object addition, and shape modification within a single model, emphasizing seamless and interactive customization. Lastly, we extend these principles to video editing with VASE and RAGME, a system that incorporates external motion patterns to enhance the realism and control of generated video dynamics. Through these contributions, we advocate for a shift from the traditional input-output paradigm, positioning users as central participants in the creative process. By integrating retrieval-based mechanisms, interactive interfaces, and object-level control, this thesis bridges technical innovation with user-centric design, fostering a more inclusive and empowering creative economy.","abstract_has_math":false,"creators":["Peruzzo, Elia"],"institution":"Università degli studi di Trento","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Sebe, Niculae"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-03-28","date_published":"2025-03-28","updated_at":"2026-07-24T05:04:45Z","subjects":["Interactive Generation, Image Editing, Video Editing, Generative Models"],"languages":["eng"],"rights":["info:eu-repo/semantics/openAccess","license:Creative commons","license uri:http://creativecommons.org/licenses/by-nc-nd/4.0/"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["http://dx.doi.org/10.15168/11572_448870","10.15168/11572_448870"],"render_values":[{"text":"http://dx.doi.org/10.15168/11572_448870","href":"http://dx.doi.org/10.15168/11572_448870","code":true},{"text":"10.15168/11572_448870","href":"https://doi.org/10.15168/11572_448870","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/11572/448870","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Peruzzo, Elia","Sebe, Niculae"]},{"key":"dc:creator","label":"Author","values":["Peruzzo, Elia"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-03-28"]},{"key":"dc:publisher","label":"Institution","values":["Università degli studi di Trento","place:TRENTO"]},{"key":"dc:relation","label":"Dc Relation","values":["firstpage:1","lastpage:110","numberofpages:110"]},{"key":"dc:type","label":"Dc Type","values":["info:eu-repo/semantics/doctoralThesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Interactive Generation, Image Editing, Video Editing, Generative Models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["info:eu-repo/semantics/openAccess","license:Creative commons","license uri:http://creativecommons.org/licenses/by-nc-nd/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/11572/448870","http://dx.doi.org/10.15168/11572_448870","10.15168/11572_448870"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The rapid expansion of the creative economy highlights the need for generative tools that empower users to produce high-quality, engaging digital content efficiently. Although advances in deep learning have significantly reduced the barriers to content creation, generative models in the visual domain often lack intuitive control and customization mechanisms, limiting their applicability in artistic and professional workflows. This thesis explores novel methodologies to enhance user-driven interaction with generative models, with a focus on improving customization, control, and creative freedom across various modalities. We propose a series of contributions addressing these challenges. First, we introduce Interactive Neural Painting, a framework that explores a brushstroke decomposition of images, enabling granular control and iterative collaboration between users and generative systems. Second, we present PAIR Diffusion, a comprehensive image editing framework that integrates diverse tasks such as inpainting, object addition, and shape modification within a single model, emphasizing seamless and interactive customization. Lastly, we extend these principles to video editing with VASE and RAGME, a system that incorporates external motion patterns to enhance the realism and control of generated video dynamics. Through these contributions, we advocate for a shift from the traditional input-output paradigm, positioning users as central participants in the creative process. By integrating retrieval-based mechanisms, interactive interfaces, and object-level control, this thesis bridges technical innovation with user-centric design, fostering a more inclusive and empowering creative economy."]},{"key":"dc:title","label":"Title","values":["Interactive and Controlled Visual Content Generation"]}]}],"canonical_facts":{"dc:contributor":["Peruzzo, Elia","Sebe, Niculae"],"dc:creator":["Peruzzo, Elia"],"dc:date":["2025-03-28"],"dc:description":["The rapid expansion of the creative economy highlights the need for generative tools that empower users to produce high-quality, engaging digital content efficiently. Although advances in deep learning have significantly reduced the barriers to content creation, generative models in the visual domain often lack intuitive control and customization mechanisms, limiting their applicability in artistic and professional workflows. This thesis explores novel methodologies to enhance user-driven interaction with generative models, with a focus on improving customization, control, and creative freedom across various modalities. We propose a series of contributions addressing these challenges. First, we introduce Interactive Neural Painting, a framework that explores a brushstroke decomposition of images, enabling granular control and iterative collaboration between users and generative systems. Second, we present PAIR Diffusion, a comprehensive image editing framework that integrates diverse tasks such as inpainting, object addition, and shape modification within a single model, emphasizing seamless and interactive customization. Lastly, we extend these principles to video editing with VASE and RAGME, a system that incorporates external motion patterns to enhance the realism and control of generated video dynamics. Through these contributions, we advocate for a shift from the traditional input-output paradigm, positioning users as central participants in the creative process. By integrating retrieval-based mechanisms, interactive interfaces, and object-level control, this thesis bridges technical innovation with user-centric design, fostering a more inclusive and empowering creative economy."],"dc:identifier":["https://hdl.handle.net/11572/448870","http://dx.doi.org/10.15168/11572_448870","10.15168/11572_448870"],"dc:language":["eng"],"dc:publisher":["Università degli studi di Trento","place:TRENTO"],"dc:relation":["firstpage:1","lastpage:110","numberofpages:110"],"dc:rights":["info:eu-repo/semantics/openAccess","license:Creative commons","license uri:http://creativecommons.org/licenses/by-nc-nd/4.0/"],"dc:subject":["Interactive Generation, Image Editing, Video Editing, Generative Models"],"dc:title":["Interactive and Controlled Visual Content Generation"],"dc:type":["info:eu-repo/semantics/doctoralThesis"]},"updated_at":"2026-07-24T05:04:45Z"}