{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/125788"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/125788","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Improving diffusion models for enhanced accuracy, control, and realism in virtual try-on","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2026-08-01","abstract_has_math":false,"creators":["Zhang, Jeffrey"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Forsyth, David Alexander","Lazebnik, Svetlana","Schwing, Alexander Gerhard","Berg, Tamara Lee"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-07-08","date_published":"2024-07-08","updated_at":"2026-07-22T22:25:02Z","subjects":["Diffusion","Virtual Try-on","Controllability","Computer Vision","Fashion"],"languages":["en","eng"],"rights":["Copyright 2024 Jeffrey Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/125788","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Forsyth, David Alexander","Lazebnik, Svetlana","Schwing, Alexander Gerhard","Berg, Tamara Lee"]},{"key":"dc:creator","label":"Author","values":["Zhang, Jeffrey"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-07-08","2024-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Diffusion","Virtual Try-on","Controllability","Computer Vision","Fashion"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Jeffrey Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/125788"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-08-01","The student, Jeffrey Zhang, accepted the attached license on 2024-07-05 at 23:28.","The student, Jeffrey Zhang, submitted this Dissertation for approval on 2024-07-05 at 23:46.","This Dissertation was approved for publication on 2024-07-08 at 14:43.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20958 on 2025-02-04 at 21:25:28","This thesis advances virtual try-on (VTON) methods, which generate images of people wearing specific garments. We prioritize garment accuracy, which is essential for accurately representing products and ensuring a reliable virtual shopping experience. Traditionally, generative adversarial networks (GANs) and warping methods have been the best means to achieve this, and we present several production-ready VTON contributions using these techniques. However, these methods often fall short in image quality compared to newer diffusion models. Diffusion methods, while providing improved image quality, frequently fail to maintain precise garment details (changing the product) and lack the ability to style garments. This thesis merges diffusion and warping techniques to improve both accuracy and image quality, while tackling other VTON challenges, such as control of garment styles, support for a wide range of people, and fast inference speed. We address challenges in standard diffusion formulations that compromise garment accuracy and control in VTON. First, diffusion models may exhibit background artifacts and shifts in image distributions during inference. To counter these issues, we introduce consistent initialization strategies to eliminate inconsistencies between training and testing procedures, resulting in consistent image distributions and artifact reductions. Second, we tackle compression errors from variational autoencoders (VAEs) that distort critical high-frequency garment details. Our automated process identifies and upsamples high-error regions during VAE processing, mitigating these errors. Finally, diffusion methods tend to hallucinate garment details, which leads to changing garment identities and unreliable generations. We introduce a novel diffusion-based VTON training scheme that uses carefully engineered control images to ensure accurate garment details, high quality, and complete garment control. Our proposed VTON diffusion method has several key advantages due to its enhanced control. Firstly, it enables multi-garment try-on, a feature seen in a handful of prior works. Secondly, it supports fine-grain layering, styling, and shoe try-ons. Few prior works support comprehensive styling and layering features, and far fewer have addressed shoe try-on. Finally, our method supports high-quality zoomed-in image generation without the need to train or run inference in higher resolution. This contribution is unique to our system, allowing for detailed close-up images of VTON. Both qualitative results and quantitative metrics, taken together with user studies, show that our method significantly outperforms others in image quality and in accurately preserving garment details (e.g. text, logos, textures, and patterns)."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Improving diffusion models for enhanced accuracy, control, and realism in virtual try-on"]}]}],"canonical_facts":{"dc:contributor":["Forsyth, David Alexander","Lazebnik, Svetlana","Schwing, Alexander Gerhard","Berg, Tamara Lee"],"dc:creator":["Zhang, Jeffrey"],"dc:date":["2024-07-08","2024-08"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-08-01","The student, Jeffrey Zhang, accepted the attached license on 2024-07-05 at 23:28.","The student, Jeffrey Zhang, submitted this Dissertation for approval on 2024-07-05 at 23:46.","This Dissertation was approved for publication on 2024-07-08 at 14:43.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20958 on 2025-02-04 at 21:25:28","This thesis advances virtual try-on (VTON) methods, which generate images of people wearing specific garments. We prioritize garment accuracy, which is essential for accurately representing products and ensuring a reliable virtual shopping experience. Traditionally, generative adversarial networks (GANs) and warping methods have been the best means to achieve this, and we present several production-ready VTON contributions using these techniques. However, these methods often fall short in image quality compared to newer diffusion models. Diffusion methods, while providing improved image quality, frequently fail to maintain precise garment details (changing the product) and lack the ability to style garments. This thesis merges diffusion and warping techniques to improve both accuracy and image quality, while tackling other VTON challenges, such as control of garment styles, support for a wide range of people, and fast inference speed. We address challenges in standard diffusion formulations that compromise garment accuracy and control in VTON. First, diffusion models may exhibit background artifacts and shifts in image distributions during inference. To counter these issues, we introduce consistent initialization strategies to eliminate inconsistencies between training and testing procedures, resulting in consistent image distributions and artifact reductions. Second, we tackle compression errors from variational autoencoders (VAEs) that distort critical high-frequency garment details. Our automated process identifies and upsamples high-error regions during VAE processing, mitigating these errors. Finally, diffusion methods tend to hallucinate garment details, which leads to changing garment identities and unreliable generations. We introduce a novel diffusion-based VTON training scheme that uses carefully engineered control images to ensure accurate garment details, high quality, and complete garment control. Our proposed VTON diffusion method has several key advantages due to its enhanced control. Firstly, it enables multi-garment try-on, a feature seen in a handful of prior works. Secondly, it supports fine-grain layering, styling, and shoe try-ons. Few prior works support comprehensive styling and layering features, and far fewer have addressed shoe try-on. Finally, our method supports high-quality zoomed-in image generation without the need to train or run inference in higher resolution. This contribution is unique to our system, allowing for detailed close-up images of VTON. Both qualitative results and quantitative metrics, taken together with user studies, show that our method significantly outperforms others in image quality and in accurately preserving garment details (e.g. text, logos, textures, and patterns)."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/125788"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Jeffrey Zhang"],"dc:subject":["Diffusion","Virtual Try-on","Controllability","Computer Vision","Fashion"],"dc:title":["Improving diffusion models for enhanced accuracy, control, and realism in virtual try-on"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}