{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124714"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124714","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards controllable and consistent 3D scene editing","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2026-05-01","abstract_has_math":false,"creators":["Dong, Jiahua"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Wang, Yuxiong"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:02Z","subjects":["3d Editing","Diffusion Model","Computer Vision"],"languages":["en","eng"],"rights":["Copyright 2024 Jiahua Dong"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124714","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Wang, Yuxiong"]},{"key":"dc:creator","label":"Author","values":["Dong, Jiahua"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-05-01"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["3d Editing","Diffusion Model","Computer Vision"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Jiahua Dong"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124714"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-05-01","The student, Jiahua Dong, accepted the attached license on 2024-04-29 at 13:14.","The student, Jiahua Dong, submitted this Thesis for approval on 2024-04-29 at 13:29.","This Thesis was approved for publication on 2024-05-01 at 09:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20693 on 2024-09-16 at 00:51:07","Recent advancements in diffusion-based generative models have unlocked the potential for transformative 3D content creation. However, existing approaches for 3D scene editing face challenges in efficiency, consistency, and controllability. In this thesis, we introduce two novel methods that address these limitations and provide more effective means for editing 3D scenes. Our first exploration focuses on global style editing with text instructions. We present ViCA-NeRF, the first view-consistency-aware method for 3D editing with text instructions. By leveraging depth information derived from neural radiance fields (NeRF) and aligning latent codes in the 2D diffusion model, ViCA-NeRF ensures multi-view consistency through geometric and learned regularization. Our two-stage approach, involving edit blending and refinement, results in more flexible, efficient, and detailed editing compared to state-of-the-art methods. Secondly, we introduce DragGaussian, an interactive framework for fine-grained 3D scene editing using intuitive drag manipulation. Our method ensures 3D consistency through Gaussian deformation-based geometric guidance and an efficient inverse-free drag editing approach. By fine-tuning 2D diffusion models with a history-aware tuning strategy, DragGaussian enables high-quality drag-based editing of geometry and appearance, supporting various manipulations such as object movement, pose adjustment, and shape modification. Experimental results demonstrate the effectiveness and versatility of our proposed methods. ViCA-NeRF achieves a 3x speedup in editing time while maintaining higher levels of consistency and detail. DragGaussian enables diverse edits within a 10-20 minute timeframe on a single GPU, showcasing its efficiency. By combining text-based instructions and interactive point-based manipulation, our methods significantly advance the field of 3D scene editing, providing more efficient, consistent, and controllable tools for content creation."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards controllable and consistent 3D scene editing"]}]}],"canonical_facts":{"dc:contributor":["Wang, Yuxiong"],"dc:creator":["Dong, Jiahua"],"dc:date":["2024-05","2024-05-01"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-05-01","The student, Jiahua Dong, accepted the attached license on 2024-04-29 at 13:14.","The student, Jiahua Dong, submitted this Thesis for approval on 2024-04-29 at 13:29.","This Thesis was approved for publication on 2024-05-01 at 09:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20693 on 2024-09-16 at 00:51:07","Recent advancements in diffusion-based generative models have unlocked the potential for transformative 3D content creation. However, existing approaches for 3D scene editing face challenges in efficiency, consistency, and controllability. In this thesis, we introduce two novel methods that address these limitations and provide more effective means for editing 3D scenes. Our first exploration focuses on global style editing with text instructions. We present ViCA-NeRF, the first view-consistency-aware method for 3D editing with text instructions. By leveraging depth information derived from neural radiance fields (NeRF) and aligning latent codes in the 2D diffusion model, ViCA-NeRF ensures multi-view consistency through geometric and learned regularization. Our two-stage approach, involving edit blending and refinement, results in more flexible, efficient, and detailed editing compared to state-of-the-art methods. Secondly, we introduce DragGaussian, an interactive framework for fine-grained 3D scene editing using intuitive drag manipulation. Our method ensures 3D consistency through Gaussian deformation-based geometric guidance and an efficient inverse-free drag editing approach. By fine-tuning 2D diffusion models with a history-aware tuning strategy, DragGaussian enables high-quality drag-based editing of geometry and appearance, supporting various manipulations such as object movement, pose adjustment, and shape modification. Experimental results demonstrate the effectiveness and versatility of our proposed methods. ViCA-NeRF achieves a 3x speedup in editing time while maintaining higher levels of consistency and detail. DragGaussian enables diverse edits within a 10-20 minute timeframe on a single GPU, showcasing its efficiency. By combining text-based instructions and interactive point-based manipulation, our methods significantly advance the field of 3D scene editing, providing more efficient, consistent, and controllable tools for content creation."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124714"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Jiahua Dong"],"dc:subject":["3d Editing","Diffusion Model","Computer Vision"],"dc:title":["Towards controllable and consistent 3D scene editing"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}