{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129184"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129184","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Modeling and editing 4D scenes by leveraging structural priors","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Lyu, Jipeng"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Wang, Yuxiong"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-04-22","date_published":"2025-04-22","updated_at":"2026-07-22T22:25:04Z","subjects":["4D Scene Understanding","Geometric Structural Priors","Semantic Video Editing","3D Gaussian Splatting","Diffusion Models"],"languages":["en","eng"],"rights":["Copyright 2025 Jipeng Lyu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129184","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Wang, Yuxiong"]},{"key":"dc:creator","label":"Author","values":["Lyu, Jipeng"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-04-22","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["4D Scene Understanding","Geometric Structural Priors","Semantic Video Editing","3D Gaussian Splatting","Diffusion Models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Jipeng Lyu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129184"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Jipeng Lyu, accepted the attached license on 2025-04-21 at 17:19.","The student, Jipeng Lyu, submitted this Thesis for approval on 2025-04-21 at 17:36.","This Thesis was approved for publication on 2025-04-22 at 15:04.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21702 on 2025-10-19 at 18:09:13","4D scene understanding powers applications ranging from AR/VR and robotics to controllable video generation. The goal is to build representations that can faithfully model and manipulate real-world environments as they evolve over time—capturing both how scenes deform (modeling) and how they respond to high-level user instructions (editing). A popular approach is to represent scenes using compact geometric primitives such as 3D Gaussians, which enable efficient rendering and temporal consistency across frames. While recent advances in 3D reconstruction and video editing have shown promising results, many existing methods still overlook a key aspect: structure. Whether in the form of spatial rigidity or semantic hierarchy, structural priors are both abundant and underexplored. This thesis investigates how incorporating such priors—geometric and semantic—into 4D scene modeling and editing can enhance efficiency, controllability, and generalization. The first part of this thesis focuses on geometric structural priors in dynamic 3D modeling. Many dynamic scenes exhibit coherent change patterns: objects often deform in groups, move rigidly or semi-rigidly, or follow interpretable part-wise trajectories. Instead of modeling motion independently for each element, we propose a structural cascaded optimization framework that organizes 3D Gaussians into a coarse-to-fine hierarchy. This structure allows us to parameterize deformation using simple transformations—rotation, translation, and scaling—substantially accelerating optimization. It also enables dense point tracking and motion-based segmentation without requiring semantic labels. These results demonstrate the potential of structured representations for fast and interpretable 4D scene modeling. The second part explores semantic structural priors in video editing. User instructions often involve multiple entangled goals that are difficult to fulfill through a single transformation. To address this, we employ large language models (LLMs) to decompose complex prompts into interpretable semantic subgoals. Each subgoal defines an editing stage, executed within a training-free diffusion-based video editing framework. To accommodate varying subgoal complexity, we further prompt the LLM to estimate editing difficulty and adapt the interpolation schedule accordingly. This results in smoother transitions and robust edits, transforming the process into a semantically grounded and interpretable sequence. Together, these contributions highlight the value of structural reasoning in 4D scene understanding. By bridging geometric modeling and semantic editing, this thesis offers unified insights into building efficient, robust, and controllable 4D systems guided by structural priors."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Modeling and editing 4D scenes by leveraging structural priors"]}]}],"canonical_facts":{"dc:contributor":["Wang, Yuxiong"],"dc:creator":["Lyu, Jipeng"],"dc:date":["2025-04-22","2025-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Jipeng Lyu, accepted the attached license on 2025-04-21 at 17:19.","The student, Jipeng Lyu, submitted this Thesis for approval on 2025-04-21 at 17:36.","This Thesis was approved for publication on 2025-04-22 at 15:04.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21702 on 2025-10-19 at 18:09:13","4D scene understanding powers applications ranging from AR/VR and robotics to controllable video generation. The goal is to build representations that can faithfully model and manipulate real-world environments as they evolve over time—capturing both how scenes deform (modeling) and how they respond to high-level user instructions (editing). A popular approach is to represent scenes using compact geometric primitives such as 3D Gaussians, which enable efficient rendering and temporal consistency across frames. While recent advances in 3D reconstruction and video editing have shown promising results, many existing methods still overlook a key aspect: structure. Whether in the form of spatial rigidity or semantic hierarchy, structural priors are both abundant and underexplored. This thesis investigates how incorporating such priors—geometric and semantic—into 4D scene modeling and editing can enhance efficiency, controllability, and generalization. The first part of this thesis focuses on geometric structural priors in dynamic 3D modeling. Many dynamic scenes exhibit coherent change patterns: objects often deform in groups, move rigidly or semi-rigidly, or follow interpretable part-wise trajectories. Instead of modeling motion independently for each element, we propose a structural cascaded optimization framework that organizes 3D Gaussians into a coarse-to-fine hierarchy. This structure allows us to parameterize deformation using simple transformations—rotation, translation, and scaling—substantially accelerating optimization. It also enables dense point tracking and motion-based segmentation without requiring semantic labels. These results demonstrate the potential of structured representations for fast and interpretable 4D scene modeling. The second part explores semantic structural priors in video editing. User instructions often involve multiple entangled goals that are difficult to fulfill through a single transformation. To address this, we employ large language models (LLMs) to decompose complex prompts into interpretable semantic subgoals. Each subgoal defines an editing stage, executed within a training-free diffusion-based video editing framework. To accommodate varying subgoal complexity, we further prompt the LLM to estimate editing difficulty and adapt the interpolation schedule accordingly. This results in smoother transitions and robust edits, transforming the process into a semantically grounded and interpretable sequence. Together, these contributions highlight the value of structural reasoning in 4D scene understanding. By bridging geometric modeling and semantic editing, this thesis offers unified insights into building efficient, robust, and controllable 4D systems guided by structural priors."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129184"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Jipeng Lyu"],"dc:subject":["4D Scene Understanding","Geometric Structural Priors","Semantic Video Editing","3D Gaussian Splatting","Diffusion Models"],"dc:title":["Modeling and editing 4D scenes by leveraging structural priors"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}