{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132482"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132482","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Generative 3D scene modeling: algorithm and applications","abstract":"The ability to generate realistic and functional 3D scenes at scale is becoming increasingly important across domains such as agriculture, autonomous driving, immersive gaming, and robot learning. However, manually constructing such environments is prohibitively expensive and often introduces domain gaps from real-world distributions. Generative modeling of 3D scenes offers a scalable alternative, with the promise of producing vast quantities of coherent, diverse, and controllable environments directly from data. This thesis aims to advance generative 3D scene modeling with a particular emphasis on large, outdoor, and multi-modal environments. We define two desiderata for an ideal generative scene model: the ability to imagine worlds beyond the observed extent (scene generation), and the ability to reinterpret observed environments with alternative variations (scene editing). Realizing these goals is far from trivial, raising three key research challenges: (1) how to enable large-scale scene generation with limited 3D data; (2) how to move beyond single-modality generation to handle diverse and integrated scene modalities; and (3) how to ensure that generative scene models bring tangible benefits to downstream applications. Our approach to these challenges is developed along two complementary dimensions: algorithms and applications. Algorithmically, the guiding philosophy is to leverage powerful 2D generative priors to overcome the scarcity of high-quality 3D data. This leads to two chapters: (1) SuperGaussian, which adapts pre-trained video models for 3D scene editing, showing that video priors can be repurposed for high-fidelity 3D super-resolution; and (2) SGAM, a progressive framework that combines generative sensor modeling with reconstruction to create globally consistent, large-scale virtual worlds from RGB-D sequences. On the application side, we demonstrate how generative scene modeling can directly support real-world use cases. In particular, Sim-on-Wheels provides a safe and realistic framework for testing autonomous driving by augmenting real-world driving scenes with virtual tra\"c scenarios, while MMCityGen showcases how multi-modal generation can power urban digital twins for planning and analysis. Together, these contributions establish a principled approach to repurposing 2D priors for 3D scene modeling, extend generation into the multi-modal domain, and demonstrate concrete uses in autonomy and urban planning.","abstract_html":"The ability to generate realistic and functional 3D scenes at scale is becoming increasingly important across domains such as agriculture, autonomous driving, immersive gaming, and robot learning. However, manually constructing such environments is prohibitively expensive and often introduces domain gaps from real-world distributions. Generative modeling of 3D scenes offers a scalable alternative, with the promise of producing vast quantities of coherent, diverse, and controllable environments directly from data. This thesis aims to advance generative 3D scene modeling with a particular emphasis on large, outdoor, and multi-modal environments. We define two desiderata for an ideal generative scene model: the ability to imagine worlds beyond the observed extent (scene generation), and the ability to reinterpret observed environments with alternative variations (scene editing). Realizing these goals is far from trivial, raising three key research challenges: (1) how to enable large-scale scene generation with limited 3D data; (2) how to move beyond single-modality generation to handle diverse and integrated scene modalities; and (3) how to ensure that generative scene models bring tangible benefits to downstream applications. Our approach to these challenges is developed along two complementary dimensions: algorithms and applications. Algorithmically, the guiding philosophy is to leverage powerful 2D generative priors to overcome the scarcity of high-quality 3D data. This leads to two chapters: (1) SuperGaussian, which adapts pre-trained video models for 3D scene editing, showing that video priors can be repurposed for high-fidelity 3D super-resolution; and (2) SGAM, a progressive framework that combines generative sensor modeling with reconstruction to create globally consistent, large-scale virtual worlds from RGB-D sequences. On the application side, we demonstrate how generative scene modeling can directly support real-world use cases. In particular, Sim-on-Wheels provides a safe and realistic framework for testing autonomous driving by augmenting real-world driving scenes with virtual tra&quot;c scenarios, while MMCityGen showcases how multi-modal generation can power urban digital twins for planning and analysis. Together, these contributions establish a principled approach to repurposing 2D priors for 3D scene modeling, extend generation into the multi-modal domain, and demonstrate concrete uses in autonomy and urban planning.","abstract_has_math":false,"creators":["Shen, Yuan"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Wang, Shenlong","Hoiem, Derek","Forsyth, David","Ceylan, Duygu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["3d scene generation","3d scene editing"],"languages":["en"],"rights":["Copyright 2025 Yuan Shen"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132482","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Wang, Shenlong","Hoiem, Derek","Forsyth, David","Ceylan, Duygu"]},{"key":"dc:creator","label":"Author","values":["Shen, Yuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-11-11"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["3d scene generation","3d scene editing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Yuan Shen"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132482"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The ability to generate realistic and functional 3D scenes at scale is becoming increasingly important across domains such as agriculture, autonomous driving, immersive gaming, and robot learning. However, manually constructing such environments is prohibitively expensive and often introduces domain gaps from real-world distributions. Generative modeling of 3D scenes offers a scalable alternative, with the promise of producing vast quantities of coherent, diverse, and controllable environments directly from data. This thesis aims to advance generative 3D scene modeling with a particular emphasis on large, outdoor, and multi-modal environments. We define two desiderata for an ideal generative scene model: the ability to imagine worlds beyond the observed extent (scene generation), and the ability to reinterpret observed environments with alternative variations (scene editing). Realizing these goals is far from trivial, raising three key research challenges: (1) how to enable large-scale scene generation with limited 3D data; (2) how to move beyond single-modality generation to handle diverse and integrated scene modalities; and (3) how to ensure that generative scene models bring tangible benefits to downstream applications. Our approach to these challenges is developed along two complementary dimensions: algorithms and applications. Algorithmically, the guiding philosophy is to leverage powerful 2D generative priors to overcome the scarcity of high-quality 3D data. This leads to two chapters: (1) SuperGaussian, which adapts pre-trained video models for 3D scene editing, showing that video priors can be repurposed for high-fidelity 3D super-resolution; and (2) SGAM, a progressive framework that combines generative sensor modeling with reconstruction to create globally consistent, large-scale virtual worlds from RGB-D sequences. On the application side, we demonstrate how generative scene modeling can directly support real-world use cases. In particular, Sim-on-Wheels provides a safe and realistic framework for testing autonomous driving by augmenting real-world driving scenes with virtual tra\"c scenarios, while MMCityGen showcases how multi-modal generation can power urban digital twins for planning and analysis. Together, these contributions establish a principled approach to repurposing 2D priors for 3D scene modeling, extend generation into the multi-modal domain, and demonstrate concrete uses in autonomy and urban planning.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Yuan Shen, accepted the attached license on 2025-11-10 at 17:29.","The student, Yuan Shen, submitted this Dissertation for approval on 2025-11-10 at 17:40.","This Dissertation was approved for publication on 2025-11-11 at 11:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22856 on 2026-02-19 at 18:24:31"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Generative 3D scene modeling: algorithm and applications"]}]}],"canonical_facts":{"dc:contributor":["Wang, Shenlong","Hoiem, Derek","Forsyth, David","Ceylan, Duygu"],"dc:creator":["Shen, Yuan"],"dc:date":["2025-12","2025-11-11"],"dc:description":["The ability to generate realistic and functional 3D scenes at scale is becoming increasingly important across domains such as agriculture, autonomous driving, immersive gaming, and robot learning. However, manually constructing such environments is prohibitively expensive and often introduces domain gaps from real-world distributions. Generative modeling of 3D scenes offers a scalable alternative, with the promise of producing vast quantities of coherent, diverse, and controllable environments directly from data. This thesis aims to advance generative 3D scene modeling with a particular emphasis on large, outdoor, and multi-modal environments. We define two desiderata for an ideal generative scene model: the ability to imagine worlds beyond the observed extent (scene generation), and the ability to reinterpret observed environments with alternative variations (scene editing). Realizing these goals is far from trivial, raising three key research challenges: (1) how to enable large-scale scene generation with limited 3D data; (2) how to move beyond single-modality generation to handle diverse and integrated scene modalities; and (3) how to ensure that generative scene models bring tangible benefits to downstream applications. Our approach to these challenges is developed along two complementary dimensions: algorithms and applications. Algorithmically, the guiding philosophy is to leverage powerful 2D generative priors to overcome the scarcity of high-quality 3D data. This leads to two chapters: (1) SuperGaussian, which adapts pre-trained video models for 3D scene editing, showing that video priors can be repurposed for high-fidelity 3D super-resolution; and (2) SGAM, a progressive framework that combines generative sensor modeling with reconstruction to create globally consistent, large-scale virtual worlds from RGB-D sequences. On the application side, we demonstrate how generative scene modeling can directly support real-world use cases. In particular, Sim-on-Wheels provides a safe and realistic framework for testing autonomous driving by augmenting real-world driving scenes with virtual tra\"c scenarios, while MMCityGen showcases how multi-modal generation can power urban digital twins for planning and analysis. Together, these contributions establish a principled approach to repurposing 2D priors for 3D scene modeling, extend generation into the multi-modal domain, and demonstrate concrete uses in autonomy and urban planning.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Yuan Shen, accepted the attached license on 2025-11-10 at 17:29.","The student, Yuan Shen, submitted this Dissertation for approval on 2025-11-10 at 17:40.","This Dissertation was approved for publication on 2025-11-11 at 11:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22856 on 2026-02-19 at 18:24:31"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132482"],"dc:language":["en"],"dc:rights":["Copyright 2025 Yuan Shen"],"dc:subject":["3d scene generation","3d scene editing"],"dc:title":["Generative 3D scene modeling: algorithm and applications"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}