{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132639"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132639","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"From objects to worlds: scalable learning of 3D assets","abstract":"Learning to reconstruct and generate the 3D world is a fundamental research problem in computer vision, with critical applications across diverse domains. However, the development of robust 3D generation and reconstruction systems is hindered by the scarcity of high-quality 3D data. This thesis aims to address this scaling challenge along several dimensions. First, we introduce ShapeClipper, which leverages semantic consistency from unlabeled 2D images to learn 3D shape reconstruction models. This enables scalable 3D learning from only single-view images, without any 3D annotations. Second, we present PointInfinity, a resolution-invariant point diffusion model for learning continuous 3D surfaces from point clouds. PointInfinity facilitates 3D learning using noisy point clouds derived from object-centric videos. Third, we introduce ZeroShape and re-examine the classical regression-based 3D reconstruction approach. We show it outperforms diffusion methods in accuracy, as well as computational and data efficiency. Finally, we explore the feasibility of learning 3D from in-the-wild videos without any 3D prior or data. As an initial yet solid step, we evaluate the 3D awareness of recent video foundation models, and find that state-of-the-art video generative models already possess strong 3D understanding. Together, this thesis makes significant advancements in scalable learning of 3D, providing practical solutions for reconstruction and generating 3D objects and worlds under limited high-quality 3D data.","abstract_html":"Learning to reconstruct and generate the 3D world is a fundamental research problem in computer vision, with critical applications across diverse domains. However, the development of robust 3D generation and reconstruction systems is hindered by the scarcity of high-quality 3D data. This thesis aims to address this scaling challenge along several dimensions. First, we introduce ShapeClipper, which leverages semantic consistency from unlabeled 2D images to learn 3D shape reconstruction models. This enables scalable 3D learning from only single-view images, without any 3D annotations. Second, we present PointInfinity, a resolution-invariant point diffusion model for learning continuous 3D surfaces from point clouds. PointInfinity facilitates 3D learning using noisy point clouds derived from object-centric videos. Third, we introduce ZeroShape and re-examine the classical regression-based 3D reconstruction approach. We show it outperforms diffusion methods in accuracy, as well as computational and data efficiency. Finally, we explore the feasibility of learning 3D from in-the-wild videos without any 3D prior or data. As an initial yet solid step, we evaluate the 3D awareness of recent video foundation models, and find that state-of-the-art video generative models already possess strong 3D understanding. Together, this thesis makes significant advancements in scalable learning of 3D, providing practical solutions for reconstruction and generating 3D objects and worlds under limited high-quality 3D data.","abstract_has_math":false,"creators":["Huang, Zixuan"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Rehg, James M.","Schwing, Alexander","Wang, Shenlong","Wu, Jiajun","Vedaldi, Andrea"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["3D Generation","3D Reconstruction","Video Generation","World Models"],"languages":["en"],"rights":["Copyright 2025 Zixuan Huang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132639","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Rehg, James M.","Schwing, Alexander","Wang, Shenlong","Wu, Jiajun","Vedaldi, Andrea"]},{"key":"dc:creator","label":"Author","values":["Huang, Zixuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-11-13"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["3D Generation","3D Reconstruction","Video Generation","World Models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Zixuan Huang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132639"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Learning to reconstruct and generate the 3D world is a fundamental research problem in computer vision, with critical applications across diverse domains. However, the development of robust 3D generation and reconstruction systems is hindered by the scarcity of high-quality 3D data. This thesis aims to address this scaling challenge along several dimensions. First, we introduce ShapeClipper, which leverages semantic consistency from unlabeled 2D images to learn 3D shape reconstruction models. This enables scalable 3D learning from only single-view images, without any 3D annotations. Second, we present PointInfinity, a resolution-invariant point diffusion model for learning continuous 3D surfaces from point clouds. PointInfinity facilitates 3D learning using noisy point clouds derived from object-centric videos. Third, we introduce ZeroShape and re-examine the classical regression-based 3D reconstruction approach. We show it outperforms diffusion methods in accuracy, as well as computational and data efficiency. Finally, we explore the feasibility of learning 3D from in-the-wild videos without any 3D prior or data. As an initial yet solid step, we evaluate the 3D awareness of recent video foundation models, and find that state-of-the-art video generative models already possess strong 3D understanding. Together, this thesis makes significant advancements in scalable learning of 3D, providing practical solutions for reconstruction and generating 3D objects and worlds under limited high-quality 3D data.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-12-01","The student, Zixuan Huang, accepted the attached license on 2025-11-12 at 15:58.","The student, Zixuan Huang, submitted this Dissertation for approval on 2025-11-12 at 16:13.","This Dissertation was approved for publication on 2025-11-13 at 10:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22864 on 2026-02-19 at 18:45:41"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["From objects to worlds: scalable learning of 3D assets"]}]}],"canonical_facts":{"dc:contributor":["Rehg, James M.","Schwing, Alexander","Wang, Shenlong","Wu, Jiajun","Vedaldi, Andrea"],"dc:creator":["Huang, Zixuan"],"dc:date":["2025-12","2025-11-13"],"dc:description":["Learning to reconstruct and generate the 3D world is a fundamental research problem in computer vision, with critical applications across diverse domains. However, the development of robust 3D generation and reconstruction systems is hindered by the scarcity of high-quality 3D data. This thesis aims to address this scaling challenge along several dimensions. First, we introduce ShapeClipper, which leverages semantic consistency from unlabeled 2D images to learn 3D shape reconstruction models. This enables scalable 3D learning from only single-view images, without any 3D annotations. Second, we present PointInfinity, a resolution-invariant point diffusion model for learning continuous 3D surfaces from point clouds. PointInfinity facilitates 3D learning using noisy point clouds derived from object-centric videos. Third, we introduce ZeroShape and re-examine the classical regression-based 3D reconstruction approach. We show it outperforms diffusion methods in accuracy, as well as computational and data efficiency. Finally, we explore the feasibility of learning 3D from in-the-wild videos without any 3D prior or data. As an initial yet solid step, we evaluate the 3D awareness of recent video foundation models, and find that state-of-the-art video generative models already possess strong 3D understanding. Together, this thesis makes significant advancements in scalable learning of 3D, providing practical solutions for reconstruction and generating 3D objects and worlds under limited high-quality 3D data.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-12-01","The student, Zixuan Huang, accepted the attached license on 2025-11-12 at 15:58.","The student, Zixuan Huang, submitted this Dissertation for approval on 2025-11-12 at 16:13.","This Dissertation was approved for publication on 2025-11-13 at 10:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22864 on 2026-02-19 at 18:45:41"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132639"],"dc:language":["en"],"dc:rights":["Copyright 2025 Zixuan Huang"],"dc:subject":["3D Generation","3D Reconstruction","Video Generation","World Models"],"dc:title":["From objects to worlds: scalable learning of 3D assets"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}