{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/156549"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/156549","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Learning Low-Level Priors from Images for Inference and Synthesis","abstract":"With the recent advancements in computer vision, scene understanding is critical for both downstream applications and photorealistic synthesis. Tasks such as image classification, semantic segmentation, and text-to-image generation parse the scene in terms of high-level properties of objects and scene. Along with understanding and creating visual media along these dimensions, it is important to understand the low-level information such as geometry, material, lighting configuration, and camera parameters. Such understanding would help us with tasks such as material acquisition, fine-grained synthesis, and robotics. In this thesis, we discuss learning priors over low-level properties to facilitate inference of geometry, static-dynamic disentanglement, and material properties. We present a self-supervised method to construct a persistent representation for inferring geometry and appearance inferred using a single image at test time. This representation can be leveraged to infer static-dynamic disentanglement and can used for 3D-aware scene editing. We employ representations from a pre-trained visual encoder for selecting similar materials in images. Additionally, we demonstrate fine-grained control over material properties for image editing using pre-trained text-to-image models. This fine-grained control is achieved by maintaining the photorealistic image ability of text-to-image models while learning control based on synthetic rendered images.","abstract_html":"With the recent advancements in computer vision, scene understanding is critical for both downstream applications and photorealistic synthesis. Tasks such as image classification, semantic segmentation, and text-to-image generation parse the scene in terms of high-level properties of objects and scene. Along with understanding and creating visual media along these dimensions, it is important to understand the low-level information such as geometry, material, lighting configuration, and camera parameters. Such understanding would help us with tasks such as material acquisition, fine-grained synthesis, and robotics. In this thesis, we discuss learning priors over low-level properties to facilitate inference of geometry, static-dynamic disentanglement, and material properties. We present a self-supervised method to construct a persistent representation for inferring geometry and appearance inferred using a single image at test time. This representation can be leveraged to infer static-dynamic disentanglement and can used for 3D-aware scene editing. We employ representations from a pre-trained visual encoder for selecting similar materials in images. Additionally, we demonstrate fine-grained control over material properties for image editing using pre-trained text-to-image models. This fine-grained control is achieved by maintaining the photorealistic image ability of text-to-image models while learning control based on synthetic rendered images.","abstract_has_math":false,"creators":["Sharma, Prafull"],"institution":"Massachusetts Institute of Technology","degree_name":"Doctoral","degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science","school":null,"contributors":[],"advisors":["Durand, Fredo","Freeman, William T."],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:22:09Z","subjects":[],"languages":[],"rights":["In Copyright - Educational Use Permitted","Copyright retained by author(s)"],"rights_urls":["https://rightsstatements.org/page/InC-EDU/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1721.1/156549","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Durand, Fredo","Freeman, William T."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science"]},{"key":"dc:creator","label":"Author","values":["Sharma, Prafull"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2024-09-03T21:06:32Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2024-09-03T21:06:32Z"]},{"key":"dc:date.issued","label":"Date","values":["2024-05"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctoral","Doctor of Philosophy"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright - Educational Use Permitted","Copyright retained by author(s)"]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://rightsstatements.org/page/InC-EDU/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1721.1/156549"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["With the recent advancements in computer vision, scene understanding is critical for both downstream applications and photorealistic synthesis. Tasks such as image classification, semantic segmentation, and text-to-image generation parse the scene in terms of high-level properties of objects and scene. Along with understanding and creating visual media along these dimensions, it is important to understand the low-level information such as geometry, material, lighting configuration, and camera parameters. Such understanding would help us with tasks such as material acquisition, fine-grained synthesis, and robotics. In this thesis, we discuss learning priors over low-level properties to facilitate inference of geometry, static-dynamic disentanglement, and material properties. We present a self-supervised method to construct a persistent representation for inferring geometry and appearance inferred using a single image at test time. This representation can be leveraged to infer static-dynamic disentanglement and can used for 3D-aware scene editing. We employ representations from a pre-trained visual encoder for selecting similar materials in images. Additionally, we demonstrate fine-grained control over material properties for image editing using pre-trained text-to-image models. This fine-grained control is achieved by maintaining the photorealistic image ability of text-to-image models while learning control based on synthetic rendered images."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Ph.D."]},{"key":"dc:title","label":"Title","values":["Learning Low-Level Priors from Images for Inference and Synthesis"]}]}],"canonical_facts":{"dc:contributor.advisor":["Durand, Fredo","Freeman, William T."],"dc:contributor.department":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science"],"dc:creator":["Sharma, Prafull"],"dc:date.accessioned":["2024-09-03T21:06:32Z"],"dc:date.available":["2024-09-03T21:06:32Z"],"dc:date.issued":["2024-05"],"dc:description.abstract":["With the recent advancements in computer vision, scene understanding is critical for both downstream applications and photorealistic synthesis. Tasks such as image classification, semantic segmentation, and text-to-image generation parse the scene in terms of high-level properties of objects and scene. Along with understanding and creating visual media along these dimensions, it is important to understand the low-level information such as geometry, material, lighting configuration, and camera parameters. Such understanding would help us with tasks such as material acquisition, fine-grained synthesis, and robotics. In this thesis, we discuss learning priors over low-level properties to facilitate inference of geometry, static-dynamic disentanglement, and material properties. We present a self-supervised method to construct a persistent representation for inferring geometry and appearance inferred using a single image at test time. This representation can be leveraged to infer static-dynamic disentanglement and can used for 3D-aware scene editing. We employ representations from a pre-trained visual encoder for selecting similar materials in images. Additionally, we demonstrate fine-grained control over material properties for image editing using pre-trained text-to-image models. This fine-grained control is achieved by maintaining the photorealistic image ability of text-to-image models while learning control based on synthetic rendered images."],"dc:description.degree":["Ph.D."],"dc:identifier.uri":["https://hdl.handle.net/1721.1/156549"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["In Copyright - Educational Use Permitted","Copyright retained by author(s)"],"dc:rights.uri":["https://rightsstatements.org/page/InC-EDU/1.0/"],"dc:title":["Learning Low-Level Priors from Images for Inference and Synthesis"],"dc:type":["Thesis"],"thesis:degree_name":["Doctoral","Doctor of Philosophy"]},"updated_at":"2026-07-22T22:22:09Z"}