{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/92750"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/92750","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Visual analysis and synthesis with physically grounded constraints","abstract":"\"The past decade has witnessed remarkable progress in image-based, data-driven vision and graphics. However, existing approaches often treat the images as pure 2D signals and not as a 2D projection of the physical 3D world. As a result, a lot of training examples are required to cover sufficiently diverse appearances and inevitably suffer from limited generalization capability. In this thesis, I propose \"\"inference-by-composition\"\" approaches to overcome these limitations by modeling and interpreting visual signals in terms of physical surface, object, and scene. I show how we can incorporate physically grounded constraints such as scene-specific geometry in a non-parametric optimization framework for (1) revealing the missing parts of an image due to removal of a foreground or background element, (2) recovering high spatial frequency details that are not resolvable in low-resolution observations. I then extend the framework from 2D images to handle spatio-temporal visual data (videos). I demonstrate that we can convincingly fill spatio-temporal holes in a temporally coherent fashion by jointly reconstructing the appearance and motion. Compared to existing approaches, our technique can synthesize physically plausible contents even in challenging videos. For visual analysis, I apply stereo camera constraints for discovering multiple approximately linear structures in extremely noisy videos with an ecological application to bird migration monitoring at night. The resulting algorithms are simple and intuitive while achieving state-of-the-art performance without the need of training on an exhaustive set of visual examples.\"","abstract_html":"&quot;The past decade has witnessed remarkable progress in image-based, data-driven vision and graphics. However, existing approaches often treat the images as pure 2D signals and not as a 2D projection of the physical 3D world. As a result, a lot of training examples are required to cover sufficiently diverse appearances and inevitably suffer from limited generalization capability. In this thesis, I propose &quot;&quot;inference-by-composition&quot;&quot; approaches to overcome these limitations by modeling and interpreting visual signals in terms of physical surface, object, and scene. I show how we can incorporate physically grounded constraints such as scene-specific geometry in a non-parametric optimization framework for (1) revealing the missing parts of an image due to removal of a foreground or background element, (2) recovering high spatial frequency details that are not resolvable in low-resolution observations. I then extend the framework from 2D images to handle spatio-temporal visual data (videos). I demonstrate that we can convincingly fill spatio-temporal holes in a temporally coherent fashion by jointly reconstructing the appearance and motion. Compared to existing approaches, our technique can synthesize physically plausible contents even in challenging videos. For visual analysis, I apply stereo camera constraints for discovering multiple approximately linear structures in extremely noisy videos with an ecological application to bird migration monitoring at night. The resulting algorithms are simple and intuitive while achieving state-of-the-art performance without the need of training on an exhaustive set of visual examples.&quot;","abstract_has_math":false,"creators":["Huang, Jia-Bin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Ahuja, Narendra","Huang, Thomas S.","Do, Minh N.","Hasegawa-Johnson, Mark","Hoiem, Derek"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-11-10T17:50:06Z","date_published":"2016-11-10T17:50:06Z","updated_at":"2026-07-22T22:26:35Z","subjects":["Computer vision","Visual synthesis","patch-based optimization","image completion","image super-resolution","video completion","visual tracking"],"languages":["en"],"rights":["Copyright 2016 Jia-Bin Huang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/92750","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ahuja, Narendra","Huang, Thomas S.","Do, Minh N.","Hasegawa-Johnson, Mark","Hoiem, Derek"]},{"key":"dc:creator","label":"Author","values":["Huang, Jia-Bin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-11-10T17:50:06Z","2016-07-06","2016-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer vision","Visual synthesis","patch-based optimization","image completion","image super-resolution","video completion","visual tracking"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2016 Jia-Bin Huang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/92750"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"The past decade has witnessed remarkable progress in image-based, data-driven vision and graphics. However, existing approaches often treat the images as pure 2D signals and not as a 2D projection of the physical 3D world. As a result, a lot of training examples are required to cover sufficiently diverse appearances and inevitably suffer from limited generalization capability. In this thesis, I propose \"\"inference-by-composition\"\" approaches to overcome these limitations by modeling and interpreting visual signals in terms of physical surface, object, and scene. I show how we can incorporate physically grounded constraints such as scene-specific geometry in a non-parametric optimization framework for (1) revealing the missing parts of an image due to removal of a foreground or background element, (2) recovering high spatial frequency details that are not resolvable in low-resolution observations. I then extend the framework from 2D images to handle spatio-temporal visual data (videos). I demonstrate that we can convincingly fill spatio-temporal holes in a temporally coherent fashion by jointly reconstructing the appearance and motion. Compared to existing approaches, our technique can synthesize physically plausible contents even in challenging videos. For visual analysis, I apply stereo camera constraints for discovering multiple approximately linear structures in extremely noisy videos with an ecological application to bird migration monitoring at night. The resulting algorithms are simple and intuitive while achieving state-of-the-art performance without the need of training on an exhaustive set of visual examples.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-11-09 without embargo terms","The student, Jia-Bin Huang, accepted the attached license on 2016-07-04 at 02:00.","The student, Jia-Bin Huang, submitted this Dissertation for approval on 2016-07-04 at 02:17.","This Dissertation was approved for publication on 2016-07-06 at 14:09.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9749 on 2016-11-09 at 10:22:39","Made available in DSpace on 2016-11-10T17:50:06Z (GMT). No. of bitstreams: 3 HUANG-DISSERTATION-2016.pdf: 41822301 bytes, checksum: 792c5fe8a31458896c5c2e37d284e871 (MD5) LICENSE.txt: 4210 bytes, checksum: 6d24520218c46f6c63a5d4206401a794 (MD5) PROQUEST_LICENSE.txt: 4556 bytes, checksum: 7476bec9b87a706036712e765c2b7fd7 (MD5) Previous issue date: 2016-07-06"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Visual analysis and synthesis with physically grounded constraints"]}]}],"canonical_facts":{"dc:contributor":["Ahuja, Narendra","Huang, Thomas S.","Do, Minh N.","Hasegawa-Johnson, Mark","Hoiem, Derek"],"dc:creator":["Huang, Jia-Bin"],"dc:date":["2016-11-10T17:50:06Z","2016-07-06","2016-08"],"dc:description":["\"The past decade has witnessed remarkable progress in image-based, data-driven vision and graphics. However, existing approaches often treat the images as pure 2D signals and not as a 2D projection of the physical 3D world. As a result, a lot of training examples are required to cover sufficiently diverse appearances and inevitably suffer from limited generalization capability. In this thesis, I propose \"\"inference-by-composition\"\" approaches to overcome these limitations by modeling and interpreting visual signals in terms of physical surface, object, and scene. I show how we can incorporate physically grounded constraints such as scene-specific geometry in a non-parametric optimization framework for (1) revealing the missing parts of an image due to removal of a foreground or background element, (2) recovering high spatial frequency details that are not resolvable in low-resolution observations. I then extend the framework from 2D images to handle spatio-temporal visual data (videos). I demonstrate that we can convincingly fill spatio-temporal holes in a temporally coherent fashion by jointly reconstructing the appearance and motion. Compared to existing approaches, our technique can synthesize physically plausible contents even in challenging videos. For visual analysis, I apply stereo camera constraints for discovering multiple approximately linear structures in extremely noisy videos with an ecological application to bird migration monitoring at night. The resulting algorithms are simple and intuitive while achieving state-of-the-art performance without the need of training on an exhaustive set of visual examples.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-11-09 without embargo terms","The student, Jia-Bin Huang, accepted the attached license on 2016-07-04 at 02:00.","The student, Jia-Bin Huang, submitted this Dissertation for approval on 2016-07-04 at 02:17.","This Dissertation was approved for publication on 2016-07-06 at 14:09.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9749 on 2016-11-09 at 10:22:39","Made available in DSpace on 2016-11-10T17:50:06Z (GMT). No. of bitstreams: 3 HUANG-DISSERTATION-2016.pdf: 41822301 bytes, checksum: 792c5fe8a31458896c5c2e37d284e871 (MD5) LICENSE.txt: 4210 bytes, checksum: 6d24520218c46f6c63a5d4206401a794 (MD5) PROQUEST_LICENSE.txt: 4556 bytes, checksum: 7476bec9b87a706036712e765c2b7fd7 (MD5) Previous issue date: 2016-07-06"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/92750"],"dc:language":["en"],"dc:rights":["Copyright 2016 Jia-Bin Huang"],"dc:subject":["Computer vision","Visual synthesis","patch-based optimization","image completion","image super-resolution","video completion","visual tracking"],"dc:title":["Visual analysis and synthesis with physically grounded constraints"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:35Z"}