{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/106236"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/106236","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Learning and evaluating image representations","abstract":"Prior to deep learning it was common to approach computer vision problems as describing a model that could be learned from a relatively small amount of data by incorporating domain knowledge. For example, image prediction tasks such as intrinsic image decomposition were approached by thinking about what reflectance and shading look like. In the case of reflectance, a Mondrian image; and in the case of shading, a smooth image. The difficult portion was how to formalize this prior domain knowledge into a model. Deep learning has changed this paradigm. While deep learning hasn’t eliminated the value of domain knowledge, for many problems we now think in terms of model architectures and losses instead. While a choice of model architecture limits the types of results possible, neural networks tend to be less task dependent than domain specific methods. In fact, for almost any problem there is a fairly simple formula for using neural networks to get good results. 1. Collect labeled data, 2. Choose a network architecture, 3. Define a loss and train. However, there are still tasks where we might not be able to collect a lot of labeled data of a particular form (Grave OCR), or tasks where we can’t easily describe an unambiguous loss on easily collected data (Intrinsic Image Decomposition, Image correction including rain, cracks and glare), or a task where we want to do many similar tasks without having to train each one independently (face adjustment). A unifying theme of my work is that generic representations can be learned from data and those learned representation can be used to make otherwise under-constrained problems tractable. Pre- deep learning this generic representation takes the form of a LEARCH-based model more recent work builds on auto-encoder representations. For authoring decompositions and removing rain, cracks, and glare, autoencoder models are learned from fake data and then shown to be applicable on real images. For learning to decompose rainy images cycle consistency losses are incorporated to learn without examples of de-rained images. In Face-to-Face transformation, an attribute sensitive image-to-image representation is pretrained and then a low dimensional representation for image attribute transformations is described. In Grave OCR we learn to generate data and learn the image decomposition model simultaneously, allowing us to learn how to predict image annotations without labeled data. Finally in evaluating intrinsic image decomposition, we explore evaluating intrinsic image models using human perception annotations. We show that human annotation evaluation has some issues and does not appear to differentiate between qualitatively different models. We propose a new task-specific procedure for evaluating intrinsic image decomposition using re- painting and reshading and show that it can be used to identify differences between model that are currently unidentified.","abstract_html":"Prior to deep learning it was common to approach computer vision problems as describing a model that could be learned from a relatively small amount of data by incorporating domain knowledge. For example, image prediction tasks such as intrinsic image decomposition were approached by thinking about what reflectance and shading look like. In the case of reflectance, a Mondrian image; and in the case of shading, a smooth image. The difficult portion was how to formalize this prior domain knowledge into a model. Deep learning has changed this paradigm. While deep learning hasn’t eliminated the value of domain knowledge, for many problems we now think in terms of model architectures and losses instead. While a choice of model architecture limits the types of results possible, neural networks tend to be less task dependent than domain specific methods. In fact, for almost any problem there is a fairly simple formula for using neural networks to get good results. 1. Collect labeled data, 2. Choose a network architecture, 3. Define a loss and train. However, there are still tasks where we might not be able to collect a lot of labeled data of a particular form (Grave OCR), or tasks where we can’t easily describe an unambiguous loss on easily collected data (Intrinsic Image Decomposition, Image correction including rain, cracks and glare), or a task where we want to do many similar tasks without having to train each one independently (face adjustment). A unifying theme of my work is that generic representations can be learned from data and those learned representation can be used to make otherwise under-constrained problems tractable. Pre- deep learning this generic representation takes the form of a LEARCH-based model more recent work builds on auto-encoder representations. For authoring decompositions and removing rain, cracks, and glare, autoencoder models are learned from fake data and then shown to be applicable on real images. For learning to decompose rainy images cycle consistency losses are incorporated to learn without examples of de-rained images. In Face-to-Face transformation, an attribute sensitive image-to-image representation is pretrained and then a low dimensional representation for image attribute transformations is described. In Grave OCR we learn to generate data and learn the image decomposition model simultaneously, allowing us to learn how to predict image annotations without labeled data. Finally in evaluating intrinsic image decomposition, we explore evaluating intrinsic image models using human perception annotations. We show that human annotation evaluation has some issues and does not appear to differentiate between qualitatively different models. We propose a new task-specific procedure for evaluating intrinsic image decomposition using re- painting and reshading and show that it can be used to identify differences between model that are currently unidentified.","abstract_has_math":false,"creators":["Rock, Jason"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Forsyth, David","Lazebnik, Svetlana","Schwing, Alexander G.","Barron, Jonathan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-03-02T21:58:21Z","date_published":"2020-03-02T21:58:21Z","updated_at":"2026-07-22T22:24:45Z","subjects":["Computer Vision","Deep Learning","Image processing","Image representation","intrinsic images","intrinsic image decomposition","rain removal","OCR"],"languages":["en"],"rights":["Copyright 2019 Jason Rock"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/106236","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Forsyth, David","Lazebnik, Svetlana","Schwing, Alexander G.","Barron, Jonathan"]},{"key":"dc:creator","label":"Author","values":["Rock, Jason"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-03-02T21:58:21Z","2019-12-03","2019-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Vision","Deep Learning","Image processing","Image representation","intrinsic images","intrinsic image decomposition","rain removal","OCR"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Jason Rock"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/106236"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Prior to deep learning it was common to approach computer vision problems as describing a model that could be learned from a relatively small amount of data by incorporating domain knowledge. For example, image prediction tasks such as intrinsic image decomposition were approached by thinking about what reflectance and shading look like. In the case of reflectance, a Mondrian image; and in the case of shading, a smooth image. The difficult portion was how to formalize this prior domain knowledge into a model. Deep learning has changed this paradigm. While deep learning hasn’t eliminated the value of domain knowledge, for many problems we now think in terms of model architectures and losses instead. While a choice of model architecture limits the types of results possible, neural networks tend to be less task dependent than domain specific methods. In fact, for almost any problem there is a fairly simple formula for using neural networks to get good results. 1. Collect labeled data, 2. Choose a network architecture, 3. Define a loss and train. However, there are still tasks where we might not be able to collect a lot of labeled data of a particular form (Grave OCR), or tasks where we can’t easily describe an unambiguous loss on easily collected data (Intrinsic Image Decomposition, Image correction including rain, cracks and glare), or a task where we want to do many similar tasks without having to train each one independently (face adjustment). A unifying theme of my work is that generic representations can be learned from data and those learned representation can be used to make otherwise under-constrained problems tractable. Pre- deep learning this generic representation takes the form of a LEARCH-based model more recent work builds on auto-encoder representations. For authoring decompositions and removing rain, cracks, and glare, autoencoder models are learned from fake data and then shown to be applicable on real images. For learning to decompose rainy images cycle consistency losses are incorporated to learn without examples of de-rained images. In Face-to-Face transformation, an attribute sensitive image-to-image representation is pretrained and then a low dimensional representation for image attribute transformations is described. In Grave OCR we learn to generate data and learn the image decomposition model simultaneously, allowing us to learn how to predict image annotations without labeled data. Finally in evaluating intrinsic image decomposition, we explore evaluating intrinsic image models using human perception annotations. We show that human annotation evaluation has some issues and does not appear to differentiate between qualitatively different models. We propose a new task-specific procedure for evaluating intrinsic image decomposition using re- painting and reshading and show that it can be used to identify differences between model that are currently unidentified.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-02-28 without embargo terms","The student, Jason Rock, accepted the attached license on 2019-12-03 at 13:47.","The student, Jason Rock, submitted this Dissertation for approval on 2019-12-03 at 14:13.","This Dissertation was approved for publication on 2019-12-03 at 16:30.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14669 on 2020-02-28 at 17:14:48","Made available in DSpace on 2020-03-02T21:58:21Z (GMT). No. of bitstreams: 3 ROCK-DISSERTATION-2019.pdf: 85579955 bytes, checksum: 4baf7ba805471cd28be7d4487dd112f4 (MD5) LICENSE.txt: 4207 bytes, checksum: 5242ac76766bf66f8d07dd9fe64e8d27 (MD5) PROQUEST_LICENSE.txt: 4553 bytes, checksum: 4e5d4291ba07d8f7dba6d668eb597e9b (MD5) Previous issue date: 2019-12-03"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Learning and evaluating image representations"]}]}],"canonical_facts":{"dc:contributor":["Forsyth, David","Lazebnik, Svetlana","Schwing, Alexander G.","Barron, Jonathan"],"dc:creator":["Rock, Jason"],"dc:date":["2020-03-02T21:58:21Z","2019-12-03","2019-12"],"dc:description":["Prior to deep learning it was common to approach computer vision problems as describing a model that could be learned from a relatively small amount of data by incorporating domain knowledge. For example, image prediction tasks such as intrinsic image decomposition were approached by thinking about what reflectance and shading look like. In the case of reflectance, a Mondrian image; and in the case of shading, a smooth image. The difficult portion was how to formalize this prior domain knowledge into a model. Deep learning has changed this paradigm. While deep learning hasn’t eliminated the value of domain knowledge, for many problems we now think in terms of model architectures and losses instead. While a choice of model architecture limits the types of results possible, neural networks tend to be less task dependent than domain specific methods. In fact, for almost any problem there is a fairly simple formula for using neural networks to get good results. 1. Collect labeled data, 2. Choose a network architecture, 3. Define a loss and train. However, there are still tasks where we might not be able to collect a lot of labeled data of a particular form (Grave OCR), or tasks where we can’t easily describe an unambiguous loss on easily collected data (Intrinsic Image Decomposition, Image correction including rain, cracks and glare), or a task where we want to do many similar tasks without having to train each one independently (face adjustment). A unifying theme of my work is that generic representations can be learned from data and those learned representation can be used to make otherwise under-constrained problems tractable. Pre- deep learning this generic representation takes the form of a LEARCH-based model more recent work builds on auto-encoder representations. For authoring decompositions and removing rain, cracks, and glare, autoencoder models are learned from fake data and then shown to be applicable on real images. For learning to decompose rainy images cycle consistency losses are incorporated to learn without examples of de-rained images. In Face-to-Face transformation, an attribute sensitive image-to-image representation is pretrained and then a low dimensional representation for image attribute transformations is described. In Grave OCR we learn to generate data and learn the image decomposition model simultaneously, allowing us to learn how to predict image annotations without labeled data. Finally in evaluating intrinsic image decomposition, we explore evaluating intrinsic image models using human perception annotations. We show that human annotation evaluation has some issues and does not appear to differentiate between qualitatively different models. We propose a new task-specific procedure for evaluating intrinsic image decomposition using re- painting and reshading and show that it can be used to identify differences between model that are currently unidentified.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-02-28 without embargo terms","The student, Jason Rock, accepted the attached license on 2019-12-03 at 13:47.","The student, Jason Rock, submitted this Dissertation for approval on 2019-12-03 at 14:13.","This Dissertation was approved for publication on 2019-12-03 at 16:30.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14669 on 2020-02-28 at 17:14:48","Made available in DSpace on 2020-03-02T21:58:21Z (GMT). No. of bitstreams: 3 ROCK-DISSERTATION-2019.pdf: 85579955 bytes, checksum: 4baf7ba805471cd28be7d4487dd112f4 (MD5) LICENSE.txt: 4207 bytes, checksum: 5242ac76766bf66f8d07dd9fe64e8d27 (MD5) PROQUEST_LICENSE.txt: 4553 bytes, checksum: 4e5d4291ba07d8f7dba6d668eb597e9b (MD5) Previous issue date: 2019-12-03"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/106236"],"dc:language":["en"],"dc:rights":["Copyright 2019 Jason Rock"],"dc:subject":["Computer Vision","Deep Learning","Image processing","Image representation","intrinsic images","intrinsic image decomposition","rain removal","OCR"],"dc:title":["Learning and evaluating image representations"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:45Z"}