{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124511"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124511","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards bridging generative and discriminative learning for visual perception","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2026-05-01","abstract_has_math":false,"creators":["Zheng, Shuhong"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Wang, Yuxiong"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:02Z","subjects":["Generative Models","Visual Perception"],"languages":["en","eng"],"rights":["Copyright 2024 Shuhong Zheng"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124511","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Wang, Yuxiong"]},{"key":"dc:creator","label":"Author","values":["Zheng, Shuhong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-04-29"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Generative Models","Visual Perception"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Shuhong Zheng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124511"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","The student, Shuhong Zheng, accepted the attached license on 2024-04-22 at 22:27.","The student, Shuhong Zheng, submitted this Thesis for approval on 2024-04-22 at 22:43.","This Thesis was approved for publication on 2024-04-29 at 08:55.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20334 on 2024-09-16 at 00:43:18","Generative models have taken the prevalence in artificial intelligence (AI) these days, as people get increasingly astonished by their impressive generative capability. However, most people solely focus on their applications of synthesizing creative content for entertainment purposes. With a different perspective, this thesis strives to explore the possibility of bridging generative and discriminative learning for visual perception tasks. Our first attempt is to leverage the neural radiance fields (NeRF) for visual perception which enables 3D awareness for scene understanding. Previous discriminative visual perception models only take in a single-view observation as input for visual attribute prediction, which overlook the underlying 3D information within the scene. To overcome this limitation, we propose a novel task called multi-task view synthesis (MTVS), which evaluates comprehensive visual perception ability of 3D scenes. We also propose an innovative framework called Multi-task and cross-view NeRF (MuvieNeRF) equipped with both multi-task and cross-view reasoning capability to solve this challenging problem. We demonstrate that the joint modeling of multi-view information and multi-modal visual attributes could benefit the performance on the challenging MTVS task. As building a NeRF representation requires multi-view observations of the scenes, which limits the real application, we move a step further to leverage diffusion models for visual perception. Diffusion models are pretrained on large-scale datasets which empower them with informative feature representations for visual input, and consequently with incredible ability of solving discriminative tasks. We propose to incorporate the discriminative capability together with their inherent generative power into one unified model for better bridging the two perspectives. Our presented Self-improving UNified Diffusion (SUNDiff) achieves this by having two sets of parameters within a single model for data generation and exploitation. We show that the unified diffusion model design brings performance gain to various types of visual perception tasks and is beneficial for diverse model architectures. Looking ahead, as generative models keep fast evolving with increasingly impressive generation capability, it is worthwhile to exploit the discriminative potential within these more powerful generative models. Meanwhile, the regime of utilizing generative models for discriminative learning can be further extended to more diverse tasks and modalities as perception and understanding on videos, point clouds, etc, providing broader benefit to different subfields of computer vision."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards bridging generative and discriminative learning for visual perception"]}]}],"canonical_facts":{"dc:contributor":["Wang, Yuxiong"],"dc:creator":["Zheng, Shuhong"],"dc:date":["2024-05","2024-04-29"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","The student, Shuhong Zheng, accepted the attached license on 2024-04-22 at 22:27.","The student, Shuhong Zheng, submitted this Thesis for approval on 2024-04-22 at 22:43.","This Thesis was approved for publication on 2024-04-29 at 08:55.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20334 on 2024-09-16 at 00:43:18","Generative models have taken the prevalence in artificial intelligence (AI) these days, as people get increasingly astonished by their impressive generative capability. However, most people solely focus on their applications of synthesizing creative content for entertainment purposes. With a different perspective, this thesis strives to explore the possibility of bridging generative and discriminative learning for visual perception tasks. Our first attempt is to leverage the neural radiance fields (NeRF) for visual perception which enables 3D awareness for scene understanding. Previous discriminative visual perception models only take in a single-view observation as input for visual attribute prediction, which overlook the underlying 3D information within the scene. To overcome this limitation, we propose a novel task called multi-task view synthesis (MTVS), which evaluates comprehensive visual perception ability of 3D scenes. We also propose an innovative framework called Multi-task and cross-view NeRF (MuvieNeRF) equipped with both multi-task and cross-view reasoning capability to solve this challenging problem. We demonstrate that the joint modeling of multi-view information and multi-modal visual attributes could benefit the performance on the challenging MTVS task. As building a NeRF representation requires multi-view observations of the scenes, which limits the real application, we move a step further to leverage diffusion models for visual perception. Diffusion models are pretrained on large-scale datasets which empower them with informative feature representations for visual input, and consequently with incredible ability of solving discriminative tasks. We propose to incorporate the discriminative capability together with their inherent generative power into one unified model for better bridging the two perspectives. Our presented Self-improving UNified Diffusion (SUNDiff) achieves this by having two sets of parameters within a single model for data generation and exploitation. We show that the unified diffusion model design brings performance gain to various types of visual perception tasks and is beneficial for diverse model architectures. Looking ahead, as generative models keep fast evolving with increasingly impressive generation capability, it is worthwhile to exploit the discriminative potential within these more powerful generative models. Meanwhile, the regime of utilizing generative models for discriminative learning can be further extended to more diverse tasks and modalities as perception and understanding on videos, point clouds, etc, providing broader benefit to different subfields of computer vision."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124511"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Shuhong Zheng"],"dc:subject":["Generative Models","Visual Perception"],"dc:title":["Towards bridging generative and discriminative learning for visual perception"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}