{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105082"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105082","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Depth aware RCNN","abstract":"Image object detection networks that depend on region proposal networks (RPN) have achieved state-of-art results. As RPN is trained to share convolutional features with the actual classification layers in the network, features learned by the convolutional backbones may have subtle impact on the RPN. A successful approach comes from RGB-D image object detection, where the convolutional layers learn not just RGB features, but also depth features. In this thesis, we study the problem of simultaneously localizing objects as well as estimating their depth. We propose to use one backbone network for two tasks and show that multi-task learning with shared weights can have reciprocating benefits. Our experiments show that when combined with depth prediction in the network, the object detection branch in our model outperforms Faster-RCNN on the challenging KITTI detection benchmark and the Cityscapes dataset. Likewise, the performance of our depth prediction branch is slightly better compared with methods using the same depth prediction architecture.","abstract_html":"Image object detection networks that depend on region proposal networks (RPN) have achieved state-of-art results. As RPN is trained to share convolutional features with the actual classification layers in the network, features learned by the convolutional backbones may have subtle impact on the RPN. A successful approach comes from RGB-D image object detection, where the convolutional layers learn not just RGB features, but also depth features. In this thesis, we study the problem of simultaneously localizing objects as well as estimating their depth. We propose to use one backbone network for two tasks and show that multi-task learning with shared weights can have reciprocating benefits. Our experiments show that when combined with depth prediction in the network, the object detection branch in our model outperforms Faster-RCNN on the challenging KITTI detection benchmark and the Cityscapes dataset. Likewise, the performance of our depth prediction branch is slightly better compared with methods using the same depth prediction architecture.","abstract_has_math":false,"creators":["Zhao, Tianxi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Shi, Honghui"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T20:36:08Z","date_published":"2019-08-23T20:36:08Z","updated_at":"2026-07-22T22:24:44Z","subjects":["Object detection","Depth estimation","Computer Vision","RCNN"],"languages":["en"],"rights":["Copyright 2019 Tianxi Zhao"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105082","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Shi, Honghui"]},{"key":"dc:creator","label":"Author","values":["Zhao, Tianxi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T20:36:08Z","2021-08-24T09:15:31Z","2019-04-24","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Object detection","Depth estimation","Computer Vision","RCNN"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Tianxi Zhao"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105082"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Image object detection networks that depend on region proposal networks (RPN) have achieved state-of-art results. As RPN is trained to share convolutional features with the actual classification layers in the network, features learned by the convolutional backbones may have subtle impact on the RPN. A successful approach comes from RGB-D image object detection, where the convolutional layers learn not just RGB features, but also depth features. In this thesis, we study the problem of simultaneously localizing objects as well as estimating their depth. We propose to use one backbone network for two tasks and show that multi-task learning with shared weights can have reciprocating benefits. Our experiments show that when combined with depth prediction in the network, the object detection branch in our model outperforms Faster-RCNN on the challenging KITTI detection benchmark and the Cityscapes dataset. Likewise, the performance of our depth prediction branch is slightly better compared with methods using the same depth prediction architecture.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-05-01","The student, Tianxi Zhao, accepted the attached license on 2019-04-24 at 08:05.","The student, Tianxi Zhao, submitted this Thesis for approval on 2019-04-24 at 08:06.","This Thesis was approved for publication on 2019-04-24 at 09:28.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13865 on 2019-08-22 at 15:07:58","Made available in DSpace on 2019-08-23T20:36:08Z (GMT). No. of bitstreams: 2 ZHAO-THESIS-2019.pdf: 12152227 bytes, checksum: 8f632c204c224a70a8ada07d01502b8f (MD5) LICENSE.txt: 4208 bytes, checksum: 050255cdbf9658c31eae4895de3316a5 (MD5) Previous issue date: 2019-04-24","Embargo set by: Seth Robbins for item 112201 Lift date: 2021-08-23T20:36:18Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 112201 on 2021-08-24T09:15:31Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Depth aware RCNN"]}]}],"canonical_facts":{"dc:contributor":["Shi, Honghui"],"dc:creator":["Zhao, Tianxi"],"dc:date":["2019-08-23T20:36:08Z","2021-08-24T09:15:31Z","2019-04-24","2019-05"],"dc:description":["Image object detection networks that depend on region proposal networks (RPN) have achieved state-of-art results. As RPN is trained to share convolutional features with the actual classification layers in the network, features learned by the convolutional backbones may have subtle impact on the RPN. A successful approach comes from RGB-D image object detection, where the convolutional layers learn not just RGB features, but also depth features. In this thesis, we study the problem of simultaneously localizing objects as well as estimating their depth. We propose to use one backbone network for two tasks and show that multi-task learning with shared weights can have reciprocating benefits. Our experiments show that when combined with depth prediction in the network, the object detection branch in our model outperforms Faster-RCNN on the challenging KITTI detection benchmark and the Cityscapes dataset. Likewise, the performance of our depth prediction branch is slightly better compared with methods using the same depth prediction architecture.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-05-01","The student, Tianxi Zhao, accepted the attached license on 2019-04-24 at 08:05.","The student, Tianxi Zhao, submitted this Thesis for approval on 2019-04-24 at 08:06.","This Thesis was approved for publication on 2019-04-24 at 09:28.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13865 on 2019-08-22 at 15:07:58","Made available in DSpace on 2019-08-23T20:36:08Z (GMT). No. of bitstreams: 2 ZHAO-THESIS-2019.pdf: 12152227 bytes, checksum: 8f632c204c224a70a8ada07d01502b8f (MD5) LICENSE.txt: 4208 bytes, checksum: 050255cdbf9658c31eae4895de3316a5 (MD5) Previous issue date: 2019-04-24","Embargo set by: Seth Robbins for item 112201 Lift date: 2021-08-23T20:36:18Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 112201 on 2021-08-24T09:15:31Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105082"],"dc:language":["en"],"dc:rights":["Copyright 2019 Tianxi Zhao"],"dc:subject":["Object detection","Depth estimation","Computer Vision","RCNN"],"dc:title":["Depth aware RCNN"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:44Z"}