{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/108603"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/108603","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Vision-based 6D object pose estimation for robot manipulation","abstract":"Vision-based 6D object pose estimation focuses on estimating the 3D translation and 3D orientation of an object with respect to the camera. Accurately estimating the 6D object pose plays a crucial role in various robotic applications such as robot manipulation and semantic navigation. In this dissertation, we study the problem of 6D object pose estimation and its application to manipulation. We first introduce PoseRBPF, a Rao-Blackwellized particle filter for tracking 6D object poses. In the framework, each particle samples 3D translation and estimates the distribution over 3D rotations conditioned on the image bonding box corresponding to the sampled translation. \\prbpf\\ compares each bounding box embedding to learned viewpoint embeddings so as to efficiently update distributions over time. We demonstrate that the tracked distributions capture both the uncertainties from the symmetry of objects and the uncertainty from object pose with RGB or RGB-D measurements. We propose a category-level extension of the PoseRBPF framework which effectively estimates the 6D poses and sizes of unseen objects. In particular, we propose a category-level auto-encoder network for depth measurements so that the feature embeddings are independent of the object instances. We extend the states in the PoseRBPF to handle the objects in different sizes. We evaluate our tracking framework on a category-level pose estimation benchmark, and achieve state-of-the-art performance. We introduce a robot system for self-supervised 6D object pose estimation. Starting from modules trained in simulation, our system is able to label real world images with accurate 6D object poses for self-supervised learning. In addition, the robot interacts with objects in the environment to change the object configuration by grasping or pushing objects. In this way, our system is able to continuously collect data and improve its pose estimation modules. We show that the self-supervised learning improves object segmentation and 6D pose estimation performance, and consequently enables the system to grasp objects more reliably.","abstract_html":"Vision-based 6D object pose estimation focuses on estimating the 3D translation and 3D orientation of an object with respect to the camera. Accurately estimating the 6D object pose plays a crucial role in various robotic applications such as robot manipulation and semantic navigation. In this dissertation, we study the problem of 6D object pose estimation and its application to manipulation. We first introduce PoseRBPF, a Rao-Blackwellized particle filter for tracking 6D object poses. In the framework, each particle samples 3D translation and estimates the distribution over 3D rotations conditioned on the image bonding box corresponding to the sampled translation. \\prbpf\\ compares each bounding box embedding to learned viewpoint embeddings so as to efficiently update distributions over time. We demonstrate that the tracked distributions capture both the uncertainties from the symmetry of objects and the uncertainty from object pose with RGB or RGB-D measurements. We propose a category-level extension of the PoseRBPF framework which effectively estimates the 6D poses and sizes of unseen objects. In particular, we propose a category-level auto-encoder network for depth measurements so that the feature embeddings are independent of the object instances. We extend the states in the PoseRBPF to handle the objects in different sizes. We evaluate our tracking framework on a category-level pose estimation benchmark, and achieve state-of-the-art performance. We introduce a robot system for self-supervised 6D object pose estimation. Starting from modules trained in simulation, our system is able to label real world images with accurate 6D object poses for self-supervised learning. In addition, the robot interacts with objects in the environment to change the object configuration by grasping or pushing objects. In this way, our system is able to continuously collect data and improve its pose estimation modules. We show that the self-supervised learning improves object segmentation and 6D pose estimation performance, and consequently enables the system to grasp objects more reliably.","abstract_has_math":false,"creators":["Deng, Xinke"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Bretl, Timothy Wolfe","Do, Minh","Fox, Dieter","Gupta, Saurabh","Hu, Bin"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-10-07T22:44:34Z","date_published":"2020-10-07T22:44:34Z","updated_at":"2026-07-22T22:24:48Z","subjects":["Robotics","State estimation","Computer vision","6D object pose estimation"],"languages":["en"],"rights":["Copyright 2020 Xinke Deng"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/108603","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Bretl, Timothy Wolfe","Do, Minh","Fox, Dieter","Gupta, Saurabh","Hu, Bin"]},{"key":"dc:creator","label":"Author","values":["Deng, Xinke"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-10-07T22:44:34Z","2022-10-07T22:44:53Z","2020-07-14","2020-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Robotics","State estimation","Computer vision","6D object pose estimation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Xinke Deng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/108603"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Vision-based 6D object pose estimation focuses on estimating the 3D translation and 3D orientation of an object with respect to the camera. Accurately estimating the 6D object pose plays a crucial role in various robotic applications such as robot manipulation and semantic navigation. In this dissertation, we study the problem of 6D object pose estimation and its application to manipulation. We first introduce PoseRBPF, a Rao-Blackwellized particle filter for tracking 6D object poses. In the framework, each particle samples 3D translation and estimates the distribution over 3D rotations conditioned on the image bonding box corresponding to the sampled translation. \\prbpf\\ compares each bounding box embedding to learned viewpoint embeddings so as to efficiently update distributions over time. We demonstrate that the tracked distributions capture both the uncertainties from the symmetry of objects and the uncertainty from object pose with RGB or RGB-D measurements. We propose a category-level extension of the PoseRBPF framework which effectively estimates the 6D poses and sizes of unseen objects. In particular, we propose a category-level auto-encoder network for depth measurements so that the feature embeddings are independent of the object instances. We extend the states in the PoseRBPF to handle the objects in different sizes. We evaluate our tracking framework on a category-level pose estimation benchmark, and achieve state-of-the-art performance. We introduce a robot system for self-supervised 6D object pose estimation. Starting from modules trained in simulation, our system is able to label real world images with accurate 6D object poses for self-supervised learning. In addition, the robot interacts with objects in the environment to change the object configuration by grasping or pushing objects. In this way, our system is able to continuously collect data and improve its pose estimation modules. We show that the self-supervised learning improves object segmentation and 6D pose estimation performance, and consequently enables the system to grasp objects more reliably.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2022-08-01","The student, Xinke Deng, accepted the attached license on 2020-07-13 at 17:23.","The student, Xinke Deng, submitted this Dissertation for approval on 2020-07-13 at 17:42.","This Dissertation was approved for publication on 2020-07-14 at 12:10.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15595 on 2020-10-02 at 15:32:59","Made available in DSpace on 2020-10-07T22:44:34Z (GMT). No. of bitstreams: 2 DENG-DISSERTATION-2020.pdf: 40870750 bytes, checksum: df9b119878f981483ff49906e6a5bf37 (MD5) LICENSE.txt: 4207 bytes, checksum: 24fac9d96529c680a0acf667ab920315 (MD5) Previous issue date: 2020-07-14","Embargo set by: Seth Robbins for item 116230 Lift date: 2022-10-07T22:44:53Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Vision-based 6D object pose estimation for robot manipulation"]}]}],"canonical_facts":{"dc:contributor":["Bretl, Timothy Wolfe","Do, Minh","Fox, Dieter","Gupta, Saurabh","Hu, Bin"],"dc:creator":["Deng, Xinke"],"dc:date":["2020-10-07T22:44:34Z","2022-10-07T22:44:53Z","2020-07-14","2020-08"],"dc:description":["Vision-based 6D object pose estimation focuses on estimating the 3D translation and 3D orientation of an object with respect to the camera. Accurately estimating the 6D object pose plays a crucial role in various robotic applications such as robot manipulation and semantic navigation. In this dissertation, we study the problem of 6D object pose estimation and its application to manipulation. We first introduce PoseRBPF, a Rao-Blackwellized particle filter for tracking 6D object poses. In the framework, each particle samples 3D translation and estimates the distribution over 3D rotations conditioned on the image bonding box corresponding to the sampled translation. \\prbpf\\ compares each bounding box embedding to learned viewpoint embeddings so as to efficiently update distributions over time. We demonstrate that the tracked distributions capture both the uncertainties from the symmetry of objects and the uncertainty from object pose with RGB or RGB-D measurements. We propose a category-level extension of the PoseRBPF framework which effectively estimates the 6D poses and sizes of unseen objects. In particular, we propose a category-level auto-encoder network for depth measurements so that the feature embeddings are independent of the object instances. We extend the states in the PoseRBPF to handle the objects in different sizes. We evaluate our tracking framework on a category-level pose estimation benchmark, and achieve state-of-the-art performance. We introduce a robot system for self-supervised 6D object pose estimation. Starting from modules trained in simulation, our system is able to label real world images with accurate 6D object poses for self-supervised learning. In addition, the robot interacts with objects in the environment to change the object configuration by grasping or pushing objects. In this way, our system is able to continuously collect data and improve its pose estimation modules. We show that the self-supervised learning improves object segmentation and 6D pose estimation performance, and consequently enables the system to grasp objects more reliably.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2022-08-01","The student, Xinke Deng, accepted the attached license on 2020-07-13 at 17:23.","The student, Xinke Deng, submitted this Dissertation for approval on 2020-07-13 at 17:42.","This Dissertation was approved for publication on 2020-07-14 at 12:10.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15595 on 2020-10-02 at 15:32:59","Made available in DSpace on 2020-10-07T22:44:34Z (GMT). No. of bitstreams: 2 DENG-DISSERTATION-2020.pdf: 40870750 bytes, checksum: df9b119878f981483ff49906e6a5bf37 (MD5) LICENSE.txt: 4207 bytes, checksum: 24fac9d96529c680a0acf667ab920315 (MD5) Previous issue date: 2020-07-14","Embargo set by: Seth Robbins for item 116230 Lift date: 2022-10-07T22:44:53Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/108603"],"dc:language":["en"],"dc:rights":["Copyright 2020 Xinke Deng"],"dc:subject":["Robotics","State estimation","Computer vision","6D object pose estimation"],"dc:title":["Vision-based 6D object pose estimation for robot manipulation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:48Z"}