{"id":{"repo_id":"queens","oai_identifier":"oai:queensu.scholaris.ca:1974/34297"},"canonical_url":"https://search.dev.ndltd.org/etd/queens/oai:queensu.scholaris.ca:1974/34297","repository":{"repo_id":"queens","name":"Queens University","base_url":"https://qspace.library.queensu.ca/server/oai/request"},"display":{"title":"Advancing 6DoF Object Pose Estimation: Keypoint Voting, Optimal Keypoint Sampling, and Bridging the Simulation-to-real Gap","abstract":"This thesis explores advancements in six-degree-of-freedom (6DoF) object pose estimation using point cloud and RGB-D data. Three novel methodologies are proposed to address distinct challenges in this field. First, RCVPose3D introduces a cascaded keypoint voting framework that separates semantic segmentation from keypoint regression, incorporating pairwise constraints and a Voter Confident Score to improve accuracy. RCVPose3D achieves state-of-the-art results on Occlusion LINEMOD (74.5%) and YCB-Video (96.9%), outperforming traditional RGB and RGB-D methods. Second, KeyGNet leverages a graph network to optimize the keypoint selection, improving accuracy and efficiency by learning dispersed, evenly distributed keypoints. KeyGNet enhances performance across all metrics, notably increasing ADD(S) on Occlusion LINEMOD by 16.4% and closing the single-to-multiobject training gap. Finally, RKHSPose addresses the simulation-to-real domain gap through a self-supervised framework using a learnable kernel in RKHS and an adapter network pre-trained on synthetic data. This approach achieves competitive results against fully supervised methods, requiring no real groundtruth annotations. Together, these contributions advance 6DoF pose estimation in accuracy, efficiency, and adaptability across diverse datasets and scenarios.","abstract_html":"This thesis explores advancements in six-degree-of-freedom (6DoF) object pose estimation using point cloud and RGB-D data. Three novel methodologies are proposed to address distinct challenges in this field. First, RCVPose3D introduces a cascaded keypoint voting framework that separates semantic segmentation from keypoint regression, incorporating pairwise constraints and a Voter Confident Score to improve accuracy. RCVPose3D achieves state-of-the-art results on Occlusion LINEMOD (74.5%) and YCB-Video (96.9%), outperforming traditional RGB and RGB-D methods. Second, KeyGNet leverages a graph network to optimize the keypoint selection, improving accuracy and efficiency by learning dispersed, evenly distributed keypoints. KeyGNet enhances performance across all metrics, notably increasing ADD(S) on Occlusion LINEMOD by 16.4% and closing the single-to-multiobject training gap. Finally, RKHSPose addresses the simulation-to-real domain gap through a self-supervised framework using a learnable kernel in RKHS and an adapter network pre-trained on synthetic data. This approach achieves competitive results against fully supervised methods, requiring no real groundtruth annotations. Together, these contributions advance 6DoF pose estimation in accuracy, efficiency, and adaptability across diverse datasets and scenarios.","abstract_has_math":false,"creators":["Wu, Yangzheng"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Electrical and Computer Engineering","school":null,"contributors":[],"advisors":["Greenspan, Michael"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-01-30","date_published":"2025-01-30","updated_at":"2026-07-27T20:35:41Z","subjects":["6DoF pose"],"languages":["eng"],"rights":["Attribution 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1974/34297","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.department","label":"Department","values":["Electrical and Computer Engineering"]},{"key":"dc:contributor.supervisor","label":"Supervisor","values":["Greenspan, Michael"]},{"key":"dc:creator","label":"Author","values":["Wu, Yangzheng"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-01-30T14:11:51Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-01-30T14:11:51Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-01-30"]},{"key":"dc:type","label":"Dc Type","values":["thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["6DoF pose"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Attribution 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1974/34297"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This thesis explores advancements in six-degree-of-freedom (6DoF) object pose estimation using point cloud and RGB-D data. Three novel methodologies are proposed to address distinct challenges in this field. First, RCVPose3D introduces a cascaded keypoint voting framework that separates semantic segmentation from keypoint regression, incorporating pairwise constraints and a Voter Confident Score to improve accuracy. RCVPose3D achieves state-of-the-art results on Occlusion LINEMOD (74.5%) and YCB-Video (96.9%), outperforming traditional RGB and RGB-D methods. Second, KeyGNet leverages a graph network to optimize the keypoint selection, improving accuracy and efficiency by learning dispersed, evenly distributed keypoints. KeyGNet enhances performance across all metrics, notably increasing ADD(S) on Occlusion LINEMOD by 16.4% and closing the single-to-multiobject training gap. Finally, RKHSPose addresses the simulation-to-real domain gap through a self-supervised framework using a learnable kernel in RKHS and an adapter network pre-trained on synthetic data. This approach achieves competitive results against fully supervised methods, requiring no real groundtruth annotations. Together, these contributions advance 6DoF pose estimation in accuracy, efficiency, and adaptability across diverse datasets and scenarios."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["PhD"]},{"key":"dc:title","label":"Title","values":["Advancing 6DoF Object Pose Estimation: Keypoint Voting, Optimal Keypoint Sampling, and Bridging the Simulation-to-real Gap"]}]}],"canonical_facts":{"dc:contributor.department":["Electrical and Computer Engineering"],"dc:contributor.supervisor":["Greenspan, Michael"],"dc:creator":["Wu, Yangzheng"],"dc:date.accessioned":["2025-01-30T14:11:51Z"],"dc:date.available":["2025-01-30T14:11:51Z"],"dc:date.issued":["2025-01-30"],"dc:description.abstract":["This thesis explores advancements in six-degree-of-freedom (6DoF) object pose estimation using point cloud and RGB-D data. Three novel methodologies are proposed to address distinct challenges in this field. First, RCVPose3D introduces a cascaded keypoint voting framework that separates semantic segmentation from keypoint regression, incorporating pairwise constraints and a Voter Confident Score to improve accuracy. RCVPose3D achieves state-of-the-art results on Occlusion LINEMOD (74.5%) and YCB-Video (96.9%), outperforming traditional RGB and RGB-D methods. Second, KeyGNet leverages a graph network to optimize the keypoint selection, improving accuracy and efficiency by learning dispersed, evenly distributed keypoints. KeyGNet enhances performance across all metrics, notably increasing ADD(S) on Occlusion LINEMOD by 16.4% and closing the single-to-multiobject training gap. Finally, RKHSPose addresses the simulation-to-real domain gap through a self-supervised framework using a learnable kernel in RKHS and an adapter network pre-trained on synthetic data. This approach achieves competitive results against fully supervised methods, requiring no real groundtruth annotations. Together, these contributions advance 6DoF pose estimation in accuracy, efficiency, and adaptability across diverse datasets and scenarios."],"dc:description.degree":["PhD"],"dc:identifier.uri":["https://hdl.handle.net/1974/34297"],"dc:language.iso":["eng"],"dc:rights":["Attribution 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by/4.0/"],"dc:subject":["6DoF pose"],"dc:title":["Advancing 6DoF Object Pose Estimation: Keypoint Voting, Optimal Keypoint Sampling, and Bridging the Simulation-to-real Gap"],"dc:type":["thesis"]},"updated_at":"2026-07-27T20:35:41Z"}