University of Cambridge
Place Recognition Anywhere Model for Efficient Large-scale Localization
Abstract
dc:description.abstractVisual localization aims to estimate the orientation and position of a given query image in a known environment. It is a fundamental task in computer vision and also a key technique of various applications including autonomous driving, robotics, and virtual/augmented reality (AR/VR). Previous visual localization methods can be roughly categorized as absolute pose regression (APR) [63, 62, 178, 169, 165], scene coordinate regression (SCR) [48, 20, 13, 17, 12], and the hierarchical method (HM) [127, 122, 157]. However, none of APRs, SCRs, and HMs can satisfy both efficiency and accuracy in large-scale environments. In this thesis, we propose the place recognition anywhere model (PRAM) for efficient and accurate visual localization. Specifically, (1) instead of defining landmarks on classic semantic labels, e.g., buildings, PRAM generates landmarks in 3D space automatically in a self-supervised manner, making any place a landmark and overcoming the low-generalization ability of commonally used semantic labels. (2) PRAM performs sparse recognition directly with transformers by taking sparse keypoints as input, resulting in efficient sparse recognition of a large number of landmarks. (3) PRAM discards global descriptors and repetitive 2D descriptors and establishes 2D-3D matches with the guidance of predicted landmark labels, so PRAM has higher memory and time efficiency. To increase the global representation ability of sparse keypoints, we introduce the work of semantic-guided feature detection and description (SFD2) with implicitly embedded semantics. Moreover, to reduce the 2D-3D matching cost, we propose the algorithm of iterative matching and pose estimation with adaptive pooling (IMP) to augment the matching quality and efficiency with geometric constraints. Both SFD2 and IMP are incorporated into PRAM to make it an efficient and accurate localization system. Our experiments on public indoor and outdoor datasets demonstrate that PRAM reduces over 90% storage of HMs and works about 2.4 times faster than HMs. Moreover, PRAM opens new directions for visual localization including localization with multi-modality signals, compact map representation, and hierarchical scene coordinate regression. These directions are also discussed at the end of this thesis as future works.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Xue, Fei
- Advisor dc:contributor.advisor
-
- Cipolla, Roberto
Subjects
dc:subject × 3Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.118056
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/383831