University of Ontario Institute of Technology
Attention-enhanced Cross-View Transformer for monocular BEV perception
Abstract
dc:description.abstractAccurately understanding road scenes from monocular images is essential for autonomous driving, yet generating reliable bird's-eye view (BEV) layouts from single front-view inputs remains challenging. This thesis presents an enhanced cross-view learning framework building upon the PYVA architecture, introducing a Convolutional Block Attention Module (CBAM) to improve feature representation. Our method incorporates spatial-channel attention to refine encoder features, followed by Cycled View Projection (CVP) and a Cross-View Transformer (CVT) for view transformation. CVP enforces geometric coherence through cycle-consistency constraints, while CVT explicitly models feature correspondence across views. We evaluate our approach on KITTI and Argoverse 1.1 datasets for both static and dynamic tasks. Results demonstrate superior performance compared to MonoLayout and PYVA, particularly in mean Average Precision (mAP). The model effectively detects small and distant vehicles while preserving fine road structures in complex scenes.
Degree
thesis:*- Name thesis:degree_name
- Master of Applied Science (MASc)
- Discipline thesis:degree_discipline
- Mechanical Engineering
- Grantor
- University of Ontario Institute of Technology
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Jiang, Yu
- Advisor dc:contributor.advisor
-
- Lang, Haoxiang
Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/10155/2049
- OAI identifier oai:identifier
- oai:ontariotechu.scholaris.ca:10155/2049