City University of New York - City College
Utilizing Deep Learning Audio Models for Blind and Low Vision Crosswalk Assistance
Abstract
dc:description.abstract<p>Navigating urban environments poses significant challenges for blind and low vision (BLV) individuals, particularly at street intersections where determining when it is safe to cross can be life-threatening. In New York City, where pedestrian fatalities are on the rise and only 2% of intersections are equipped with Accessible Pedestrian Signals (APS), alternative solutions are urgently needed. This thesis proposes an audio-based deep learning approach to support BLV individuals at crosswalks by detecting traffic movement direction and idling states using spatial sound. With 4-channel audio capturing capabilities of wearables, such as Meta Project Aria glasses, we explore state-of-the-art sound event localization and detection (SELD) models, specifically ResNet-based feature extractors and transformer-enhanced architectures such as EINV2. To train these models, a synthetic dataset of quadraphonic audio from simulated 4-way traffic scenes is generated using Unity Engine, addressing both the scalability and privacy concerns of real-world data collection. This work begins an exploration into a hands-free, audio-centric crosswalk navigation aid for BLV individuals.</p>
Degree
thesis:*- Name thesis:degree_name
- Master of Science (M.S.)
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Engineering
- Year dc:date.available
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Lam, Wayne
- Contributors dc:contributor
-
- Hao Tang
- Zhigang Zhu
Subjects
dc:subject × 5Identifiers
dc:identifier.*- Repository record dc:identifier
- https://academicworks.cuny.edu/cc_etds_theses/1273
- OAI identifier oai:identifier
- oai:academicworks.cuny.edu:cc_etds_theses-2325