{"id":{"repo_id":"cuny","oai_identifier":"oai:academicworks.cuny.edu:cc_etds_theses-2325"},"canonical_url":"https://search.dev.ndltd.org/etd/cuny/oai:academicworks.cuny.edu:cc_etds_theses-2325","repository":{"repo_id":"cuny","name":"City University of New York - City College","base_url":"https://academicworks.cuny.edu/do/oai/"},"display":{"title":"Utilizing Deep Learning Audio Models for Blind and Low Vision Crosswalk Assistance","abstract":"<p>Navigating urban environments poses significant challenges for blind and low vision (BLV) individuals, particularly at street intersections where determining when it is safe to cross can be life-threatening. In New York City, where pedestrian fatalities are on the rise and only 2% of intersections are equipped with Accessible Pedestrian Signals (APS), alternative solutions are urgently needed. This thesis proposes an audio-based deep learning approach to support BLV individuals at crosswalks by detecting traffic movement direction and idling states using spatial sound. With 4-channel audio capturing capabilities of wearables, such as Meta Project Aria glasses, we explore state-of-the-art sound event localization and detection (SELD) models, specifically ResNet-based feature extractors and transformer-enhanced architectures such as EINV2. To train these models, a synthetic dataset of quadraphonic audio from simulated 4-way traffic scenes is generated using Unity Engine, addressing both the scalability and privacy concerns of real-world data collection. This work begins an exploration into a hands-free, audio-centric crosswalk navigation aid for BLV individuals.</p>","abstract_html":"&lt;p&gt;Navigating urban environments poses significant challenges for blind and low vision (BLV) individuals, particularly at street intersections where determining when it is safe to cross can be life-threatening. In New York City, where pedestrian fatalities are on the rise and only 2% of intersections are equipped with Accessible Pedestrian Signals (APS), alternative solutions are urgently needed. This thesis proposes an audio-based deep learning approach to support BLV individuals at crosswalks by detecting traffic movement direction and idling states using spatial sound. With 4-channel audio capturing capabilities of wearables, such as Meta Project Aria glasses, we explore state-of-the-art sound event localization and detection (SELD) models, specifically ResNet-based feature extractors and transformer-enhanced architectures such as EINV2. To train these models, a synthetic dataset of quadraphonic audio from simulated 4-way traffic scenes is generated using Unity Engine, addressing both the scalability and privacy concerns of real-world data collection. This work begins an exploration into a hands-free, audio-centric crosswalk navigation aid for BLV individuals.&lt;/p&gt;","abstract_has_math":false,"creators":["Lam, Wayne"],"institution":null,"degree_name":"Master of Science (M.S.)","degree_level":"Thesis","degree_discipline":"Engineering","degree_department":null,"school":null,"contributors":["Hao Tang","Zhigang Zhu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-01-01T08:00:00Z","date_published":"2025-01-01T08:00:00Z","updated_at":"2026-07-24T01:58:13Z","subjects":["Machine Learning","Audio Data Processing","Assistive Navigation","Blind or Low Vision","Data Science"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://academicworks.cuny.edu/cc_etds_theses/1273","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hao Tang","Zhigang Zhu"]},{"key":"dc:creator","label":"Author","values":["Lam, Wayne"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2025-05-20T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (M.S.)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","Audio Data Processing","Assistive Navigation","Blind or Low Vision","Data Science"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://academicworks.cuny.edu/cc_etds_theses/1273"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Navigating urban environments poses significant challenges for blind and low vision (BLV) individuals, particularly at street intersections where determining when it is safe to cross can be life-threatening. In New York City, where pedestrian fatalities are on the rise and only 2% of intersections are equipped with Accessible Pedestrian Signals (APS), alternative solutions are urgently needed. This thesis proposes an audio-based deep learning approach to support BLV individuals at crosswalks by detecting traffic movement direction and idling states using spatial sound. With 4-channel audio capturing capabilities of wearables, such as Meta Project Aria glasses, we explore state-of-the-art sound event localization and detection (SELD) models, specifically ResNet-based feature extractors and transformer-enhanced architectures such as EINV2. To train these models, a synthetic dataset of quadraphonic audio from simulated 4-way traffic scenes is generated using Unity Engine, addressing both the scalability and privacy concerns of real-world data collection. This work begins an exploration into a hands-free, audio-centric crosswalk navigation aid for BLV individuals.</p>"]},{"key":"dc:title","label":"Title","values":["Utilizing Deep Learning Audio Models for Blind and Low Vision Crosswalk Assistance"]}]}],"canonical_facts":{"dc:contributor":["Hao Tang","Zhigang Zhu"],"dc:creator":["Lam, Wayne"],"dc:date.available":["2025-05-20T07:00:00Z"],"dc:description.abstract":["<p>Navigating urban environments poses significant challenges for blind and low vision (BLV) individuals, particularly at street intersections where determining when it is safe to cross can be life-threatening. In New York City, where pedestrian fatalities are on the rise and only 2% of intersections are equipped with Accessible Pedestrian Signals (APS), alternative solutions are urgently needed. This thesis proposes an audio-based deep learning approach to support BLV individuals at crosswalks by detecting traffic movement direction and idling states using spatial sound. With 4-channel audio capturing capabilities of wearables, such as Meta Project Aria glasses, we explore state-of-the-art sound event localization and detection (SELD) models, specifically ResNet-based feature extractors and transformer-enhanced architectures such as EINV2. To train these models, a synthetic dataset of quadraphonic audio from simulated 4-way traffic scenes is generated using Unity Engine, addressing both the scalability and privacy concerns of real-world data collection. This work begins an exploration into a hands-free, audio-centric crosswalk navigation aid for BLV individuals.</p>"],"dc:identifier":["https://academicworks.cuny.edu/cc_etds_theses/1273"],"dc:subject":["Machine Learning","Audio Data Processing","Assistive Navigation","Blind or Low Vision","Data Science"],"dc:title":["Utilizing Deep Learning Audio Models for Blind and Low Vision Crosswalk Assistance"],"thesis:degree_discipline":["Engineering"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Master of Science (M.S.)"]},"updated_at":"2026-07-24T01:58:13Z"}