Abstract
Sign language is the primary form of communication among Deaf and Hard of Hearing (DHH) individuals.One of the goal of translation is to develop a system that can convert a video sequence of signing into spoken language, and vice versa—spoken language into signing videos. Sign language is a rich and expressive form of communication that relies on a combination of body gestures, facial expressions, and hand movements. This dissertation presents approaches for sign language translation from video inputs, focusing on converting sign language gestures into corresponding spoken language text. The main challenge in this domain is the limited availability of large-scale annotated datasets, as well as variations in signing styles which further complicate the development of accurate translation systems. We leverage advancements in computer vision, machine learning, and natural language processing to develop efficient, high-accuracy models for fingerspelling, word-level recognition, and gloss-to-text translation. For fingerspelling, we propose a pose-based encoder-decoder transformer trained with Connectionist Temporal Classification (CTC), cross-entropy, and length-prediction losses. For gloss-to-text translation task, we leverage LLMs and effective data augmentation techniques. To improve the handling of ambiguous labels, we introduce a novel Semantically Aware Label Smoothing (SALS) method. Finally, we leverage existing self-supervised masked video encoders for pre-training with a tailored hand-masking strategy specific to sign language data. This method applies masking to regions containing hands, where the model should reconstruct these occluded areas.
Author and committee
dc:creator, dc:contributor.*- Author
-
- Fayyazsanavi, Pooya
Identifiers
dc:identifier.*- Identifier
- hdl:1920/14742
- OAI identifier oai:identifier
- oai:MARS:1920/14742