Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 24 for “"speech-to-text"”.
-
Learning shared semantic space for speech-to-text translation
Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01
-
A speech recognition module for speech-to-text language translation
Thesis (S.B. and M.Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1998.
-
Enforcing constraints for multi-lingual and cross-lingual speech-to-text systems
The student, Junrui Ni, accepted the attached license on 2021-12-08 at 18:02.
-
The Effect of Speech-to-Text Software on Learning a New Writing Strategy
… Mertens, 2013). Several studies have shown that speech-to-text (STT) software can improve students' writing on a specific text (Higgins & Raskind, 1997; MacArthur & Cavalier, 2004; Quinlan, 2004); however, the question of whether STT can be used to teach writing strategies has been neglected. …
-
Does Speech-To-Text Assistive Technology Improve the Written Expression of Students with Traumatic Brain Injury?
… Brain Injury outcomes vary by individual due to age at the onset of injury, the location of the injury, and the degree to which the deficits appear to be pronounced, among other factors. As an acquired injury to the brain, the neurophysiological consequences are not homogenous; they are as …
-
Constructing Low Resource Approaches to Improve Speech-to-text Translation from Modern Standard Arabic to English
This thesis explores novel approaches to the Arabic-English speech-to-text translation task. First, we construct a novel Modern Standard Arabic speech and English text parallel dataset. Second, we propose a novel framework for leveraging unsupervised machine translation to improve speech-to-text …
-
Strategic Selection of Training Data for Domain-Specific Speech Recognition
<p>Speech recognition is now a key topic in computer science with the proliferation of voice-activated assistants, and voice-enabled devices. Many companies over a speech recognition service for developers to use to enable smart devices and services. These speech-to-text systems, however, have …
-
Text Entry in Virtual Reality; Implementation of FLIK method and Text Entry Test-bed
We present new testing software for text entry techniques. Text entry is the act of entering text via some interaction technique into a machine or monitor, common text entry techniques are the typical QWERTY keyboard, smartphone touch virtual keyboards, speech to text techniques, and others. The …
-
Recognizing Speech with Large Language Models
… has shown that large language models can be made to parse the contents of non-text embeddings and use those contents to perform various tasks. However, work focusing on audio inputs to large language models has thus far focused on either training a joint audio-text model from scratch on a lot of …
-
Fact-based visual question answering using knowledge graph embeddings
Humans have a remarkable capability to learn new concepts, process them in relation to their existing mental models of the world, and seamlessly leverage their knowledge and experiences while reasoning about the outside world perceived through vision and language. Fact-based Visual Question …
-
Detection of the uniqueness of a human voice: towards machine learning for improved data efficiency
The aim of this thesis is to characterise voice characteristics that can establish the identity of the person who is speaking, independent of the language used. The fundamental goal of the work is to understand how humans recognise a speaker. The voice parameters such as: speech rate, natural …
-
Learning Audio-Video Language Representations
Automatic speech recognition has seen recent advancements powered by machine learning, but it is still only available for a small fraction of the more than 7,000 languages spoken worldwide due to the reliance on manually annotated speech data. Unlabeled multi-modal data, such as videos, are now …
-
Exploring the provision of learning support for adolescents with special needs within an educational setting
… the significance of delivering learning support to enhance educational outcomes for adolescents with special needs. The primary components of learning support encompass personalised assistance, inclusive practices, and robust collaboration among educators and families. Access to specialised …
-
Effects of Training on Intent, Ease, Self-Efficacy, Frequency, and Usefulness in Multimedia-Based Feedback for University-Level Instructors Using Canvas® LMS
<p>The purpose of this study was to investigate how training and professional development effected university-level instructors’ perceived usefulness, perceived ease of use, behavioral intent to use, perception of self-efficacy, and frequency of use of audio-, video-, and …
-
Unsupervised learning of cross-modal mappings between speech and text
… vision, natural language processing, and speech and audio processing. Current deep learning models, however, rely on signicant amounts of supervision for training to achieve exceptional performance. For example, commercial speech recognition systems are usually trained on tens of thousands …
-
Semi-supervised cycle-consistency training for end-to-end ASR using unpaired speech
… colleagues (2019), which introduces a new method to train end-to-end automatic speech recognition (ASR) models using unpaired speech. In general, large amounts of paired data (speech and text) are needed to train an end-to-end automatic speech recognition system. To alleviate the problem of …
-
Transfer Learning For Spoken Language Processing
… we tackle domain adaptation in the context of Automatic Speech Recognition (ASR) and Cross-Lingual Learning in Automatic Speech Translation (AST). The first part of the thesis develops an algorithm for unsupervised domain adaptation of End-to-End ASR models. In recent years, ASR …
-
Large Scale Online Aggregation Via Distributed Systems
From movie recommendations to fraud detection to personalized health care, there is growing need to analyze huge amounts of data quickly. To deal with huge amounts of data, many analysts use MapReduce, a software framework that parallelizes computations across a compute cluster. However, due to the …
-
Socio-Cultural Communication System--a Communication Mechanism For Multi-Media Information Access System for Non-Literate and Linguistically Diverse Users
… of languages. However, there is a growing need to provide network services to non-literate or linguistically diverse users. The Socio-cultural communication system (SoCCS) is a communication mechanism for building an information access system that helps both literate and non-literate users to …
-
Multimodal Eye-Tracking and Predictive Gaze Modelling for Assistive Human–Computer Interaction in Augmentative and Alternative Communication Systems
… Amyotrophic Lateral Sclerosis (ALS) and other motor impairments. These systems provide a means to express intent when conventional communication methods such as natural speech or manual typing are unavailable. Current eye-tracking solutions frequently fail due to three critical issues: low input …
Page 1 of 2