Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 24 for “"speech-to-text"”.

  1. Learning shared semantic space for speech-to-text translation

    Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01

    uiuc Repository record for Learning shared semantic space for speech-to-text translation (opens in a new tab)

  2. A speech recognition module for speech-to-text language translation

    Thesis (S.B. and M.Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1998.

    mit Repository record for A speech recognition module for speech-to-text language translation (opens in a new tab)

  3. Enforcing constraints for multi-lingual and cross-lingual speech-to-text systems

    The student, Junrui Ni, accepted the attached license on 2021-12-08 at 18:02.

    uiuc Repository record for Enforcing constraints for multi-lingual and cross-lingual speech-to-text systems (opens in a new tab)

  4. The Effect of Speech-to-Text Software on Learning a New Writing Strategy

    … Mertens, 2013). Several studies have shown that speech-to-text (STT) software can improve students' writing on a specific text (Higgins & Raskind, 1997; MacArthur & Cavalier, 2004; Quinlan, 2004); however, the question of whether STT can be used to teach writing strategies has been neglected. …

    uwo Repository record for The Effect of Speech-to-Text Software on Learning a New Writing Strategy (opens in a new tab)

  5. Does Speech-To-Text Assistive Technology Improve the Written Expression of Students with Traumatic Brain Injury?

    … Brain Injury outcomes vary by individual due to age at the onset of injury, the location of the injury, and the degree to which the deficits appear to be pronounced, among other factors. As an acquired injury to the brain, the neurophysiological consequences are not homogenous; they are as …

    duquesne Repository record for Does Speech-To-Text Assistive Technology Improve the Written Expression of Students with Traumatic Brain Injury? (opens in a new tab)

  6. Constructing Low Resource Approaches to Improve Speech-to-text Translation from Modern Standard Arabic to English

    This thesis explores novel approaches to the Arabic-English speech-to-text translation task. First, we construct a novel Modern Standard Arabic speech and English text parallel dataset. Second, we propose a novel framework for leveraging unsupervised machine translation to improve speech-to-text

    mit Repository record for Constructing Low Resource Approaches to Improve Speech-to-text Translation from Modern Standard Arabic to English (opens in a new tab)

  7. Strategic Selection of Training Data for Domain-Specific Speech Recognition

    <p>Speech recognition is now a key topic in computer science with the proliferation of voice-activated assistants, and voice-enabled devices. Many companies over a speech recognition service for developers to use to enable smart devices and services. These speech-to-text systems, however, have …

    calpoly Repository record for Strategic Selection of Training Data for Domain-Specific Speech Recognition (opens in a new tab)

  8. Text Entry in Virtual Reality; Implementation of FLIK method and Text Entry Test-bed

    We present new testing software for text entry techniques. Text entry is the act of entering text via some interaction technique into a machine or monitor, common text entry techniques are the typical QWERTY keyboard, smartphone touch virtual keyboards, speech to text techniques, and others. The …

    carleton Repository record for Text Entry in Virtual Reality; Implementation of FLIK method and Text Entry Test-bed (opens in a new tab)

  9. Recognizing Speech with Large Language Models

    … has shown that large language models can be made to parse the contents of non-text embeddings and use those contents to perform various tasks. However, work focusing on audio inputs to large language models has thus far focused on either training a joint audio-text model from scratch on a lot of …

    mit Repository record for Recognizing Speech with Large Language Models (opens in a new tab)

  10. Fact-based visual question answering using knowledge graph embeddings

    Humans have a remarkable capability to learn new concepts, process them in relation to their existing mental models of the world, and seamlessly leverage their knowledge and experiences while reasoning about the outside world perceived through vision and language. Fact-based Visual Question …

    uiuc Repository record for Fact-based visual question answering using knowledge graph embeddings (opens in a new tab)

  11. Detection of the uniqueness of a human voice: towards machine learning for improved data efficiency

    The aim of this thesis is to characterise voice characteristics that can establish the identity of the person who is speaking, independent of the language used. The fundamental goal of the work is to understand how humans recognise a speaker. The voice parameters such as: speech rate, natural …

    greenwich Repository record for Detection of the uniqueness of a human voice: towards machine learning for improved data efficiency (opens in a new tab)

  12. Learning Audio-Video Language Representations

    Automatic speech recognition has seen recent advancements powered by machine learning, but it is still only available for a small fraction of the more than 7,000 languages spoken worldwide due to the reliance on manually annotated speech data. Unlabeled multi-modal data, such as videos, are now …

    mit Repository record for Learning Audio-Video Language Representations (opens in a new tab)

  13. Exploring the provision of learning support for adolescents with special needs within an educational setting

    … the significance of delivering learning support to enhance educational outcomes for adolescents with special needs. The primary components of learning support encompass personalised assistance, inclusive practices, and robust collaboration among educators and families. Access to specialised …

    western-cape Repository record for Exploring the provision of learning support for adolescents with special needs within an educational setting (opens in a new tab)

  14. Effects of Training on Intent, Ease, Self-Efficacy, Frequency, and Usefulness in Multimedia-Based Feedback for University-Level Instructors Using Canvas® LMS

    <p>The purpose of this study was to investigate how training and professional development effected university-level instructors’ perceived usefulness, perceived ease of use, behavioral intent to use, perception of self-efficacy, and frequency of use of audio-, video-, and …

    usfca Repository record for Effects of Training on Intent, Ease, Self-Efficacy, Frequency, and Usefulness in Multimedia-Based Feedback for University-Level Instructors Using Canvas® LMS (opens in a new tab)

  15. Unsupervised learning of cross-modal mappings between speech and text

    … vision, natural language processing, and speech and audio processing. Current deep learning models, however, rely on signicant amounts of supervision for training to achieve exceptional performance. For example, commercial speech recognition systems are usually trained on tens of thousands …

    mit Repository record for Unsupervised learning of cross-modal mappings between speech and text (opens in a new tab)

  16. Semi-supervised cycle-consistency training for end-to-end ASR using unpaired speech

    … colleagues (2019), which introduces a new method to train end-to-end automatic speech recognition (ASR) models using unpaired speech. In general, large amounts of paired data (speech and text) are needed to train an end-to-end automatic speech recognition system. To alleviate the problem of …

    uiuc Repository record for Semi-supervised cycle-consistency training for end-to-end ASR using unpaired speech (opens in a new tab)

  17. Transfer Learning For Spoken Language Processing

    … we tackle domain adaptation in the context of Automatic Speech Recognition (ASR) and Cross-Lingual Learning in Automatic Speech Translation (AST). The first part of the thesis develops an algorithm for unsupervised domain adaptation of End-to-End ASR models. In recent years, ASR …

    mit Repository record for Transfer Learning For Spoken Language Processing (opens in a new tab)

  18. Large Scale Online Aggregation Via Distributed Systems

    From movie recommendations to fraud detection to personalized health care, there is growing need to analyze huge amounts of data quickly. To deal with huge amounts of data, many analysts use MapReduce, a software framework that parallelizes computations across a compute cluster. However, due to the …

    rice Repository record for Large Scale Online Aggregation Via Distributed Systems (opens in a new tab)

  19. Socio-Cultural Communication System--a Communication Mechanism For Multi-Media Information Access System for Non-Literate and Linguistically Diverse Users

    … of languages. However, there is a growing need to provide network services to non-literate or linguistically diverse users. The Socio-cultural communication system (SoCCS) is a communication mechanism for building an information access system that helps both literate and non-literate users to

    umkc Repository record for Socio-Cultural Communication System--a Communication Mechanism For Multi-Media Information Access System for Non-Literate and Linguistically Diverse Users (opens in a new tab)

  20. Multimodal Eye-Tracking and Predictive Gaze Modelling for Assistive Human–Computer Interaction in Augmentative and Alternative Communication Systems

    … Amyotrophic Lateral Sclerosis (ALS) and other motor impairments. These systems provide a means to express intent when conventional communication methods such as natural speech or manual typing are unavailable. Current eye-tracking solutions frequently fail due to three critical issues: low input …

    northampton Repository record for Multimodal Eye-Tracking and Predictive Gaze Modelling for Assistive Human–Computer Interaction in Augmentative and Alternative Communication Systems (opens in a new tab)

Page 1 of 2