Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 11 of 11 for “"Speech technologies"”.

  1. Spoke : a framework for building speech-enabled websites

    … a JavaScript framework for building interactive speech-enabled web applications. This project was motivated by the need for a consolidated framework for integrating custom speech technologies into website backends to demonstrate their power. Spoke achieves this by providing a Node.js server-side …

    mit Repository record for Spoke : a framework for building speech-enabled websites (opens in a new tab)

  2. A Predictive Model of Prosody Through Grammatical Interface: A Computational Approach

    … has a direct application in the development of speech technologies that incorporate linguistic models of prosody, including text-to-speech and automatic speech recognition systems.

    uiuc Repository record for A Predictive Model of Prosody Through Grammatical Interface: A Computational Approach (opens in a new tab)

  3. Speech processing with less supervision : learning from weak labels and multiple modalities

    … learning has achieved great success in speech processing with powerful neural network models and vast quantities of in-domain labeled data. However, collecting a labeled dataset covering all domains can be either expensive due to the diversity of speech or almost impossible for some …

    mit Repository record for Speech processing with less supervision : learning from weak labels and multiple modalities (opens in a new tab)

  4. Learning Audio-Video Language Representations

    Automatic speech recognition has seen recent advancements powered by machine learning, but it is still only available for a small fraction of the more than 7,000 languages spoken worldwide due to the reliance on manually annotated speech data. Unlabeled multi-modal data, such as videos, are now …

    mit Repository record for Learning Audio-Video Language Representations (opens in a new tab)

  5. Automatic subtitling: A new paradigm

    … and timed by humans. With recent developments in speech translation (ST), the time is ripe for extended automation in subtitling, with end-to-end solutions for obtaining target language subtitles directly from the source speech. In this thesis, we address the key steps for accomplishing the new …

    trento Repository record for Automatic subtitling: A new paradigm (opens in a new tab)

  6. Language technologies in speech-enabled second language learning games : from reading to dialogue

    … new opportunities for self-learning. The use of speech technologies is especially attractive to offer students unlimited chances for speaking exercises. To create helpful and intelligent speaking exercises on a computer, it is necessary for the computer to not only recognize the acoustics, but …

    mit Repository record for Language technologies in speech-enabled second language learning games : from reading to dialogue (opens in a new tab)

  7. Deep unsupervised learning from speech

    Automatic speech recognition (ASR) systems have become hugely successful in recent years - we have become accustomed to speech interfaces across all kinds of devices. However, despite the huge impact ASR has had on the way we interact with technology, it is out of reach for a significant portion of …

    mit Repository record for Deep unsupervised learning from speech (opens in a new tab)

  8. Lietuvių ir latvių tarmių monoftongų priegaidžių akustiniai požymiai: lyginamoji analizė /

    … of its positive contribution to perfecting speech technologies as well as to other similar research into other tone languages.

    vilnius Repository record for Lietuvių ir latvių tarmių monoftongų priegaidžių akustiniai požymiai: lyginamoji analizė / (opens in a new tab)

  9. Navigator- A navigation system for the visually impaired

    … This system would use voice-to-text and text-to-speech technologies to communicate with the user easily effectively. The system would also give turn-by-turn voice navigation while detecting any obstacles on the way. Two user studies were conducted to better understand these characteristics and …

    uiuc Repository record for Navigator- A navigation system for the visually impaired (opens in a new tab)

  10. Speech Foundation Models for Audio Processing

    … shaped the development of natural language and speech technologies. Foundation speech models such as Whisper have shown strong performance across a variety of audio processing tasks, including automatic speech recognition (ASR) and speech translation. Unlike traditional systems that require …

    cambridge Repository record for Speech Foundation Models for Audio Processing (opens in a new tab)