Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 70 for “"automatic speech recognition (ASR)"”.

  1. Incorporating pitch features for tone modeling in automatic recognition of Mandarin Chinese

    … of pitch data should improve the performance of automatic speech recognition (ASR) systems on Mandarin Chinese. The focus of this thesis is to improve the performance of a non-tonal automatic speech recognition (ASR) system on a Mandarin Chinese corpus by implementing modifications to the system …

    mit Repository record for Incorporating pitch features for tone modeling in automatic recognition of Mandarin Chinese (opens in a new tab)

  2. Improving searchability of automatically transcribed lectures through dynamic language modelling

    … can enable faster navigation and searching. Automatic speech recognition (ASR) technologies may be used to create automated transcripts, to avoid the significant time and cost involved in manual transcription.

    cape-town Repository record for Improving searchability of automatically transcribed lectures through dynamic language modelling (opens in a new tab)

  3. Improving performance of a GSM-based speech recognizer

    … if we could get the machines to understand human speech, we could be able to communicate with people (and even make communicating with computers open to many people). This is the main motivation behind Automatic Speech Recognition (ASR) as a field of research, to enable machines to recognize human …

    cape-town Repository record for Improving performance of a GSM-based speech recognizer (opens in a new tab)

  4. Visual Speech Recognition Using a 3D Convolutional Neural Network

    <p>Main stream automatic speech recognition (ASR) makes use of audio data to identify spoken words, however visual speech recognition (VSR) has recently been of increased interest to researchers. VSR is used when audio data is corrupted or missing entirely and also to further enhance the accuracy …

    calpoly Repository record for Visual Speech Recognition Using a 3D Convolutional Neural Network (opens in a new tab)

  5. Lip Detection and Adaptive Tracking

    <p>Performance of automatic speech recognition (ASR) systems utilizing only acoustic information degrades significantly in noisy environments such as a car cabins. Incorporating audio and visual information together can improve performance in these situations. This work proposes a lip detection and …

    calpoly Repository record for Lip Detection and Adaptive Tracking (opens in a new tab)

  6. Evaluation of the usability and usefulness of automatic speech recognition among users in South Africa

    An automatic speech recognition (ASR) system is a software application which recognizes human speech, processes it as input, and displays a text version of the speech as output or uses the input as commands for another application's usage. ASR can either be speaker-dependent or speaker-independent. …

    cape-town Repository record for Evaluation of the usability and usefulness of automatic speech recognition among users in South Africa (opens in a new tab)

  7. Toward More Intelligible Simultaneous Multi-channel Speech Enhancement and Recognition

    <p>Traditional single-channel speech enhancement and separation methods focus on enhancing the target speech signal by suppressing the noise and interfering speech signal. The methods suffer from nonlinear distortion brought by the algorithm, which hurts the intelligibility of the speech and also …

    cuny-grad Repository record for Toward More Intelligible Simultaneous Multi-channel Speech Enhancement and Recognition (opens in a new tab)

  8. Self-Supervised Audio-Visual Speech Diarization and Recognition

    Many real world use cases of automatic speech recognition (ASR) contain video and multiple speakers, such as TV broadcasts and video conferences. However, state-of-the-art end-to-end multimodal ASR models generally do not support diarization. This thesis extends one such model, AV-HuBERT, to …

    mit Repository record for Self-Supervised Audio-Visual Speech Diarization and Recognition (opens in a new tab)

  9. TASK- AND DOMAIN-DEPENDENT NATURAL LANGUAGE PROCESSING FOR SUPPORTING SPOKEN DOCUMENT USE

    … prototypes. The presented field studies identify Automatic Speech Recognition (ASR) and Sentiment Analysis (SA) as two requisite NLP technologies for supporting information seekers in the studied domain. The dissertation also demonstrates how SA models can be optimized with TDD feedback taken into …

    toronto-retro Repository record for TASK- AND DOMAIN-DEPENDENT NATURAL LANGUAGE PROCESSING FOR SUPPORTING SPOKEN DOCUMENT USE (opens in a new tab)

  10. Analyzing Icelandic conversation using State-of-the-Art ASR models

    … performance and fine-tuning of transformer-based Automatic Speech Recognition (ASR) systems applied to Icelandic conversational speech, with a primary focus on OpenAI’s Whisper model and a secondary fo- cus on Meta’s Wav2Vec 2.0. The study uses the Spjallrómur dataset, a 21-hour unscripted …

    reykjavik Repository record for Analyzing Icelandic conversation using State-of-the-Art ASR models (opens in a new tab)

  11. Towards an end-to-end music transcription system using neural networks

    … It has notable parallels with the task of Automatic Speech Recognition (ASR) and indeed from this connection arises some natural Machine Learning-based approaches. However, these methods usually involve carefully designed preprocessing steps, or transcription into less flexible …

    uiuc Repository record for Towards an end-to-end music transcription system using neural networks (opens in a new tab)

  12. Adaptation of hybrid deep neural network-hidden Markov model speech recognition system using a sub-space approach

    The performance of automatic speech recognition (ASR) system can be enhanced by adaptation of the ASR for a particular speaker or a group of speakers. In ASR, training and testing data often do not follow the same statistics; they are often mismatched, which leads to a gap in performance. The …

    gatech Repository record for Adaptation of hybrid deep neural network-hidden Markov model speech recognition system using a sub-space approach (opens in a new tab)

  13. Effective automatic speech recognition data collection for under–resourced languages

    As building transcribed speech corpora for under-resourced languages plays a pivotal role in developing automatic speech recognition (ASR) technologies for such languages, a key step in developing these technologies is the effective collection of ASR data, consisting of transcribed audio and …

    nwu-za Repository record for Effective automatic speech recognition data collection for under–resourced languages (opens in a new tab)

  14. Machine learning and deep learning techniques for natural language processing with application to audio recordings

    … of the recordings to text was done using Automatic Speech Recognition (ASR), followed by data cleaning and the transcribed text was represented in numerical form using the Term Frequency-Inverse Document Frequency (TF- IDF) and the Count Vectorizer. The study then compared the accuracy of …

    nwu-za Repository record for Machine learning and deep learning techniques for natural language processing with application to audio recordings (opens in a new tab)

  15. Robust Unconstrained Face Detection and Lip Localization Using Gabor Filters

    Automatic speech recognition (ASR) is a well-researched field of study aimed at augmenting the man-machine interface through interpretation of the spoken word. From in-car voice recognition systems to automated telephone directories, automatic speech recognition technology is becoming increasingly …

    calpoly Repository record for Robust Unconstrained Face Detection and Lip Localization Using Gabor Filters (opens in a new tab)

  16. Parts-based models and local features for automatic speech recognition

    While automatic speech recognition (ASR) systems have steadily improved and are now in widespread use, their accuracy continues to lag behind human performance, particularly in adverse conditions. This thesis revisits the basic acoustic modeling assumptions common to most ASR systems and argues …

    mit Repository record for Parts-based models and local features for automatic speech recognition (opens in a new tab)

  17. Language Modeling from Visually Grounded Speech

    … language processing have significantly reduced automatic speech recognition (ASR) error rates, driven by large-scale supervised training on paired speech–text data and, more recently, self-supervised pre-training on unpaired speech and audio. These methods have facilitated robust transfer …

    mit Repository record for Language Modeling from Visually Grounded Speech (opens in a new tab)

  18. Evaluating the Effects of Automatic Speech Recognition Word Accuracy

    Automatic Speech Recognition (ASR) research has been primarily focused towards large-scale systems and industry, while other areas that require attention are often over-looked by researchers. For this reason, this research looked at automatic speech recognition at the consumer level. Many …

    vt Repository record for Evaluating the Effects of Automatic Speech Recognition Word Accuracy (opens in a new tab)

  19. Implementing a distributed approach for speech resource and system development

    The range of applications for high-quality automatic speech recognition (ASR) systems has grown dramatically with the advent of smart phones, in which speech recognition can greatly enhance the user experience. Currently, the languages with extensive ASR support on these devices are languages that …

    nwu-za Repository record for Implementing a distributed approach for speech resource and system development (opens in a new tab)

Page 1 of 4