Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 15 of 15 for “"out-of-vocabulary"”.

  1. Modelling out-of-vocabulary words for robust speech recognition

    This thesis concerns the problem of unknown or out-of-vocabulary (OOV) words in continuous speech recognition. Most of today's state-of-the-art speech recognition systems can recognize only words that belong to some predefined finite word vocabulary. When encountering an OOV word, a speech …

    mit Repository record for Modelling out-of-vocabulary words for robust speech recognition (opens in a new tab)

  2. A study on out-of-vocabulary word modelling for a segment-based keyword spotting system

    Thesis (M.S.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1996.

    mit Repository record for A study on out-of-vocabulary word modelling for a segment-based keyword spotting system (opens in a new tab)

  3. A characterization of the problem of new, out-of-vocabulary words in continuous-speech recognition and understanding

    Thesis (Ph. D.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1995.

    mit Repository record for A characterization of the problem of new, out-of-vocabulary words in continuous-speech recognition and understanding (opens in a new tab)

  4. Incorporate Out-of-Vocabulary Words for Psycholinguistic Analysis using Social Media Texts - An OOV-aware Data Curation Process and a Hybrid Approach

    … differences such as personality. SMT are often written in an informal way, and thus contain lexical variants such as nonstandard spellings, capitalizations, and abbreviations. These lexical variants are referred as out-of-vocabulary (OOV) words. They are not captured in standard …

    claremont Repository record for Incorporate Out-of-Vocabulary Words for Psycholinguistic Analysis using Social Media Texts - An OOV-aware Data Curation Process and a Hybrid Approach (opens in a new tab)

  5. Word alignment and smoothing methods in statistical machine translation: Noise, prior knowledge and overfitting

    … an SMT system. Although one important category of linguistic knowledge is that obtained by a constituent / dependency parser, a POS / super tagger, and a morphological analyser, linguistic knowledge here includes larger domains than this: Multi-Word Expressions, Out-Of-Vocabulary words, …

    dcu Repository record for Word alignment and smoothing methods in statistical machine translation: Noise, prior knowledge and overfitting (opens in a new tab)

  6. Towards multi-domain speech understanding with flexible and dynamic vocabulary

    … systems, we foresee future systems capable of supporting multiple domains and flexible vocabulary. Users can pursue several topics of interest within a single telephone call, and the system is able to switch transparently among domains within a single dialog. This system is able to detect …

    mit Repository record for Towards multi-domain speech understanding with flexible and dynamic vocabulary (opens in a new tab)

  7. Lexical and Language Modeling of Diacritics and Morphemes in Arabic Automatic Speech Recognition

    … rarely displays diacritics. These two features of the language pose challenges when building Automatic Speech Recognition (ASR) systems. Morphological complexity leads to many possible combinations of stems and affixes to form words, and produces texts with high Out Of Vocabulary (OOV) rates. In …

    mit Repository record for Lexical and Language Modeling of Diacritics and Morphemes in Arabic Automatic Speech Recognition (opens in a new tab)

  8. Online Adaptive Neural Machine Translation: from single- to multi-domain scenarios

    … application scenarios related to the use of MT in computer assisted translation (CAT), where human translators post-edit MT outputs. In particular, we investigate (in chronological order) MT adaptation under two working conditions: single-domain and multi-domain. In the former, we assume …

    trento Repository record for Online Adaptive Neural Machine Translation: from single- to multi-domain scenarios (opens in a new tab)

  9. Hospital readmission prediction with long clinical notes

    … institutions, resulting in large amounts of diverse information that can be analysed for diagnosis, prognosis, treatment and prevention of disease. One type of data captured by EHRs are clinical notes, which are unstructured data written in natural language. We can leverage Natural …

    cape-town Repository record for Hospital readmission prediction with long clinical notes (opens in a new tab)

  10. Linguistically-motivated sub-word modeling with applications to speech recognition

    Despite the proliferation of speech-enabled applications and devices, speech-driven human-machine interaction still faces several challenges. One of theses issues is the new word or the out-of-vocabulary (OOV) problem, which occurs when the underlying automatic speech recognizer (ASR) encounters a …

    mit Repository record for Linguistically-motivated sub-word modeling with applications to speech recognition (opens in a new tab)

  11. Robust machine translation for multi-domain tasks

    … to improved concepts and algorithms, the quality of the generated translation hypotheses has been significantly improved in recent years. Still, the translation quality leaves a lot to be desired when going beyond traditional translation tasks, such as newswire articles, and when addressing more …

    aachen Repository record for Robust machine translation for multi-domain tasks (opens in a new tab)

  12. Text mining at multiple granularity: leveraging subwords, words, phrases, and sentences

    With the rapid digitization of information, large quantities of text-heavy data is being constantly generated in many languages and across domains such as web documents, research papers, business reviews, news, and social posts. As such, efficiently and effectively searching, organizing, and …

    uiuc Repository record for Text mining at multiple granularity: leveraging subwords, words, phrases, and sentences (opens in a new tab)

  13. Deep Learning for Unstructured Data by Leveraging Domain Knowledge

    … a sentence is first converted to a vector of word counts, and then fed into a classification algorithm such as logistic regression and support vector machine. The creation of such numerical vectors is very challenging and difficult. Recent progress in deep learning provides us a new way to …

    temple Repository record for Deep Learning for Unstructured Data by Leveraging Domain Knowledge (opens in a new tab)

  14. Hybrid Words Representation for the classification of low quality text

    … in its contents. The effective capturing of such content relationships is thereby crucial for a better understanding of text representations. This is especially challenging in the environments where the text messages are short, informal and noisy, and involves natural language ambiguities. …

    uts Repository record for Hybrid Words Representation for the classification of low quality text (opens in a new tab)

  15. Data quality in the deep learning era: Active semi-supervised learning and text normalization for natural language understanding

    Deep Learning, a growing sub-field of machine learning, has been applied with tremendous success in a variety of domains, opening opportunities for achieving human level performance in many applications. However, Deep Learning methods depend on large quantities of data with millions of annotated …

    uiuc Repository record for Data quality in the deep learning era: Active semi-supervised learning and text normalization for natural language understanding (opens in a new tab)