Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 15 of 15 for “"out-of-vocabulary"”.
-
Modelling out-of-vocabulary words for robust speech recognition
This thesis concerns the problem of unknown or out-of-vocabulary (OOV) words in continuous speech recognition. Most of today's state-of-the-art speech recognition systems can recognize only words that belong to some predefined finite word vocabulary. When encountering an OOV word, a speech …
-
A study on out-of-vocabulary word modelling for a segment-based keyword spotting system
Thesis (M.S.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1996.
-
A characterization of the problem of new, out-of-vocabulary words in continuous-speech recognition and understanding
Thesis (Ph. D.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1995.
-
Incorporate Out-of-Vocabulary Words for Psycholinguistic Analysis using Social Media Texts - An OOV-aware Data Curation Process and a Hybrid Approach
… differences such as personality. SMT are often written in an informal way, and thus contain lexical variants such as nonstandard spellings, capitalizations, and abbreviations. These lexical variants are referred as out-of-vocabulary (OOV) words. They are not captured in standard …
-
Word alignment and smoothing methods in statistical machine translation: Noise, prior knowledge and overfitting
… an SMT system. Although one important category of linguistic knowledge is that obtained by a constituent / dependency parser, a POS / super tagger, and a morphological analyser, linguistic knowledge here includes larger domains than this: Multi-Word Expressions, Out-Of-Vocabulary words, …
-
Towards multi-domain speech understanding with flexible and dynamic vocabulary
… systems, we foresee future systems capable of supporting multiple domains and flexible vocabulary. Users can pursue several topics of interest within a single telephone call, and the system is able to switch transparently among domains within a single dialog. This system is able to detect …
-
Lexical and Language Modeling of Diacritics and Morphemes in Arabic Automatic Speech Recognition
… rarely displays diacritics. These two features of the language pose challenges when building Automatic Speech Recognition (ASR) systems. Morphological complexity leads to many possible combinations of stems and affixes to form words, and produces texts with high Out Of Vocabulary (OOV) rates. In …
-
Online Adaptive Neural Machine Translation: from single- to multi-domain scenarios
… application scenarios related to the use of MT in computer assisted translation (CAT), where human translators post-edit MT outputs. In particular, we investigate (in chronological order) MT adaptation under two working conditions: single-domain and multi-domain. In the former, we assume …
-
Hospital readmission prediction with long clinical notes
… institutions, resulting in large amounts of diverse information that can be analysed for diagnosis, prognosis, treatment and prevention of disease. One type of data captured by EHRs are clinical notes, which are unstructured data written in natural language. We can leverage Natural …
-
Linguistically-motivated sub-word modeling with applications to speech recognition
Despite the proliferation of speech-enabled applications and devices, speech-driven human-machine interaction still faces several challenges. One of theses issues is the new word or the out-of-vocabulary (OOV) problem, which occurs when the underlying automatic speech recognizer (ASR) encounters a …
-
Robust machine translation for multi-domain tasks
… to improved concepts and algorithms, the quality of the generated translation hypotheses has been significantly improved in recent years. Still, the translation quality leaves a lot to be desired when going beyond traditional translation tasks, such as newswire articles, and when addressing more …
-
Text mining at multiple granularity: leveraging subwords, words, phrases, and sentences
With the rapid digitization of information, large quantities of text-heavy data is being constantly generated in many languages and across domains such as web documents, research papers, business reviews, news, and social posts. As such, efficiently and effectively searching, organizing, and …
-
Deep Learning for Unstructured Data by Leveraging Domain Knowledge
… a sentence is first converted to a vector of word counts, and then fed into a classification algorithm such as logistic regression and support vector machine. The creation of such numerical vectors is very challenging and difficult. Recent progress in deep learning provides us a new way to …
-
Hybrid Words Representation for the classification of low quality text
… in its contents. The effective capturing of such content relationships is thereby crucial for a better understanding of text representations. This is especially challenging in the environments where the text messages are short, informal and noisy, and involves natural language ambiguities. …
-
Data quality in the deep learning era: Active semi-supervised learning and text normalization for natural language understanding
Deep Learning, a growing sub-field of machine learning, has been applied with tremendous success in a variety of domains, opening opportunities for achieving human level performance in many applications. However, Deep Learning methods depend on large quantities of data with millions of annotated …