Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 18 of 18 for “"Word segmentation"”.

  1. Chinese word segmentation with a maximum entropy approach

    … we present a maximum entropy approach to Chinese word segmentation. Besides using features derived from gold-standard word-segmented training data, we also used an external dictionary and additional training corpora of different segmentation standards to further improve segmentation accuracy. The …

    nus Repository record for Chinese word segmentation with a maximum entropy approach (opens in a new tab)

  2. Syllabification and Visual Word Segmentation in Spanish–English Bilinguals

    … Spanish syllable structure and whether or not word segmentation strategies are affected by these syllabic intuitions. The study utilizes monolingual Spanish speakers, L1 Spanish speakers whoare L2 learners of English and L1 English speakers who are L2 learners of Spanish. For Spanish syllabic …

    arizona-thes Repository record for Syllabification and Visual Word Segmentation in Spanish–English Bilinguals (opens in a new tab)

  3. The role of first and second language speech rhythm in syntactic ambiguity processing and musical rhythmic aptitude

    … events, e.g., sounds and pauses, together into words, making their boundaries acoustically prominent and aiding word segmentation and recognition by the hearer. After word recognition, the hearer is able to retrieve word meaning form his mental lexicon, integrating it with information from other …

    potsdam-diss Repository record for The role of first and second language speech rhythm in syntactic ambiguity processing and musical rhythmic aptitude (opens in a new tab)

  4. Semi-supervised learning for natural language

    … the performance of a number of tasks, e.g. word sense disambiguation, information extraction, and natural language parsing. In this thesis, we focus on two segmentation tasks, named-entity recognition and Chinese word segmentation. The goal of named-entity recognition is to detect and …

    mit Repository record for Semi-supervised learning for natural language (opens in a new tab)

  5. Implicit graphemic cues in Thai reading

    … not use space or other punctuation to demarcate word boundaries typically show improvements in early reading time measures when spaces are inserted. However, it seems that Thai may not follow this same trend (Winskel, Radach, & Luksaneeyanawin, 2009). One potential explanation is that Thai …

    uiuc Repository record for Implicit graphemic cues in Thai reading (opens in a new tab)

  6. Word based off-line handwritten Arabic classification and recognition. Design of automatic recognition system for large vocabulary offline handwritten Arabic words using machine learning approaches.

    The design of a machine which reads unconstrained words still remains an unsolved problem. For example, automatic interpretation of handwritten documents by a computer is still under research. Most systems attempt to segment words into letters and read words one character at a time. However, …

    bradford Repository record for Word based off-line handwritten Arabic classification and recognition. Design of automatic recognition system for large vocabulary offline handwritten Arabic words using machine learning approaches. (opens in a new tab)

  7. The effects of character predictability on eye-movement control in reading Mandarin Chinese texts

    … studies of alphabetic scripts, hypothesize that “word” is the basic unit of processing, but there is no consensus on whether this assumption is also true in Chinese reading. As word boundaries are not visually denoted in Chinese text, “word” is hard to define. Chinese readers do not agree with …

    uiuc Repository record for The effects of character predictability on eye-movement control in reading Mandarin Chinese texts (opens in a new tab)

  8. Context-dependent type-level models for unsupervised morpho-syntactic induction

    … part-of-speech (POS) induction and morphological word segmentation by modeling linguistic phenomena previously not used. For both tasks, we realize these linguistic intuitions with Bayesian generative models that first create a latent lexicon before generating unannotated tokens in the input …

    mit Repository record for Context-dependent type-level models for unsupervised morpho-syntactic induction (opens in a new tab)

  9. Word boundary detection using landmarks : a survey of consonants

    … consonants that would help in identifying their word positions in running speech. A database of sentences containing word pairs (e.g. "lay keys" vs. "lake ease" for /k/) of thirteen consonants (six stops, two affricates, three fricatives, and two nasals), controlled for prosodic boundaries, pitch …

    mit Repository record for Word boundary detection using landmarks : a survey of consonants (opens in a new tab)

  10. Biases in segmenting non-concatenative morphology

    Segmentation of words containing non-concatenative morphology into their component morphemes, such as Arabic /kita:b/ 'book' into root [check symbol]ktb and vocalism /i-a:/ (McCarthy, 1979, 1981), is a difficult task due to the size of its search space of possibilities, which grows exponentially as …

    mit Repository record for Biases in segmenting non-concatenative morphology (opens in a new tab)

  11. A Formal Model of Ambiguity and its Applications in Machine Translation

    … instance and use it to learn a model of compound word segmentation is also introduced, which is used as a preprocessing step in machine translation.

    maryland Repository record for A Formal Model of Ambiguity and its Applications in Machine Translation (opens in a new tab)

  12. Direct Speech Translation Toward High-Quality, Inclusive, and Augmented Systems

    … effects caused by the suboptimal automatic audio segmentation in two ways: on one side, we created models robust to this condition; on the other, we enhanced the audio segmentation itself. The good results achieved in terms of overall translation quality allowed us to investigate specific …

    trento Repository record for Direct Speech Translation Toward High-Quality, Inclusive, and Augmented Systems (opens in a new tab)

  13. Time-varying networks estimation and Chinese words segmentation

    … time-varying networks estimation and Chinese words segmentation. Chapter 1 introduces the background of the time-varying networks and the structure of Chinese language, followed by the motivations and goals for the research work. In many biomedical and social science studies, it is important …

    uiuc Repository record for Time-varying networks estimation and Chinese words segmentation (opens in a new tab)

  14. Phonological Representations in Language Models

    … input representation consisting of discrete subword tokens derived from web-scraped orthographic text. Understanding and interpreting these models is crucial not only for advancing NLP technology but also for gaining insights into the structure of language itself. Recently, smaller language …

    cambridge Repository record for Phonological Representations in Language Models (opens in a new tab)

  15. ASSESSING L2 CHINESE LISTENING USING AUTHENTICATED SPOKEN TEXTS

    … even though the authenticated group had more word segmentation difficulties, this type of listening difficulty seemed to be less severe for both groups compared to the difficulty of phonological decoding; and third, other features commonly found in unscripted spoken Chinese such as filled …

    temple Repository record for ASSESSING L2 CHINESE LISTENING USING AUTHENTICATED SPOKEN TEXTS (opens in a new tab)

  16. Hybrid Words Representation for the classification of low quality text

    … natural language; to be more precise, we use words while communicating with others. However, in today's world, we wish to communicate with computers, just like humans. It is not an easy task because human communicate in an unstructured and informal way, whereas computers need structured and …

    uts Repository record for Hybrid Words Representation for the classification of low quality text (opens in a new tab)

  17. The acquisition of resyllabification in Spanish by English speakers

    … second language learners produce and recognize words in continuous speech when speech segmentation strategies differ between languages? This thesis explores phrases that have been affected by the process of resyllabification of consonants across word boundaries, where syllable and word

    uiuc Repository record for The acquisition of resyllabification in Spanish by English speakers (opens in a new tab)

  18. A Pointillism Approach for Natural Language Processing of Social Media

    … tasks typically start with the basic unit of words, and then from words and their meanings a big picture is constructed about what the meanings of documents or other larger constructs are in terms of the topics discussed. Social media is very challenging for natural language processing because …

    unm Repository record for A Pointillism Approach for Natural Language Processing of Social Media (opens in a new tab)