Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 14 of 14 for “"low-resource language"”.

  1. Crowdsourcing a text corpus for a low resource language

    Low resourced languages, such as South Africa's isiXhosa, have a limited number of digitised texts, making it challenging to build language corpora and the information retrieval services, such as search and translation that depend on them. Researchers have been unable to assemble isiXhosa corpora …

    cape-town Repository record for Crowdsourcing a text corpus for a low resource language (opens in a new tab)

  2. Data augmentation and data efficiency for low-resource language processing

    Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-05-01

    uiuc Repository record for Data augmentation and data efficiency for low-resource language processing (opens in a new tab)

  3. Consonant (De)gradation in Ingrian?

    … present a dual method toward data enrichment for low-resource languages. Using Yoyodyne -- a Fairseq-inspired neural library for small-vocabulary sequence-to-sequence generation -- a morphological generation task was tested across labeled data encompassing multiple stages of enrichment for the …

    cuny-grad Repository record for Consonant (De)gradation in Ingrian? (opens in a new tab)

  4. Improving neural language models on low-resource creole languages

    When using neural models for NLP tasks, like language modelling, it is difficult to utilize a language with little data, also known as a low-resource language. Creole languages are frequently low-resource and as such it is difficult to train neural language models for them well. Creole languages …

    uiuc Repository record for Improving neural language models on low-resource creole languages (opens in a new tab)

  5. Modeling phones, keywords, topics and intents in spoken languages

    Spoken Language Understanding for both rich-resource languages (RRL) and low-resource languages (LRL) is an important research area for academia and the commercial world. In the conversational situations where either the language used in speech is a minority one, or the environment is noisy, …

    uiuc Repository record for Modeling phones, keywords, topics and intents in spoken languages (opens in a new tab)

  6. Multilingual techniques for low resource automatic speech recognition

    Out of the approximately 7000 languages spoken around the world, there are only about 100 languages with Automatic Speech Recognition (ASR) capability. This is due to the fact that a vast amount of resources is required to build a speech recognizer. This often includes thousands of hours of …

    mit Repository record for Multilingual techniques for low resource automatic speech recognition (opens in a new tab)

  7. Crosslingual Sharing for Low-Resource Natural Language Processing

    … successful at a wide variety of tasks, including language modeling and structured prediction problems such as syntactic and semantic parsing. This is due in large part to the use of supervised neural networks and more recently to unsupervised contextualized representations. However, these …

    washington Repository record for Crosslingual Sharing for Low-Resource Natural Language Processing (opens in a new tab)

  8. Grammatical analysis of Maltese text

    … analysis of Maltese words. Maltese is a hybrid language, with Semitic influence mostly visible in its grammar rules (following a root-and-pattern conjucation system) and a strong lexical influence from Italian and English. Given these influences, this project looks at grammatical inference from …

    malta Repository record for Grammatical analysis of Maltese text (opens in a new tab)

  9. Rimarju

    Rhyming English words using Natural Language Processing has been done using multiple techniques over recent years, however not many studies have carried out the same process on different languages. The Maltese language and culture has multiple aspects in which rhyming words comes into play, from …

    malta Repository record for Rimarju (opens in a new tab)

  10. On the Ethics and Linguistic Impacts of Using the Bible as Training Data for Yucatec Maya-to-Spanish Machine Translation

    … Bible, are commonly used as training data for low-resource machine translation (MT) systems because they constitute some of the most extensive and systematically digitized parallel corpora available for many languages. However, this practice raises both linguistic and ethical concerns, …

    washington Repository record for On the Ethics and Linguistic Impacts of Using the Bible as Training Data for Yucatec Maya-to-Spanish Machine Translation (opens in a new tab)

  11. Analyzing Networks with Hypergraphs: Detection, Classification, and Prediction

    … tailored for sarcasm detection across multiple low-resource languages. Our model excels in interpreting the subtle and context-dependent nature of sarcasm in short texts by exploiting the power of hypergraph structures to capture complex, high-order relationships among words. Through the …

    vt Repository record for Analyzing Networks with Hypergraphs: Detection, Classification, and Prediction (opens in a new tab)

  12. Unsupervised speech technology for low-resource languages

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms

    uiuc Repository record for Unsupervised speech technology for low-resource languages (opens in a new tab)

  13. Unsupervised speech technology for low-resource languages

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms

    uiuc Repository record for Unsupervised speech technology for low-resource languages (opens in a new tab)