Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 14 of 14 for “"parallel corpora"”.

  1. A corpus-driven study on translation units in an English-Chinese parallel corpus

    … such units, and that they can be identified in parallel corpora. It aims to show that these translation units and their target language equivalents can be extracted from parallel corpora and can be re-used to facilitate new translations. The concept of translation units and their equivalents …

    birmingham Repository record for A corpus-driven study on translation units in an English-Chinese parallel corpus (opens in a new tab)

  2. Cross-Domain and Cross-Language Porting of Shallow Parsing

    … developed mainly for English. Developing similar corpora and tools for other languages is an important issue. However, this requires significant amount of effort. Recently, Statistical Machine Translation (SMT) techniques and parallel corpora were used to transfer annotations from a linguistic …

    trento Repository record for Cross-Domain and Cross-Language Porting of Shallow Parsing (opens in a new tab)

  3. New Data-Driven Approaches to Text Simplification

    … splitting and deletion decisions, built upon parallel corpora of original and manually simplified Spanish texts, which outperform the existing similar systems. Our experiments in adaptation of those methods to different text genres and target populations report promising results, thus offering …

    wlv Repository record for New Data-Driven Approaches to Text Simplification (opens in a new tab)

  4. Rethinking translation unit size: an empirical study of an English-Japanese newswire corpus

    … issues of translation units with the help of a parallel corpus: the ARC (the Alignment of Reuters Corpora), an English-Japanese newswire corpus. The main achievements of this study are: the identification of five variables associated with translation unit size; the establishment of an unbiased, …

    birmingham Repository record for Rethinking translation unit size: an empirical study of an English-Japanese newswire corpus (opens in a new tab)

  5. Problems and strategies in translating legal texts from arabic to english

    … a mixed-methods approach. That is, it utilises parallel corpora for legal genres. The parallel corpus is not only used for a qualitative translation quality assessment of the problems and strategies involved in the translation of legal texts but it can also be used in providing some quantitative …

    western-cape Repository record for Problems and strategies in translating legal texts from arabic to english (opens in a new tab)

  6. Automatic identification and translation of multiword expressions

    … extract translation equivalents from comparable corpora which are an alternative resource to costly parallel corpora. We finally conduct a series of experiments to investigate the effects of size and quality of comparable corpora on automatic extraction of translation equivalents.

    wlv Repository record for Automatic identification and translation of multiword expressions (opens in a new tab)

  7. On the Ethics and Linguistic Impacts of Using the Bible as Training Data for Yucatec Maya-to-Spanish Machine Translation

    … the most extensive and systematically digitized parallel corpora available for many languages. However, this practice raises both linguistic and ethical concerns, particularly for Indigenous language communities for whom Bible translation has historically been intertwined with colonialism and …

    washington Repository record for On the Ethics and Linguistic Impacts of Using the Bible as Training Data for Yucatec Maya-to-Spanish Machine Translation (opens in a new tab)

  8. Crosslingual Sharing for Low-Resource Natural Language Processing

    … in turn require crosslingual resources such as parallel corpora, which again may not be available. All of these factors pose challenges to natural language processing in low-resource languages. This thesis argues that even with little, indirect or absent crosslingual supervision, sharing …

    washington Repository record for Crosslingual Sharing for Low-Resource Natural Language Processing (opens in a new tab)

  9. Empathetic Text Style Transfer Layer for Conversational Agents a modular architecture and new resource for the italian language

    … annotated resources and the absence of parallel corpora for controlled stylistic generation. Moreover, incorporating empathetic capabilities directly into existing domain‐specific conversational agents entails significant risks for system reliability and imposes high retraining and …

    trento Repository record for Empathetic Text Style Transfer Layer for Conversational Agents a modular architecture and new resource for the italian language (opens in a new tab)

  10. Joint analysis of user-generated content and product information to enhance user experience in e-commerce

    … implicit intent in user status text leveraging parallel corpora we build from social media, and we recommend products whose descriptions satisfy the inferred intent. In order to help users make purchase decisions, we first propose to generate augmented product specifications leveraging user …

    uiuc Repository record for Joint analysis of user-generated content and product information to enhance user experience in e-commerce (opens in a new tab)

  11. Crowdsourcing a text corpus for a low resource language

    … texts, making it challenging to build language corpora and the information retrieval services, such as search and translation that depend on them. Researchers have been unable to assemble isiXhosa corpora of sufficient size and quality to produce working machine translation systems and it has …

    cape-town Repository record for Crowdsourcing a text corpus for a low resource language (opens in a new tab)

  12. Self-supervised text sentiment transfer with rationale predictions and pretrained transformers

    … methods are inadequate due to the dearth of parallel corpora for sentiment transfer. Thus, sentiment transfer can be posed as an unsupervised learning problem where a model must learn to transfer from one sentiment to another in the absence of parallel sentences. Given that the sentiment of a …

    cape-town Repository record for Self-supervised text sentiment transfer with rationale predictions and pretrained transformers (opens in a new tab)

  13. Does it have to be trees? : Data-driven dependency parsing with incomplete and noisy training data

    … from one language to another within a parallel corpus. However, the output tends to be noisy and incomplete due to cross-lingual non-parallelism and error-prone word alignments. This makes the projected annotations a suitable test bed for our fragment parsers. Our results show that (i) …

    potsdam-diss Repository record for Does it have to be trees? : Data-driven dependency parsing with incomplete and noisy training data (opens in a new tab)