Back to search

Publikationsserver der RWTH Aachen University

Improving statistical machine translation using morpho-syntactic information

Abstract

dc:description

In the framework of statistical machine translation (SMT), correspondences between the words in the source and the target language are learned from bilingual corpora, and often little or no linguistic knowledge is used to structure the underlying models. The work presented in this thesis is motivated by the observation that training data typically does not sufficiently represent the phenomena in natural languages. Various methods of incorporating morphological and syntactic information into SMT systems are proposed and systematically assessed. The overall goal is to improve translation quality and to reduce the amount of parallel text necessary to train the model parameters. Large differences in word order between corresponding sentences are difficult to capture for automatic alignment algorithms. In this work, sentence level restructuring transformations are introduced which are motivated by the sentence structure in the involved languages. These transformations aim at the assimilation of word orders in related sentences. Their application results in better alignments and as a consequence in cleaner probabilistic lexica, broader applicability of multi-word phrase pairs and a better coverage of the language model. Existing SMT systems often treat different inflected forms of the same lemma as if they were independent of each other. A better exploitation of the bilingual training data can be achieved by taking into account the interdependencies of related inflected forms. In this work a hierarchy of equivalence classes is defined on the basis of morphological and syntactic information. Features from those hierarchy levels are combined to form hierarchical lexicon models which can replace the standard lexicon used in most SMT systems. The benefit from these combined models is twofold: Firstly, the translation of unseen word forms can be derived by considering information from lower levels in the hierarchy. Secondly, category ambiguity can be resolved, because syntactical context information is made locally accessible. It is a costly and time consuming task to gather large texts and have them translated to form bilingual corpora. In this work the amount of bilingual data required to achieve an acceptable translation quality is systematically investigated. All the methods presented in this thesis contribute to a better exploitation of the available data and thus to improving translation quality in frameworks with scarce resources. The combination of the suggested methods results in substantial improvements on tasks from the projects Verbmobil, Nespole and EuTrans, for German to English and English to German translation and for text and speech input. The second focus of this thesis is on evaluation of MT quality. A tool for the evaluation of translation quality which accounts for the requirements in a research environment is developed. Evaluation criteria which are more adequate than edit distance are defined. The measurement along these quality criteria is performed semi-automatically in a fast, convenient and consistent way using the graphical user interface. The quality criteria themselves are systematically assessed.

Degree

thesis:*
Grantor dc:publisher
Publikationsserver der RWTH Aachen University
Year dc:date
2002

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Nießen, Sonja
Contributors dc:contributor
  • Ney, Hermann

Subjects

dc:subject × 7

Rights

dc:rights
Statement dc:rights
  • info:eu-repo/semantics/openAccess
Language dc:language
eng

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:publications.rwth-aachen.de:58844

Chain of custody

source
Harvested from
RWTH Aachen University
Base URL
publications.rwth-aachen.de/oai2d
Last updated
2026-07-30
Source record
OAI-PMH GetRecord
citation

Nießen, Sonja. Improving statistical machine translation using morpho-syntactic information. Publikationsserver der RWTH Aachen University, 2002. https://publications.rwth-aachen.de/record/58844