Publikationsserver der RWTH Aachen University
Improving statistical machine translation using morpho-syntactic information
Abstract
dc:descriptionIn the framework of statistical machine translation (SMT), correspondences between the words in the source and the target language are learned from bilingual corpora, and often little or no linguistic knowledge is used to structure the underlying models. The work presented in this thesis is motivated by the observation that training data typically does not sufficiently represent the phenomena in natural languages. Various methods of incorporating morphological and syntactic information into SMT systems are proposed and systematically assessed. The overall goal is to improve translation quality and to reduce the amount of parallel text necessary to train the model parameters. Large differences in word order between corresponding sentences are difficult to capture for automatic alignment algorithms. In this work, sentence level restructuring transformations are introduced which are motivated by the sentence structure in the involved languages. These transformations aim at the assimilation of word orders in related sentences. Their application results in better alignments and as a consequence in cleaner probabilistic lexica, broader applicability of multi-word phrase pairs and a better coverage of the language model. Existing SMT systems often treat different inflected forms of the same lemma as if they were independent of each other. A better exploitation of the bilingual training data can be achieved by taking into account the interdependencies of related inflected forms. In this work a hierarchy of equivalence classes is defined on the basis of morphological and syntactic information. Features from those hierarchy levels are combined to form hierarchical lexicon models which can replace the standard lexicon used in most SMT systems. The benefit from these combined models is twofold: Firstly, the translation of unseen word forms can be derived by considering information from lower levels in the hierarchy. Secondly, category ambiguity can be resolved, because syntactical context information is made locally accessible. It is a costly and time consuming task to gather large texts and have them translated to form bilingual corpora. In this work the amount of bilingual data required to achieve an acceptable translation quality is systematically investigated. All the methods presented in this thesis contribute to a better exploitation of the available data and thus to improving translation quality in frameworks with scarce resources. The combination of the suggested methods results in substantial improvements on tasks from the projects Verbmobil, Nespole and EuTrans, for German to English and English to German translation and for text and speech input. The second focus of this thesis is on evaluation of MT quality. A tool for the evaluation of translation quality which accounts for the requirements in a research environment is developed. Evaluation criteria which are more adequate than edit distance are defined. The measurement along these quality criteria is performed semi-automatically in a fast, convenient and consistent way using the graphical user interface. The quality criteria themselves are systematically assessed.
Degree
thesis:*- Grantor dc:publisher
- Publikationsserver der RWTH Aachen University
- Year dc:date
- 2002
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Nießen, Sonja
- Contributors dc:contributor
-
- Ney, Hermann
Subjects
dc:subject × 7Rights
dc:rights- Statement dc:rights
-
- info:eu-repo/semantics/openAccess
- Language dc:language
- eng
Identifiers
dc:identifier.*- OAI identifier oai:identifier
- oai:publications.rwth-aachen.de:58844