Back to results

Publikationsserver der RWTH Aachen University

Phrase based statistical machine translation : models, search, training

Abstract

dc:description

Machine translation is the task of automatically translating a text from one natural language into another. In this work, we describe and analyze the phrase-based approach to statistical machine translation. In any statistical approach to machine translation, we have to address three problems: the modeling problem, i.e. how to structure the dependencies of source and target language sentences; the search problem, i.e. how to find the best translation candidate among all possible target language sentences; the training problem, i.e. how to estimate the free parameters of the model from the training data. We present improved alignment and translation models. We present alignment models which improve the alignment quality significantly. We extend the standard IBM word alignment models with a symmetric lexicon and we describe the corresponding training procedure. Furthermore, we reduce the word alignment problem to the minimum-weight edge cover problem and present an efficient algorithm to solve this problem. Both approaches result in significantly reduced word alignment error rates on the German-English Verbmobil and French-English Canadian Hansards tasks. We describe several phrase translation models and analyze their contribution to the overall translation quality. We formulate the search problem for phrase-based statistical machine translation and present different search algorithms in detail. The search problem consists of two main components: lexical choice, i.e. selecting the correct target language words, and reordering, i.e. selecting the correct word order of the target language sentence. We analyze the search and show that it is important to focus on alternative reorderings, whereas on the other hand, already a small number of lexical alternatives are sufficient to achieve good translation quality. The reordering problem in machine translation is difficult for two reasons: first, it is computationally expensive to explore all possible permutations; second, it is hard to select a good permutation. We compare different reordering constraints to solve this problem efficiently and introduce a lexicalized reordering model to find better reorderings. Standard phrase-based machine translation systems store the whole phrase table in memory. This requires large amounts of memory. A common approach is to filter the phrase table for a specific text. This is a time-consuming task and it has to be repeated whenever a new text has to be translated. We introduce a phrase-table representation that is loaded on-demand from disk. This enables online translation, i.e. the translation of arbitrary text without time-consuming pre-filtering of the phrase table. Additionally, the memory requirements are reduced by two to three orders of a magnitude. The implementation was made available as part of a public open-source toolkit. Often, machine translation is a component of a larger pipeline; then, the input to the machine translation component is the output of another natural language processing tool. A typical example is a speech translation system where the input to the machine translation system is the output of an automatic speech recognizer. As the output of the speech recognizer may be erroneous, it is preferable to take multiple alternatives into account. On the other hand, taking a large number of input alternatives into account is computationally problematic. The alternative input sentences can be represented as a lattice. The phrase matching of input lattice and phrase table can be done efficiently by exploiting the prefix tree structure of the phrase table. Using this algorithm on the Spanish-English TC-Star task, we achieve a significant speed-up compared to previous work and generate translations of better quality. We investigate alternative training criteria for phrase-based statistical machine translation. We compare the common minimum error rate training with maximum likelihood training as well as a maximization of the expected accuracy. In this context, we generalize the known word posterior probabilities to n-gram posterior probabilities. Additionally, we introduce a sentence length posterior probability. These can be also used in a rescoring/reranking framework. The resulting machine translation system achieves state-of-the-art performance on the large scale Chinese-English NIST task. Furthermore, the system was ranked first in the official TC-Star evaluations in 2005, 2006 and 2007 for the Chinese-English broadcast news speech translation task.

Degree

thesis:*
Grantor dc:publisher
Publikationsserver der RWTH Aachen University
Year dc:date
2008

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zens, Richard
Contributors dc:contributor
  • Ney, Hermann

Subjects

dc:subject × 6

Rights

dc:rights
Statement dc:rights
  • info:eu-repo/semantics/openAccess
Language dc:language
eng

Identifiers

dc:identifier.*

Chain of custody

source
Harvested from
RWTH Aachen University
Base URL
publications.rwth-aachen.de/oai2d
Last updated
2026-07-30
Source record
OAI-PMH GetRecord
citation

Zens, Richard. Phrase based statistical machine translation : models, search, training. Publikationsserver der RWTH Aachen University, 2008. https://publications.rwth-aachen.de/record/50042