Back to search

Publikationsserver der RWTH Aachen University

Statistical machine translation with cascaded probabilistic transducers

Abstract

dc:description

Statistical machine translation is based on the idea to extract information from bilingual corpora, which can be used to generate new translations. The current work combines aspects from example-based machine translation and from grammar-based approaches, esp. bilingual regular grammars, to develop a statistical translation system based on cascaded transducers. These transducers can be constructed manually, semi-automatically, or – in restricted form – fully automatically. A training method for these cascaded transducers is developed based on an extension of the HMM alignment model to the alignment of graphs. To generate new translations using the trained models a decoder is needed. This is essentially a search for the translation with the highest probability. A decoder had been developed which is based on Dynamic Programming and which allows for pruning to control runtime. Recombination of hypotheses can be based on different criteria: coverage of the source word positions, the most recent target words, the number of generated target words, and any combination thereof. Additional aspects covered in this dissertation include:1. Segmentation of long sentences based on minimizing the perplexity of the underlying word alignment models.2. This technique is then extended into a new and robust phrase alignment. To find the target phrase for a given phrase in a source sentence the algorithm searches for the segmentation of the target sentence, which gives the highest word alignment probability under the constraints of the segmentation.3. The use and integration of manual dictionaries, including the addition of automatically generated word forms for which probabilities are estimated from the bilingual corpora.Experiments are described in which these different methods had been tested. Corpora of different sizes and for different language pairs are used. Cascaded transducers are tested esp. for small corpora, while the word-based phrase alignment are applied to large corpora. In addition – and for the situation of very restricted bilingual data – a comparison is done between the statistical translation approach and an Interlingua-based translation system, and it is shown that even in this scenario statistical translation can give comparable translation quality.

Degree

thesis:*
Grantor dc:publisher
Publikationsserver der RWTH Aachen University
Year dc:date
2005

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Vogel, Stephan
Contributors dc:contributor
  • Ney, Hermann

Subjects

dc:subject × 13

Rights

dc:rights
Statement dc:rights
  • info:eu-repo/semantics/openAccess
Language dc:language
eng

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:publications.rwth-aachen.de:62303

Chain of custody

source
Harvested from
RWTH Aachen University
Base URL
publications.rwth-aachen.de/oai2d
Last updated
2026-07-30
Source record
OAI-PMH GetRecord
citation

Vogel, Stephan. Statistical machine translation with cascaded probabilistic transducers. Publikationsserver der RWTH Aachen University, 2005. https://publications.rwth-aachen.de/record/62303