{"id":{"repo_id":"aachen","oai_identifier":"oai:publications.rwth-aachen.de:62303"},"canonical_url":"https://search.dev.ndltd.org/etd/aachen/oai:publications.rwth-aachen.de:62303","repository":{"repo_id":"aachen","name":"RWTH Aachen University","base_url":"https://publications.rwth-aachen.de/oai2d"},"display":{"title":"Statistical machine translation with cascaded probabilistic transducers","abstract":"Statistical machine translation is based on the idea to extract information from bilingual corpora, which can be used to generate new translations. The current work combines aspects from example-based machine translation and from grammar-based approaches, esp. bilingual regular grammars, to develop a statistical translation system based on cascaded transducers. These transducers can be constructed manually, semi-automatically, or – in restricted form – fully automatically. A training method for these cascaded transducers is developed based on an extension of the HMM alignment model to the alignment of graphs. To generate new translations using the trained models a decoder is needed. This is essentially a search for the translation with the highest probability. A decoder had been developed which is based on Dynamic Programming and which allows for pruning to control runtime. Recombination of hypotheses can be based on different criteria: coverage of the source word positions, the most recent target words, the number of generated target words, and any combination thereof. Additional aspects covered in this dissertation include:1. Segmentation of long sentences based on minimizing the perplexity of the underlying word alignment models.2. This technique is then extended into a new and robust phrase alignment. To find the target phrase for a given phrase in a source sentence the algorithm searches for the segmentation of the target sentence, which gives the highest word alignment probability under the constraints of the segmentation.3. The use and integration of manual dictionaries, including the addition of automatically generated word forms for which probabilities are estimated from the bilingual corpora.Experiments are described in which these different methods had been tested. Corpora of different sizes and for different language pairs are used. Cascaded transducers are tested esp. for small corpora, while the word-based phrase alignment are applied to large corpora. In addition – and for the situation of very restricted bilingual data – a comparison is done between the statistical translation approach and an Interlingua-based translation system, and it is shown that even in this scenario statistical translation can give comparable translation quality.","abstract_html":"Statistical machine translation is based on the idea to extract information from bilingual corpora, which can be used to generate new translations. The current work combines aspects from example-based machine translation and from grammar-based approaches, esp. bilingual regular grammars, to develop a statistical translation system based on cascaded transducers. These transducers can be constructed manually, semi-automatically, or – in restricted form – fully automatically. A training method for these cascaded transducers is developed based on an extension of the HMM alignment model to the alignment of graphs. To generate new translations using the trained models a decoder is needed. This is essentially a search for the translation with the highest probability. A decoder had been developed which is based on Dynamic Programming and which allows for pruning to control runtime. Recombination of hypotheses can be based on different criteria: coverage of the source word positions, the most recent target words, the number of generated target words, and any combination thereof. Additional aspects covered in this dissertation include:1. Segmentation of long sentences based on minimizing the perplexity of the underlying word alignment models.2. This technique is then extended into a new and robust phrase alignment. To find the target phrase for a given phrase in a source sentence the algorithm searches for the segmentation of the target sentence, which gives the highest word alignment probability under the constraints of the segmentation.3. The use and integration of manual dictionaries, including the addition of automatically generated word forms for which probabilities are estimated from the bilingual corpora.Experiments are described in which these different methods had been tested. Corpora of different sizes and for different language pairs are used. Cascaded transducers are tested esp. for small corpora, while the word-based phrase alignment are applied to large corpora. In addition – and for the situation of very restricted bilingual data – a comparison is done between the statistical translation approach and an Interlingua-based translation system, and it is shown that even in this scenario statistical translation can give comparable translation quality.","abstract_has_math":false,"creators":["Vogel, Stephan"],"institution":"Publikationsserver der RWTH Aachen University","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Ney, Hermann"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2005,"date_issued":"2005","date_published":"2005","updated_at":"2026-07-30T19:43:28Z","subjects":["info:eu-repo/classification/ddc/004","Automatische Übersetzung","Informatik","Maschinelle Übersetzung","statistische Übersetzung","hierarchisches Alignment","Phrasen-Alignment","Suche","Statistical machine translation","word alignment","phrase alignment","cascaded transducer","decoder"],"languages":["eng"],"rights":["info:eu-repo/semantics/openAccess"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-123878%22"],"render_values":[{"text":"https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-123878%22","href":"https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-123878%22","code":true}]}]},"links":{"outbound_url":"https://publications.rwth-aachen.de/record/62303","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ney, Hermann"]},{"key":"dc:creator","label":"Author","values":["Vogel, Stephan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:coverage","label":"Dc Coverage","values":["DE"]},{"key":"dc:date","label":"Dc Date","values":["2005"]},{"key":"dc:publisher","label":"Institution","values":["Publikationsserver der RWTH Aachen University"]},{"key":"dc:relation","label":"Dc Relation","values":["info:eu-repo/semantics/altIdentifier/urn/urn:nbn:de:hbz:82-opus-19044"]},{"key":"dc:type","label":"Dc Type","values":["info:eu-repo/semantics/doctoralThesis","info:eu-repo/semantics/publishedVersion"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["info:eu-repo/classification/ddc/004","Automatische Übersetzung","Informatik","Maschinelle Übersetzung","statistische Übersetzung","hierarchisches Alignment","Phrasen-Alignment","Suche","Statistical machine translation","word alignment","phrase alignment","cascaded transducer","decoder"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["info:eu-repo/semantics/openAccess"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://publications.rwth-aachen.de/record/62303","https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-123878%22"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Statistical machine translation is based on the idea to extract information from bilingual corpora, which can be used to generate new translations. The current work combines aspects from example-based machine translation and from grammar-based approaches, esp. bilingual regular grammars, to develop a statistical translation system based on cascaded transducers. These transducers can be constructed manually, semi-automatically, or – in restricted form – fully automatically. A training method for these cascaded transducers is developed based on an extension of the HMM alignment model to the alignment of graphs. To generate new translations using the trained models a decoder is needed. This is essentially a search for the translation with the highest probability. A decoder had been developed which is based on Dynamic Programming and which allows for pruning to control runtime. Recombination of hypotheses can be based on different criteria: coverage of the source word positions, the most recent target words, the number of generated target words, and any combination thereof. Additional aspects covered in this dissertation include:1. Segmentation of long sentences based on minimizing the perplexity of the underlying word alignment models.2. This technique is then extended into a new and robust phrase alignment. To find the target phrase for a given phrase in a source sentence the algorithm searches for the segmentation of the target sentence, which gives the highest word alignment probability under the constraints of the segmentation.3. The use and integration of manual dictionaries, including the addition of automatically generated word forms for which probabilities are estimated from the bilingual corpora.Experiments are described in which these different methods had been tested. Corpora of different sizes and for different language pairs are used. Cascaded transducers are tested esp. for small corpora, while the word-based phrase alignment are applied to large corpora. In addition – and for the situation of very restricted bilingual data – a comparison is done between the statistical translation approach and an Interlingua-based translation system, and it is shown that even in this scenario statistical translation can give comparable translation quality."]},{"key":"dc:source","label":"Dc Source","values":["Aachen : Publikationsserver der RWTH Aachen University XII, 123 S. : graph. Darst. (2005). = Aachen, Techn. Hochsch., Diss., 2005"]},{"key":"dc:title","label":"Title","values":["Statistical machine translation with cascaded probabilistic transducers"]}]}],"canonical_facts":{"dc:contributor":["Ney, Hermann"],"dc:coverage":["DE"],"dc:creator":["Vogel, Stephan"],"dc:date":["2005"],"dc:description":["Statistical machine translation is based on the idea to extract information from bilingual corpora, which can be used to generate new translations. The current work combines aspects from example-based machine translation and from grammar-based approaches, esp. bilingual regular grammars, to develop a statistical translation system based on cascaded transducers. These transducers can be constructed manually, semi-automatically, or – in restricted form – fully automatically. A training method for these cascaded transducers is developed based on an extension of the HMM alignment model to the alignment of graphs. To generate new translations using the trained models a decoder is needed. This is essentially a search for the translation with the highest probability. A decoder had been developed which is based on Dynamic Programming and which allows for pruning to control runtime. Recombination of hypotheses can be based on different criteria: coverage of the source word positions, the most recent target words, the number of generated target words, and any combination thereof. Additional aspects covered in this dissertation include:1. Segmentation of long sentences based on minimizing the perplexity of the underlying word alignment models.2. This technique is then extended into a new and robust phrase alignment. To find the target phrase for a given phrase in a source sentence the algorithm searches for the segmentation of the target sentence, which gives the highest word alignment probability under the constraints of the segmentation.3. The use and integration of manual dictionaries, including the addition of automatically generated word forms for which probabilities are estimated from the bilingual corpora.Experiments are described in which these different methods had been tested. Corpora of different sizes and for different language pairs are used. Cascaded transducers are tested esp. for small corpora, while the word-based phrase alignment are applied to large corpora. In addition – and for the situation of very restricted bilingual data – a comparison is done between the statistical translation approach and an Interlingua-based translation system, and it is shown that even in this scenario statistical translation can give comparable translation quality."],"dc:identifier":["https://publications.rwth-aachen.de/record/62303","https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-123878%22"],"dc:language":["eng"],"dc:publisher":["Publikationsserver der RWTH Aachen University"],"dc:relation":["info:eu-repo/semantics/altIdentifier/urn/urn:nbn:de:hbz:82-opus-19044"],"dc:rights":["info:eu-repo/semantics/openAccess"],"dc:source":["Aachen : Publikationsserver der RWTH Aachen University XII, 123 S. : graph. Darst. (2005). = Aachen, Techn. Hochsch., Diss., 2005"],"dc:subject":["info:eu-repo/classification/ddc/004","Automatische Übersetzung","Informatik","Maschinelle Übersetzung","statistische Übersetzung","hierarchisches Alignment","Phrasen-Alignment","Suche","Statistical machine translation","word alignment","phrase alignment","cascaded transducer","decoder"],"dc:title":["Statistical machine translation with cascaded probabilistic transducers"],"dc:type":["info:eu-repo/semantics/doctoralThesis","info:eu-repo/semantics/publishedVersion"]},"updated_at":"2026-07-30T19:43:28Z"}