{"id":{"repo_id":"aachen","oai_identifier":"oai:publications.rwth-aachen.de:59767"},"canonical_url":"https://search.dev.ndltd.org/etd/aachen/oai:publications.rwth-aachen.de:59767","repository":{"repo_id":"aachen","name":"RWTH Aachen University","base_url":"https://publications.rwth-aachen.de/oai2d"},"display":{"title":"Statistische Auswahl von Wortabhängigkeiten in der automatischen Spracherkennung","abstract":"This PhD thesis studies the overall effect of statistical language modeling on perplexity and word error rate in automatic speech recognition. A trigram language model with a standard smoothing method is extended by complex state-of-the-art language modeling techniques, including: the comparison of different smoothing methods, namely linear vs. absolute discounting, interpolation vs. backing-off, and back-off functions based on relative frequencies vs. singleton events, the effect of complex language model techniques by using distant trigrams and word classes, word phrases and varigrams that are automatically selected using a maximum likelihood criterion (i.e. minimum perplexity), the overall gain of the combined application of the above techniques, as opposed to their separate assessment in past publications, and perplexity results on the Wall Street Journal corpus and perplexity and word error rate results for both the North American Business corpus (NAB) with a training text of about 240 million words and the German Verbmobil corpus.","abstract_html":"This PhD thesis studies the overall effect of statistical language modeling on perplexity and word error rate in automatic speech recognition. A trigram language model with a standard smoothing method is extended by complex state-of-the-art language modeling techniques, including: the comparison of different smoothing methods, namely linear vs. absolute discounting, interpolation vs. backing-off, and back-off functions based on relative frequencies vs. singleton events, the effect of complex language model techniques by using distant trigrams and word classes, word phrases and varigrams that are automatically selected using a maximum likelihood criterion (i.e. minimum perplexity), the overall gain of the combined application of the above techniques, as opposed to their separate assessment in past publications, and perplexity results on the Wall Street Journal corpus and perplexity and word error rate results for both the North American Business corpus (NAB) with a training text of about 240 million words and the German Verbmobil corpus.","abstract_has_math":false,"creators":["Martin, Sven Carl"],"institution":"Publikationsserver der RWTH Aachen University","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Ney, Hermann"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2000,"date_issued":"2000","date_published":"2000","updated_at":"2026-07-30T19:42:48Z","subjects":["info:eu-repo/classification/ddc/004","Informatik","Automatische Spracherkennung","Wortstellung","A-priori-Verteilung","Maximum-Likelihood-Schätzung"],"languages":["ger"],"rights":["info:eu-repo/semantics/openAccess"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-121521%22"],"render_values":[{"text":"https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-121521%22","href":"https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-121521%22","code":true}]}]},"links":{"outbound_url":"https://publications.rwth-aachen.de/record/59767","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ney, Hermann"]},{"key":"dc:creator","label":"Author","values":["Martin, Sven Carl"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:coverage","label":"Dc Coverage","values":["DE"]},{"key":"dc:date","label":"Dc Date","values":["2000"]},{"key":"dc:publisher","label":"Institution","values":["Publikationsserver der RWTH Aachen University"]},{"key":"dc:relation","label":"Dc Relation","values":["info:eu-repo/semantics/altIdentifier/urn/urn:nbn:de:hbz:82-opus-503"]},{"key":"dc:type","label":"Dc Type","values":["info:eu-repo/semantics/doctoralThesis","info:eu-repo/semantics/publishedVersion"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["info:eu-repo/classification/ddc/004","Informatik","Automatische Spracherkennung","Wortstellung","A-priori-Verteilung","Maximum-Likelihood-Schätzung"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["ger"]},{"key":"dc:rights","label":"Dc Rights","values":["info:eu-repo/semantics/openAccess"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://publications.rwth-aachen.de/record/59767","https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-121521%22"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This PhD thesis studies the overall effect of statistical language modeling on perplexity and word error rate in automatic speech recognition. A trigram language model with a standard smoothing method is extended by complex state-of-the-art language modeling techniques, including: the comparison of different smoothing methods, namely linear vs. absolute discounting, interpolation vs. backing-off, and back-off functions based on relative frequencies vs. singleton events, the effect of complex language model techniques by using distant trigrams and word classes, word phrases and varigrams that are automatically selected using a maximum likelihood criterion (i.e. minimum perplexity), the overall gain of the combined application of the above techniques, as opposed to their separate assessment in past publications, and perplexity results on the Wall Street Journal corpus and perplexity and word error rate results for both the North American Business corpus (NAB) with a training text of about 240 million words and the German Verbmobil corpus."]},{"key":"dc:source","label":"Dc Source","values":["Aachen : Publikationsserver der RWTH Aachen University XII, 132 S. (2000). = Aachen, Techn. Hochsch., Diss., 2000"]},{"key":"dc:title","label":"Title","values":["Statistische Auswahl von Wortabhängigkeiten in der automatischen Spracherkennung"]}]}],"canonical_facts":{"dc:contributor":["Ney, Hermann"],"dc:coverage":["DE"],"dc:creator":["Martin, Sven Carl"],"dc:date":["2000"],"dc:description":["This PhD thesis studies the overall effect of statistical language modeling on perplexity and word error rate in automatic speech recognition. A trigram language model with a standard smoothing method is extended by complex state-of-the-art language modeling techniques, including: the comparison of different smoothing methods, namely linear vs. absolute discounting, interpolation vs. backing-off, and back-off functions based on relative frequencies vs. singleton events, the effect of complex language model techniques by using distant trigrams and word classes, word phrases and varigrams that are automatically selected using a maximum likelihood criterion (i.e. minimum perplexity), the overall gain of the combined application of the above techniques, as opposed to their separate assessment in past publications, and perplexity results on the Wall Street Journal corpus and perplexity and word error rate results for both the North American Business corpus (NAB) with a training text of about 240 million words and the German Verbmobil corpus."],"dc:identifier":["https://publications.rwth-aachen.de/record/59767","https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-121521%22"],"dc:language":["ger"],"dc:publisher":["Publikationsserver der RWTH Aachen University"],"dc:relation":["info:eu-repo/semantics/altIdentifier/urn/urn:nbn:de:hbz:82-opus-503"],"dc:rights":["info:eu-repo/semantics/openAccess"],"dc:source":["Aachen : Publikationsserver der RWTH Aachen University XII, 132 S. (2000). = Aachen, Techn. Hochsch., Diss., 2000"],"dc:subject":["info:eu-repo/classification/ddc/004","Informatik","Automatische Spracherkennung","Wortstellung","A-priori-Verteilung","Maximum-Likelihood-Schätzung"],"dc:title":["Statistische Auswahl von Wortabhängigkeiten in der automatischen Spracherkennung"],"dc:type":["info:eu-repo/semantics/doctoralThesis","info:eu-repo/semantics/publishedVersion"]},"updated_at":"2026-07-30T19:42:48Z"}