Publikationsserver der RWTH Aachen University
Statistische Auswahl von Wortabhängigkeiten in der automatischen Spracherkennung
Abstract
dc:descriptionThis PhD thesis studies the overall effect of statistical language modeling on perplexity and word error rate in automatic speech recognition. A trigram language model with a standard smoothing method is extended by complex state-of-the-art language modeling techniques, including: the comparison of different smoothing methods, namely linear vs. absolute discounting, interpolation vs. backing-off, and back-off functions based on relative frequencies vs. singleton events, the effect of complex language model techniques by using distant trigrams and word classes, word phrases and varigrams that are automatically selected using a maximum likelihood criterion (i.e. minimum perplexity), the overall gain of the combined application of the above techniques, as opposed to their separate assessment in past publications, and perplexity results on the Wall Street Journal corpus and perplexity and word error rate results for both the North American Business corpus (NAB) with a training text of about 240 million words and the German Verbmobil corpus.
Degree
thesis:*- Grantor dc:publisher
- Publikationsserver der RWTH Aachen University
- Year dc:date
- 2000
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Martin, Sven Carl
- Contributors dc:contributor
-
- Ney, Hermann
Subjects
dc:subject × 6Rights
dc:rights- Statement dc:rights
-
- info:eu-repo/semantics/openAccess
- Language dc:language
- ger
Identifiers
dc:identifier.*- OAI identifier oai:identifier
- oai:publications.rwth-aachen.de:59767