{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/16031"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/16031","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Estimation problems in speech and natural language","abstract":"This dissertation is a study of two problems on estimation in the areas of natural language and speech. In the first problem we revisit the classical problem of estimating the size of unseen elements which we study in the context of a regime that is characterized by a large number of rare events, natural language being one. We propose an estimator of the size of the vocabulary of the underlying population that generates an observation and show that it has theoretical guarantees of optimal performance. Using natural language corpora from different languages we show that the performance of our estimator compares favorably with that of state-of-the-art estimators. In the second problem, we explore the effect of vocabulary size and temporal aspects of speech production on perceptions of second language fluency with the aim of designing objective methods of fluency assessment from spontaneous speech. We show that articulation rate, phonation-time ratio, mean length of silent pauses and the number of silent pauses per second are aspects of speech production that are well correlated with human assigned scores of fluency. The measures of lexical use that we found to correlate well with fluency scores were the total number of words spoken (word tokens), the number of different words uttered (word types) and the number of words spoken once ({\\em hapax legomena}). With the goal of objective fluency assessment without the use of automatic speech recognition, we show the utility of measures of temporal aspects of speech production that were obtained from direct signal-level measurements. Their use in a logistic regression framework for predicting fluency scores showed high agreement with scores assigned by human raters. An interesting experiment was exploring the difference in automatic assessment based on random snippets of the spoken utterance and that based on the complete utterance. Although the differences are not seen to be statistically significant at the 1\\% level, this opens avenues for further experimentation.","abstract_html":"This dissertation is a study of two problems on estimation in the areas of natural language and speech. In the first problem we revisit the classical problem of estimating the size of unseen elements which we study in the context of a regime that is characterized by a large number of rare events, natural language being one. We propose an estimator of the size of the vocabulary of the underlying population that generates an observation and show that it has theoretical guarantees of optimal performance. Using natural language corpora from different languages we show that the performance of our estimator compares favorably with that of state-of-the-art estimators. In the second problem, we explore the effect of vocabulary size and temporal aspects of speech production on perceptions of second language fluency with the aim of designing objective methods of fluency assessment from spontaneous speech. We show that articulation rate, phonation-time ratio, mean length of silent pauses and the number of silent pauses per second are aspects of speech production that are well correlated with human assigned scores of fluency. The measures of lexical use that we found to correlate well with fluency scores were the total number of words spoken (word tokens), the number of different words uttered (word types) and the number of words spoken once ({\\em hapax legomena}). With the goal of objective fluency assessment without the use of automatic speech recognition, we show the utility of measures of temporal aspects of speech production that were obtained from direct signal-level measurements. Their use in a logistic regression framework for predicting fluency scores showed high agreement with scores assigned by human raters. An interesting experiment was exploring the difference in automatic assessment based on random snippets of the spoken utterance and that based on the complete utterance. Although the differences are not seen to be statistically significant at the 1\\% level, this opens avenues for further experimentation.","abstract_has_math":false,"creators":["Bhat, Suma P."],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Sproat, Richard W.","Church, Kenneth W.","Hasegawa-Johnson, Mark A.","Roth, Dan","Levinson, Stephen E."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2010,"date_issued":"2010-05-19T18:32:47Z","date_published":"2010-05-19T18:32:47Z","updated_at":"2026-07-22T22:25:08Z","subjects":["Vocabulary Size Estimation","Estimation of the number of unseen events","Automatic Fluency Assessment","Variable Selection","Predictors of Oral Fluency","Click-Through Rate Prediction"],"languages":["en"],"rights":["Copyright 2010 Suma P. Bhat"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/16031","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Sproat, Richard W.","Church, Kenneth W.","Hasegawa-Johnson, Mark A.","Roth, Dan","Levinson, Stephen E."]},{"key":"dc:creator","label":"Author","values":["Bhat, Suma P."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2010-05-19T18:32:47Z","2010-05"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Vocabulary Size Estimation","Estimation of the number of unseen events","Automatic Fluency Assessment","Variable Selection","Predictors of Oral Fluency","Click-Through Rate Prediction"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2010 Suma P. Bhat"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/16031"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This dissertation is a study of two problems on estimation in the areas of natural language and speech. In the first problem we revisit the classical problem of estimating the size of unseen elements which we study in the context of a regime that is characterized by a large number of rare events, natural language being one. We propose an estimator of the size of the vocabulary of the underlying population that generates an observation and show that it has theoretical guarantees of optimal performance. Using natural language corpora from different languages we show that the performance of our estimator compares favorably with that of state-of-the-art estimators. In the second problem, we explore the effect of vocabulary size and temporal aspects of speech production on perceptions of second language fluency with the aim of designing objective methods of fluency assessment from spontaneous speech. We show that articulation rate, phonation-time ratio, mean length of silent pauses and the number of silent pauses per second are aspects of speech production that are well correlated with human assigned scores of fluency. The measures of lexical use that we found to correlate well with fluency scores were the total number of words spoken (word tokens), the number of different words uttered (word types) and the number of words spoken once ({\\em hapax legomena}). With the goal of objective fluency assessment without the use of automatic speech recognition, we show the utility of measures of temporal aspects of speech production that were obtained from direct signal-level measurements. Their use in a logistic regression framework for predicting fluency scores showed high agreement with scores assigned by human raters. An interesting experiment was exploring the difference in automatic assessment based on random snippets of the spoken utterance and that based on the complete utterance. Although the differences are not seen to be statistically significant at the 1\\% level, this opens avenues for further experimentation.","Item withdrawn by Alexis Thompson (athmpsn1@illinois.edu) on 2010-04-21T23:08:39Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Bhat_Suma.pdf: 3751250 bytes, checksum: 323b1e6ac5ddd32dccab593378b52e66 (MD5)","Made available in DSpace on 2010-05-19T18:32:47Z (GMT). No. of bitstreams: 2 Bhat_Suma.pdf: 3751250 bytes, checksum: 323b1e6ac5ddd32dccab593378b52e66 (MD5) license.txt: 4058 bytes, checksum: 119d6f50798d8f039b23ba854504f216 (MD5)"]},{"key":"dc:title","label":"Title","values":["Estimation problems in speech and natural language"]}]}],"canonical_facts":{"dc:contributor":["Sproat, Richard W.","Church, Kenneth W.","Hasegawa-Johnson, Mark A.","Roth, Dan","Levinson, Stephen E."],"dc:creator":["Bhat, Suma P."],"dc:date":["2010-05-19T18:32:47Z","2010-05"],"dc:description":["This dissertation is a study of two problems on estimation in the areas of natural language and speech. In the first problem we revisit the classical problem of estimating the size of unseen elements which we study in the context of a regime that is characterized by a large number of rare events, natural language being one. We propose an estimator of the size of the vocabulary of the underlying population that generates an observation and show that it has theoretical guarantees of optimal performance. Using natural language corpora from different languages we show that the performance of our estimator compares favorably with that of state-of-the-art estimators. In the second problem, we explore the effect of vocabulary size and temporal aspects of speech production on perceptions of second language fluency with the aim of designing objective methods of fluency assessment from spontaneous speech. We show that articulation rate, phonation-time ratio, mean length of silent pauses and the number of silent pauses per second are aspects of speech production that are well correlated with human assigned scores of fluency. The measures of lexical use that we found to correlate well with fluency scores were the total number of words spoken (word tokens), the number of different words uttered (word types) and the number of words spoken once ({\\em hapax legomena}). With the goal of objective fluency assessment without the use of automatic speech recognition, we show the utility of measures of temporal aspects of speech production that were obtained from direct signal-level measurements. Their use in a logistic regression framework for predicting fluency scores showed high agreement with scores assigned by human raters. An interesting experiment was exploring the difference in automatic assessment based on random snippets of the spoken utterance and that based on the complete utterance. Although the differences are not seen to be statistically significant at the 1\\% level, this opens avenues for further experimentation.","Item withdrawn by Alexis Thompson (athmpsn1@illinois.edu) on 2010-04-21T23:08:39Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Bhat_Suma.pdf: 3751250 bytes, checksum: 323b1e6ac5ddd32dccab593378b52e66 (MD5)","Made available in DSpace on 2010-05-19T18:32:47Z (GMT). No. of bitstreams: 2 Bhat_Suma.pdf: 3751250 bytes, checksum: 323b1e6ac5ddd32dccab593378b52e66 (MD5) license.txt: 4058 bytes, checksum: 119d6f50798d8f039b23ba854504f216 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/16031"],"dc:language":["en"],"dc:rights":["Copyright 2010 Suma P. Bhat"],"dc:subject":["Vocabulary Size Estimation","Estimation of the number of unseen events","Automatic Fluency Assessment","Variable Selection","Predictors of Oral Fluency","Click-Through Rate Prediction"],"dc:title":["Estimation problems in speech and natural language"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:08Z"}