{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105767"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105767","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Modeling phones, keywords, topics and intents in spoken languages","abstract":"Spoken Language Understanding for both rich-resource languages (RRL) and low-resource languages (LRL) is an important research area for academia and the commercial world. In the conversational situations where either the language used in speech is a minority one, or the environment is noisy, barriers will emerge between the communicators. Essentially, people would like to understand the basic components of any language spoken by others who they meet in their daily lives. On the other hand, machines can also be trained to learn the process of modeling the basic language components such as phones, keywords, topics and intents during both human/machine interactions and human/human communications. Eventually, if we can develop a machine assistant for people to understand the basic meaning of any language in speech, we could make the human world much more efficient and harmonious. This thesis addresses the problem with the help of mismatched-crowdsourcing- based distant supervision, linguistic knowledge, and corpus-based transfer learning. First we analyze the usefulness of mismatched transcripts and distinctive features, and then propose phone recognition based on the optimized inference of the phone set in the low-resource language from the clustering of the mismatched transcripts. Subsequently, the keyword discovery from the phone-level results is explored. The topic information collected in the corpus is then used as the additional knowledge for topic classification and further improving phone recognition. Based on the keyword sequence, the intents of the speaker are also eventually obtained. The experimental results show that with the help of data collection design and existing knowledge, we can achieve reasonably good machine language understanding for languages whose phones, keywords, topics, and intents were not learned before. This work will lead to further investigations in the area of spoken language understanding in any language.","abstract_html":"Spoken Language Understanding for both rich-resource languages (RRL) and low-resource languages (LRL) is an important research area for academia and the commercial world. In the conversational situations where either the language used in speech is a minority one, or the environment is noisy, barriers will emerge between the communicators. Essentially, people would like to understand the basic components of any language spoken by others who they meet in their daily lives. On the other hand, machines can also be trained to learn the process of modeling the basic language components such as phones, keywords, topics and intents during both human/machine interactions and human/human communications. Eventually, if we can develop a machine assistant for people to understand the basic meaning of any language in speech, we could make the human world much more efficient and harmonious. This thesis addresses the problem with the help of mismatched-crowdsourcing- based distant supervision, linguistic knowledge, and corpus-based transfer learning. First we analyze the usefulness of mismatched transcripts and distinctive features, and then propose phone recognition based on the optimized inference of the phone set in the low-resource language from the clustering of the mismatched transcripts. Subsequently, the keyword discovery from the phone-level results is explored. The topic information collected in the corpus is then used as the additional knowledge for topic classification and further improving phone recognition. Based on the keyword sequence, the intents of the speaker are also eventually obtained. The experimental results show that with the help of data collection design and existing knowledge, we can achieve reasonably good machine language understanding for languages whose phones, keywords, topics, and intents were not learned before. This work will lead to further investigations in the area of spoken language understanding in any language.","abstract_has_math":false,"creators":["Chen, Wenda"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark","Li, Haizhou","Levinson, Stephen E.","Varshney, Lav R."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-11-26T20:49:16Z","date_published":"2019-11-26T20:49:16Z","updated_at":"2026-07-22T22:24:45Z","subjects":["Clustering","Distant Supervision, Spoken language understanding","Spoken term detection","Speech recognition","Low-resource languages","Transfer learning"],"languages":["en"],"rights":["Copyright 2019 Wenda Chen"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105767","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark","Li, Haizhou","Levinson, Stephen E.","Varshney, Lav R."]},{"key":"dc:creator","label":"Author","values":["Chen, Wenda"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-11-26T20:49:16Z","2021-11-27T10:15:23Z","2019-07-02","2019-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Clustering","Distant Supervision, Spoken language understanding","Spoken term detection","Speech recognition","Low-resource languages","Transfer learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Wenda Chen"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105767"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Spoken Language Understanding for both rich-resource languages (RRL) and low-resource languages (LRL) is an important research area for academia and the commercial world. In the conversational situations where either the language used in speech is a minority one, or the environment is noisy, barriers will emerge between the communicators. Essentially, people would like to understand the basic components of any language spoken by others who they meet in their daily lives. On the other hand, machines can also be trained to learn the process of modeling the basic language components such as phones, keywords, topics and intents during both human/machine interactions and human/human communications. Eventually, if we can develop a machine assistant for people to understand the basic meaning of any language in speech, we could make the human world much more efficient and harmonious. This thesis addresses the problem with the help of mismatched-crowdsourcing- based distant supervision, linguistic knowledge, and corpus-based transfer learning. First we analyze the usefulness of mismatched transcripts and distinctive features, and then propose phone recognition based on the optimized inference of the phone set in the low-resource language from the clustering of the mismatched transcripts. Subsequently, the keyword discovery from the phone-level results is explored. The topic information collected in the corpus is then used as the additional knowledge for topic classification and further improving phone recognition. Based on the keyword sequence, the intents of the speaker are also eventually obtained. The experimental results show that with the help of data collection design and existing knowledge, we can achieve reasonably good machine language understanding for languages whose phones, keywords, topics, and intents were not learned before. This work will lead to further investigations in the area of spoken language understanding in any language.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-08-01","The student, Wenda Chen, accepted the attached license on 2019-07-01 at 09:42.","The student, Wenda Chen, submitted this Dissertation for approval on 2019-07-01 at 17:16.","This Dissertation was approved for publication on 2019-07-02 at 17:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14097 on 2019-11-26 at 13:03:56","Made available in DSpace on 2019-11-26T20:49:16Z (GMT). No. of bitstreams: 4 CHEN-DISSERTATION-2019.pdf: 3650605 bytes, checksum: 8ace7c7d743f553bf04cd14b593f0a1b (MD5) LICENSE.txt: 4207 bytes, checksum: 010c1fe8495035a0726e697086da8473 (MD5) PROQUEST_LICENSE.txt: 4553 bytes, checksum: 9198c0209e12202a92007e3363b0be92 (MD5) letter.pdf: 16053 bytes, checksum: d47579f1873c3c9748b52be214c47ec8 (MD5) Previous issue date: 2019-07-02","Embargo set by: Seth Robbins for item 112912 Lift date: 2021-11-26T20:49:41Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 112912 on 2021-11-27T10:15:23Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Modeling phones, keywords, topics and intents in spoken languages"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark","Li, Haizhou","Levinson, Stephen E.","Varshney, Lav R."],"dc:creator":["Chen, Wenda"],"dc:date":["2019-11-26T20:49:16Z","2021-11-27T10:15:23Z","2019-07-02","2019-08"],"dc:description":["Spoken Language Understanding for both rich-resource languages (RRL) and low-resource languages (LRL) is an important research area for academia and the commercial world. In the conversational situations where either the language used in speech is a minority one, or the environment is noisy, barriers will emerge between the communicators. Essentially, people would like to understand the basic components of any language spoken by others who they meet in their daily lives. On the other hand, machines can also be trained to learn the process of modeling the basic language components such as phones, keywords, topics and intents during both human/machine interactions and human/human communications. Eventually, if we can develop a machine assistant for people to understand the basic meaning of any language in speech, we could make the human world much more efficient and harmonious. This thesis addresses the problem with the help of mismatched-crowdsourcing- based distant supervision, linguistic knowledge, and corpus-based transfer learning. First we analyze the usefulness of mismatched transcripts and distinctive features, and then propose phone recognition based on the optimized inference of the phone set in the low-resource language from the clustering of the mismatched transcripts. Subsequently, the keyword discovery from the phone-level results is explored. The topic information collected in the corpus is then used as the additional knowledge for topic classification and further improving phone recognition. Based on the keyword sequence, the intents of the speaker are also eventually obtained. The experimental results show that with the help of data collection design and existing knowledge, we can achieve reasonably good machine language understanding for languages whose phones, keywords, topics, and intents were not learned before. This work will lead to further investigations in the area of spoken language understanding in any language.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-08-01","The student, Wenda Chen, accepted the attached license on 2019-07-01 at 09:42.","The student, Wenda Chen, submitted this Dissertation for approval on 2019-07-01 at 17:16.","This Dissertation was approved for publication on 2019-07-02 at 17:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14097 on 2019-11-26 at 13:03:56","Made available in DSpace on 2019-11-26T20:49:16Z (GMT). No. of bitstreams: 4 CHEN-DISSERTATION-2019.pdf: 3650605 bytes, checksum: 8ace7c7d743f553bf04cd14b593f0a1b (MD5) LICENSE.txt: 4207 bytes, checksum: 010c1fe8495035a0726e697086da8473 (MD5) PROQUEST_LICENSE.txt: 4553 bytes, checksum: 9198c0209e12202a92007e3363b0be92 (MD5) letter.pdf: 16053 bytes, checksum: d47579f1873c3c9748b52be214c47ec8 (MD5) Previous issue date: 2019-07-02","Embargo set by: Seth Robbins for item 112912 Lift date: 2021-11-26T20:49:41Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 112912 on 2021-11-27T10:15:23Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105767"],"dc:language":["en"],"dc:rights":["Copyright 2019 Wenda Chen"],"dc:subject":["Clustering","Distant Supervision, Spoken language understanding","Spoken term detection","Speech recognition","Low-resource languages","Transfer learning"],"dc:title":["Modeling phones, keywords, topics and intents in spoken languages"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:45Z"}