{"id":{"repo_id":"cuny-grad","oai_identifier":"oai:academicworks.cuny.edu:gc_etds-1272"},"canonical_url":"https://search.dev.ndltd.org/etd/cuny-grad/oai:academicworks.cuny.edu:gc_etds-1272","repository":{"repo_id":"cuny-grad","name":"City University of New York - Graduate Center","base_url":"https://academicworks.cuny.edu/do/oai/"},"display":{"title":"Echolocation: Using Word-Burst Analysis to Rescore Keyword Search Candidates in Low-Resource Languages","abstract":"<p>State of the art technologies for speech recognition are very accurate for heavily studied languages like English. They perform poorly, though, for languages wherein the recorded archives of speech data available to researchers are relatively scant. In the context of these low-resource languages, the task of keyword search within recorded speech is formidable. We demonstrate a method that generates more accurate keyword search results on low-resource languages by studying a pattern not exploited by the speech recognizer. The word-burst, or burstiness, pattern is the tendency for word utterances to appear together in bursts as conversational topics fluctuate. We give evidence that the burstiness phenomenon exhibits itself across varied languages. Using burstiness features to train a machine-learning algorithm, we are able to assess the likelihood that a hypothesized keyword location is correct and adjust its confidence score accordingly, yielding improvements in the efficacy of keyword search in low-resource languages. </p>","abstract_html":"&lt;p&gt;State of the art technologies for speech recognition are very accurate for heavily studied languages like English. They perform poorly, though, for languages wherein the recorded archives of speech data available to researchers are relatively scant. In the context of these low-resource languages, the task of keyword search within recorded speech is formidable. We demonstrate a method that generates more accurate keyword search results on low-resource languages by studying a pattern not exploited by the speech recognizer. The word-burst, or burstiness, pattern is the tendency for word utterances to appear together in bursts as conversational topics fluctuate. We give evidence that the burstiness phenomenon exhibits itself across varied languages. Using burstiness features to train a machine-learning algorithm, we are able to assess the likelihood that a hypothesized keyword location is correct and adjust its confidence score accordingly, yielding improvements in the efficacy of keyword search in low-resource languages. &lt;/p&gt;","abstract_has_math":false,"creators":["Richards, Justin"],"institution":"The Graduate School and University Center of The City University of New York","degree_name":"Master of Arts","degree_level":"Master","degree_discipline":"Linguistics","degree_department":null,"school":null,"contributors":[],"advisors":["Andrew Rosenberg"],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-06-03T07:00:00Z","date_published":"2014-06-03T07:00:00Z","updated_at":"2026-07-24T01:59:38Z","subjects":["Computer Sciences","Linguistics","Babel","burstiness","cache","keyword search","spoken term detection","word-burst"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://academicworks.cuny.edu/gc_etds/273","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Andrew Rosenberg"]},{"key":"dc:creator","label":"Author","values":["Richards, Justin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2014-12-03T08:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Linguistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Master"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Arts"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["The Graduate School and University Center of The City University of New York"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Sciences","Linguistics","Babel","burstiness","cache","keyword search","spoken term detection","word-burst"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://academicworks.cuny.edu/gc_etds/273"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>State of the art technologies for speech recognition are very accurate for heavily studied languages like English. They perform poorly, though, for languages wherein the recorded archives of speech data available to researchers are relatively scant. In the context of these low-resource languages, the task of keyword search within recorded speech is formidable. We demonstrate a method that generates more accurate keyword search results on low-resource languages by studying a pattern not exploited by the speech recognizer. The word-burst, or burstiness, pattern is the tendency for word utterances to appear together in bursts as conversational topics fluctuate. We give evidence that the burstiness phenomenon exhibits itself across varied languages. Using burstiness features to train a machine-learning algorithm, we are able to assess the likelihood that a hypothesized keyword location is correct and adjust its confidence score accordingly, yielding improvements in the efficacy of keyword search in low-resource languages. </p>"]},{"key":"dc:title","label":"Title","values":["Echolocation: Using Word-Burst Analysis to Rescore Keyword Search Candidates in Low-Resource Languages"]}]}],"canonical_facts":{"dc:contributor.advisor":["Andrew Rosenberg"],"dc:creator":["Richards, Justin"],"dc:date.available":["2014-12-03T08:00:00Z"],"dc:description.abstract":["<p>State of the art technologies for speech recognition are very accurate for heavily studied languages like English. They perform poorly, though, for languages wherein the recorded archives of speech data available to researchers are relatively scant. In the context of these low-resource languages, the task of keyword search within recorded speech is formidable. We demonstrate a method that generates more accurate keyword search results on low-resource languages by studying a pattern not exploited by the speech recognizer. The word-burst, or burstiness, pattern is the tendency for word utterances to appear together in bursts as conversational topics fluctuate. We give evidence that the burstiness phenomenon exhibits itself across varied languages. Using burstiness features to train a machine-learning algorithm, we are able to assess the likelihood that a hypothesized keyword location is correct and adjust its confidence score accordingly, yielding improvements in the efficacy of keyword search in low-resource languages. </p>"],"dc:identifier":["https://academicworks.cuny.edu/gc_etds/273"],"dc:subject":["Computer Sciences","Linguistics","Babel","burstiness","cache","keyword search","spoken term detection","word-burst"],"dc:title":["Echolocation: Using Word-Burst Analysis to Rescore Keyword Search Candidates in Low-Resource Languages"],"thesis:degree_discipline":["Linguistics"],"thesis:degree_level":["Master"],"thesis:degree_name":["Master of Arts"],"thesis:institution_name":["The Graduate School and University Center of The City University of New York"]},"updated_at":"2026-07-24T01:59:38Z"}