{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/81157"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/81157","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Generative Models for Retrieval of Video, Audio and Text Data","abstract":"\"We propose a general approach to audio segment retrieval, which allows a user to query audio data by an example audio segment of a short duration and to find similar segments. The basic idea of our approach is to first train a hidden Markov model (HMM) using the given example, it is called the theme HMM. The total audio data available is used to train a background HMM. We combine these individual HMMs to form a synthesized \"\"background-theme-background\"\" HMM. This synthesized HMM can then be applied to any audio stream as a parser to detect the most likely theme segment. A major advantage of this approach is that it does not assume any predefined segment boundaries as in previous work and thus can be expected to retrieve theme segments with more accurate boundaries. We overcome the problem of only being able to use a short duration query to train a theme HMM by using the maximum a posteriori estimator with the background model as a prior model. Evaluation of the proposed retrieval scheme using short duration example audio clips of narration as queries gives quite promising results.\"","abstract_html":"&quot;We propose a general approach to audio segment retrieval, which allows a user to query audio data by an example audio segment of a short duration and to find similar segments. The basic idea of our approach is to first train a hidden Markov model (HMM) using the given example, it is called the theme HMM. The total audio data available is used to train a background HMM. We combine these individual HMMs to form a synthesized &quot;&quot;background-theme-background&quot;&quot; HMM. This synthesized HMM can then be applied to any audio stream as a parser to detect the most likely theme segment. A major advantage of this approach is that it does not assume any predefined segment boundaries as in previous work and thus can be expected to retrieve theme segments with more accurate boundaries. We overcome the problem of only being able to use a short duration query to train a theme HMM by using the maximum a posteriori estimator with the background model as a prior model. Evaluation of the proposed retrieval scheme using short duration example audio clips of narration as queries gives quite promising results.&quot;","abstract_has_math":false,"creators":["Velivelli, Atulya"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical and Computer Engineering","degree_department":null,"school":null,"contributors":["Huang, Thomas S."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-09-25T20:09:48Z","date_published":"2015-09-25T20:09:48Z","updated_at":"2026-07-22T22:26:15Z","subjects":["Engineering, Electronics and Electrical"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["(MiAaPQ)AAI3425399"],"render_values":[{"text":"(MiAaPQ)AAI3425399","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/81157","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Huang, Thomas S."]},{"key":"dc:creator","label":"Author","values":["Velivelli, Atulya"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-09-25T20:09:48Z","10000-01-01","2010"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical and Computer Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Engineering, Electronics and Electrical"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/81157","(MiAaPQ)AAI3425399"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"We propose a general approach to audio segment retrieval, which allows a user to query audio data by an example audio segment of a short duration and to find similar segments. The basic idea of our approach is to first train a hidden Markov model (HMM) using the given example, it is called the theme HMM. The total audio data available is used to train a background HMM. We combine these individual HMMs to form a synthesized \"\"background-theme-background\"\" HMM. This synthesized HMM can then be applied to any audio stream as a parser to detect the most likely theme segment. A major advantage of this approach is that it does not assume any predefined segment boundaries as in previous work and thus can be expected to retrieve theme segments with more accurate boundaries. We overcome the problem of only being able to use a short duration query to train a theme HMM by using the maximum a posteriori estimator with the background model as a prior model. Evaluation of the proposed retrieval scheme using short duration example audio clips of narration as queries gives quite promising results.\"","Made available in DSpace on 2015-09-25T20:09:48Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3425399.pdf: 5225473 bytes, checksum: 489737a84ef857d254a94c49d22d686c (MD5) Previous issue date: 2010","Embargo set by: Seth Robbins for item 82438 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","120 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2010."]},{"key":"dc:title","label":"Title","values":["Generative Models for Retrieval of Video, Audio and Text Data"]}]}],"canonical_facts":{"dc:contributor":["Huang, Thomas S."],"dc:creator":["Velivelli, Atulya"],"dc:date":["2015-09-25T20:09:48Z","10000-01-01","2010"],"dc:description":["\"We propose a general approach to audio segment retrieval, which allows a user to query audio data by an example audio segment of a short duration and to find similar segments. The basic idea of our approach is to first train a hidden Markov model (HMM) using the given example, it is called the theme HMM. The total audio data available is used to train a background HMM. We combine these individual HMMs to form a synthesized \"\"background-theme-background\"\" HMM. This synthesized HMM can then be applied to any audio stream as a parser to detect the most likely theme segment. A major advantage of this approach is that it does not assume any predefined segment boundaries as in previous work and thus can be expected to retrieve theme segments with more accurate boundaries. We overcome the problem of only being able to use a short duration query to train a theme HMM by using the maximum a posteriori estimator with the background model as a prior model. Evaluation of the proposed retrieval scheme using short duration example audio clips of narration as queries gives quite promising results.\"","Made available in DSpace on 2015-09-25T20:09:48Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3425399.pdf: 5225473 bytes, checksum: 489737a84ef857d254a94c49d22d686c (MD5) Previous issue date: 2010","Embargo set by: Seth Robbins for item 82438 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","120 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2010."],"dc:identifier":["http://hdl.handle.net/2142/81157","(MiAaPQ)AAI3425399"],"dc:language":["eng"],"dc:subject":["Engineering, Electronics and Electrical"],"dc:title":["Generative Models for Retrieval of Video, Audio and Text Data"],"dc:type":["text"],"thesis:degree_discipline":["Electrical and Computer Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:15Z"}