{"id":{"repo_id":"houston","oai_identifier":"oai:uh-ir.tdl.org:10657/1666"},"canonical_url":"https://search.dev.ndltd.org/etd/houston/oai:uh-ir.tdl.org:10657/1666","repository":{"repo_id":"houston","name":"University of Houston","base_url":"https://uh-ir.tdl.org/server/oai/request"},"display":{"title":"Evaluation of Speech And Text-Based Indexing For Classroom Lecture Videos","abstract":"Lecture videos are useful and great learning resources. At the University of Houston, videos are widely used throughout departments within the College of Natural Sciences and Mathematics such as Computer Science, Biology and Biochemistry, Earth and Atmospheric Sciences, etc. Since most videos are very long, it is difficult to directly access the required topic within a video. The ICS (indexed, captioned, and searchable) videos project provides students direct access to a topic within video lectures by providing index points representing the topic. These index points are generated using text from the extracted images using OCR (optical character recognition) technology. Index points are assigned with the assistance of an indexing algorithm that determines topic change based on text similarity. We present a topic-based lecture video segmentation using speech text/captions. The purpose of this thesis is to utilize the spoken text of a lecture video to assign index points using an underlying text-based indexing algorithm. To achieve this goal, a set of twenty-five lecture videos was taken from various departments at the University of Houston and Coursera website. The captions were produced with the assistance of the YouTube Speech Recognition System. The performances and limitations of OCR text, uncorrected/original speech text, and corrected speech text-based indexing was analyzed. The results indicate that slide text-based indexing yields 4% better results than spoken text-based indexing. The corrected speech text/caption provides better indexing results (11%) where OCR text fails to perform and the results closely matched the ground truth. The error analysis done on speech texts and slide texts prove that poor OCR text and caption quality are some of the main issues that hamper indexing accuracy.","abstract_html":"Lecture videos are useful and great learning resources. At the University of Houston, videos are widely used throughout departments within the College of Natural Sciences and Mathematics such as Computer Science, Biology and Biochemistry, Earth and Atmospheric Sciences, etc. Since most videos are very long, it is difficult to directly access the required topic within a video. The ICS (indexed, captioned, and searchable) videos project provides students direct access to a topic within video lectures by providing index points representing the topic. These index points are generated using text from the extracted images using OCR (optical character recognition) technology. Index points are assigned with the assistance of an indexing algorithm that determines topic change based on text similarity. We present a topic-based lecture video segmentation using speech text/captions. The purpose of this thesis is to utilize the spoken text of a lecture video to assign index points using an underlying text-based indexing algorithm. To achieve this goal, a set of twenty-five lecture videos was taken from various departments at the University of Houston and Coursera website. The captions were produced with the assistance of the YouTube Speech Recognition System. The performances and limitations of OCR text, uncorrected/original speech text, and corrected speech text-based indexing was analyzed. The results indicate that slide text-based indexing yields 4% better results than spoken text-based indexing. The corrected speech text/caption provides better indexing results (11%) where OCR text fails to perform and the results closely matched the ground truth. The error analysis done on speech texts and slide texts prove that poor OCR text and caption quality are some of the main issues that hamper indexing accuracy.","abstract_has_math":false,"creators":["Joshi, Mahima 1990-"],"institution":"University of Houston","degree_name":"Master of Science","degree_level":"Masters","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Subhlok, Jaspal"],"committee_chairs":[],"committee_members":["Johnson, Olin","Barr, Christopher D."],"year":2014,"date_issued":"2014-12","date_published":"2014-12","updated_at":"2026-07-24T02:32:27Z","subjects":["ICS videos","Transition points","Indexing","Captioning"],"languages":["eng"],"rights":["The author of this work is the copyright owner. UH Libraries and the Texas Digital Library have their permission to store and provide access to this work. Further transmission, reproduction, or presentation of this work is prohibited except with permission of the author(s)."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/10657/1666","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Subhlok, Jaspal"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Johnson, Olin","Barr, Christopher D."]},{"key":"dc:creator","label":"Author","values":["Joshi, Mahima 1990-"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2017-04-09T23:29:58Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2017-04-09T23:29:58Z"]},{"key":"dc:date.issued","label":"Date","values":["2014-12"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Houston"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["ICS videos","Transition points","Indexing","Captioning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["The author of this work is the copyright owner. UH Libraries and the Texas Digital Library have their permission to store and provide access to this work. Further transmission, reproduction, or presentation of this work is prohibited except with permission of the author(s)."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/10657/1666"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Lecture videos are useful and great learning resources. At the University of Houston, videos are widely used throughout departments within the College of Natural Sciences and Mathematics such as Computer Science, Biology and Biochemistry, Earth and Atmospheric Sciences, etc. Since most videos are very long, it is difficult to directly access the required topic within a video. The ICS (indexed, captioned, and searchable) videos project provides students direct access to a topic within video lectures by providing index points representing the topic. These index points are generated using text from the extracted images using OCR (optical character recognition) technology. Index points are assigned with the assistance of an indexing algorithm that determines topic change based on text similarity. We present a topic-based lecture video segmentation using speech text/captions. The purpose of this thesis is to utilize the spoken text of a lecture video to assign index points using an underlying text-based indexing algorithm. To achieve this goal, a set of twenty-five lecture videos was taken from various departments at the University of Houston and Coursera website. The captions were produced with the assistance of the YouTube Speech Recognition System. The performances and limitations of OCR text, uncorrected/original speech text, and corrected speech text-based indexing was analyzed. The results indicate that slide text-based indexing yields 4% better results than spoken text-based indexing. The corrected speech text/caption provides better indexing results (11%) where OCR text fails to perform and the results closely matched the ground truth. The error analysis done on speech texts and slide texts prove that poor OCR text and caption quality are some of the main issues that hamper indexing accuracy."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Evaluation of Speech And Text-Based Indexing For Classroom Lecture Videos"]}]}],"canonical_facts":{"dc:contributor.advisor":["Subhlok, Jaspal"],"dc:contributor.committeemember":["Johnson, Olin","Barr, Christopher D."],"dc:creator":["Joshi, Mahima 1990-"],"dc:date.accessioned":["2017-04-09T23:29:58Z"],"dc:date.available":["2017-04-09T23:29:58Z"],"dc:date.issued":["2014-12"],"dc:description.abstract":["Lecture videos are useful and great learning resources. At the University of Houston, videos are widely used throughout departments within the College of Natural Sciences and Mathematics such as Computer Science, Biology and Biochemistry, Earth and Atmospheric Sciences, etc. Since most videos are very long, it is difficult to directly access the required topic within a video. The ICS (indexed, captioned, and searchable) videos project provides students direct access to a topic within video lectures by providing index points representing the topic. These index points are generated using text from the extracted images using OCR (optical character recognition) technology. Index points are assigned with the assistance of an indexing algorithm that determines topic change based on text similarity. We present a topic-based lecture video segmentation using speech text/captions. The purpose of this thesis is to utilize the spoken text of a lecture video to assign index points using an underlying text-based indexing algorithm. To achieve this goal, a set of twenty-five lecture videos was taken from various departments at the University of Houston and Coursera website. The captions were produced with the assistance of the YouTube Speech Recognition System. The performances and limitations of OCR text, uncorrected/original speech text, and corrected speech text-based indexing was analyzed. The results indicate that slide text-based indexing yields 4% better results than spoken text-based indexing. The corrected speech text/caption provides better indexing results (11%) where OCR text fails to perform and the results closely matched the ground truth. The error analysis done on speech texts and slide texts prove that poor OCR text and caption quality are some of the main issues that hamper indexing accuracy."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["http://hdl.handle.net/10657/1666"],"dc:language.iso":["eng"],"dc:rights":["The author of this work is the copyright owner. UH Libraries and the Texas Digital Library have their permission to store and provide access to this work. Further transmission, reproduction, or presentation of this work is prohibited except with permission of the author(s)."],"dc:subject":["ICS videos","Transition points","Indexing","Captioning"],"dc:title":["Evaluation of Speech And Text-Based Indexing For Classroom Lecture Videos"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["University of Houston"]},"updated_at":"2026-07-24T02:32:27Z"}