{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/101210"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/101210","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Adopting the two-branch network to video-text tasks","abstract":"This Thesis was approved for publication on 2018-04-23 at 16:33.","abstract_html":"This Thesis was approved for publication on 2018-04-23 at 16:33.","abstract_has_math":false,"creators":["Chang, Hsiao-Ching"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Lazebnik, Svetlana"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-09-04T20:36:51Z","date_published":"2018-09-04T20:36:51Z","updated_at":"2026-07-22T22:24:38Z","subjects":["Computer vision","Video captioning"],"languages":["en"],"rights":["Copyright 2018 Hsiao-Ching Chang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/101210","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lazebnik, Svetlana"]},{"key":"dc:creator","label":"Author","values":["Chang, Hsiao-Ching"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-09-04T20:36:51Z","2020-09-05T09:15:20Z","2018-04-23","2018-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer vision","Video captioning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Hsiao-Ching Chang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/101210"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This Thesis was approved for publication on 2018-04-23 at 16:33.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12431 on 2018-08-31 at 17:21:14","Made available in DSpace on 2018-09-04T20:36:51Z (GMT). No. of bitstreams: 2 CHANG-THESIS-2018.pdf: 15447494 bytes, checksum: 19c9cab37c0e29586fc0af1a3fd27846 (MD5) LICENSE.txt: 4214 bytes, checksum: 7282ca76dd914b160456ecc3dc7e4409 (MD5) Previous issue date: 2018-04-23","Modeling visual context and its corresponding text description with a joint embedding network has been an effective way to enable cross-modal retrieval. However, while abundant work has been done for image-text tasks, not much exists with regards to the video domain. We hope to adopt a nonlinear embedding model, the two-branch network, to the video-text tasks in order to show its robustness. Two kinds of tasks are explored, bidirectional video-sentence retrieval and video description generation. For the retrieval task, we use nearest neighbor search to get the corresponding video or text with respect to the query. For video captioning, we incorporate the two-branch network in a traditional LSTM model with an additional embedding loss term in order to demonstrate its ability of preserving a semantic structure between video and text.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2020-05-01","The student, Hsiao-Ching Chang, accepted the attached license on 2018-04-23 at 16:08.","The student, Hsiao-Ching Chang, submitted this Thesis for approval on 2018-04-23 at 16:14.","Embargo set by: Seth Robbins for item 107294 Lift date: 2020-09-04T20:37:00Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107294 Lift date: 2020-09-04T20:42:08Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 107294 on 2020-09-05T09:15:20Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Adopting the two-branch network to video-text tasks"]}]}],"canonical_facts":{"dc:contributor":["Lazebnik, Svetlana"],"dc:creator":["Chang, Hsiao-Ching"],"dc:date":["2018-09-04T20:36:51Z","2020-09-05T09:15:20Z","2018-04-23","2018-05"],"dc:description":["This Thesis was approved for publication on 2018-04-23 at 16:33.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12431 on 2018-08-31 at 17:21:14","Made available in DSpace on 2018-09-04T20:36:51Z (GMT). No. of bitstreams: 2 CHANG-THESIS-2018.pdf: 15447494 bytes, checksum: 19c9cab37c0e29586fc0af1a3fd27846 (MD5) LICENSE.txt: 4214 bytes, checksum: 7282ca76dd914b160456ecc3dc7e4409 (MD5) Previous issue date: 2018-04-23","Modeling visual context and its corresponding text description with a joint embedding network has been an effective way to enable cross-modal retrieval. However, while abundant work has been done for image-text tasks, not much exists with regards to the video domain. We hope to adopt a nonlinear embedding model, the two-branch network, to the video-text tasks in order to show its robustness. Two kinds of tasks are explored, bidirectional video-sentence retrieval and video description generation. For the retrieval task, we use nearest neighbor search to get the corresponding video or text with respect to the query. For video captioning, we incorporate the two-branch network in a traditional LSTM model with an additional embedding loss term in order to demonstrate its ability of preserving a semantic structure between video and text.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2020-05-01","The student, Hsiao-Ching Chang, accepted the attached license on 2018-04-23 at 16:08.","The student, Hsiao-Ching Chang, submitted this Thesis for approval on 2018-04-23 at 16:14.","Embargo set by: Seth Robbins for item 107294 Lift date: 2020-09-04T20:37:00Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107294 Lift date: 2020-09-04T20:42:08Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 107294 on 2020-09-05T09:15:20Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/101210"],"dc:language":["en"],"dc:rights":["Copyright 2018 Hsiao-Ching Chang"],"dc:subject":["Computer vision","Video captioning"],"dc:title":["Adopting the two-branch network to video-text tasks"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:38Z"}