University of Illinois at Urbana-Champaign
Adopting the two-branch network to video-text tasks
Abstract
dc:descriptionModeling visual context and its corresponding text description with a joint embedding network has been an effective way to enable cross-modal retrieval. However, while abundant work has been done for image-text tasks, not much exists with regards to the video domain. We hope to adopt a nonlinear embedding model, the two-branch network, to the video-text tasks in order to show its robustness. Two kinds of tasks are explored, bidirectional video-sentence retrieval and video description generation. For the retrieval task, we use nearest neighbor search to get the corresponding video or text with respect to the query. For video captioning, we incorporate the two-branch network in a traditional LSTM model with an additional embedding loss term in order to demonstrate its ability of preserving a semantic structure between video and text.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2018
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Chang, Hsiao-Ching
- Contributors dc:contributor
-
- Lazebnik, Svetlana
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2018 Hsiao-Ching Chang
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/101210