{"id":{"repo_id":"toronto-retro","oai_identifier":"oai:utoronto.scholaris.ca:1807/89798"},"canonical_url":"https://search.dev.ndltd.org/etd/toronto-retro/oai:utoronto.scholaris.ca:1807/89798","repository":{"repo_id":"toronto-retro","name":"University of Toronto","base_url":"https://utoronto.scholaris.ca/server/oai/request"},"display":{"title":"Conditional Neural Language Models for Multimodal Learning and Natural Language Understanding","abstract":"In this thesis we introduce conditional neural language models based on log-bilinear and recurrent neural networks with applications to multimodal learning and natural language understanding. We first introduce a LSTM encoder for learning visual-semantic embeddings for ranking the relevance of text to images in a joint embedding space. Next we introduce three log-bilinear models for generating image descriptions that integrate both additive and multiplicative interactions. Beyond image conditioning, we describe a multiplicative conditional neural language model for learning distributed representations of attributes and meta data. Our model allows for contextual word relatedness comparisons through decompositions of a word embedding tensor. Finally we show how we can abstract the skip-gram model for learning word representations to a conditional recurrent neural language model for unsupervised learning of sentence representations. We introduce a family of models called contextual encoder-decoders and demonstrate how our models can be used to induce generic sentence representations as well as unaligned generation of short stories conditioned on images. This thesis closes by highlighting several open areas of future work.","abstract_html":"In this thesis we introduce conditional neural language models based on log-bilinear and recurrent neural networks with applications to multimodal learning and natural language understanding. We first introduce a LSTM encoder for learning visual-semantic embeddings for ranking the relevance of text to images in a joint embedding space. Next we introduce three log-bilinear models for generating image descriptions that integrate both additive and multiplicative interactions. Beyond image conditioning, we describe a multiplicative conditional neural language model for learning distributed representations of attributes and meta data. Our model allows for contextual word relatedness comparisons through decompositions of a word embedding tensor. Finally we show how we can abstract the skip-gram model for learning word representations to a conditional recurrent neural language model for unsupervised learning of sentence representations. We introduce a family of models called contextual encoder-decoders and demonstrate how our models can be used to induce generic sentence representations as well as unaligned generation of short stories conditioned on images. This thesis closes by highlighting several open areas of future work.","abstract_has_math":false,"creators":["Kiros, Jamie Ryan"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Computer Science","school":null,"contributors":[],"advisors":["Zemel, Richard","Salakhutdinov, Ruslan"],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-06","date_published":"2018-06","updated_at":"2026-07-27T21:28:02Z","subjects":["Computer Vision","Deep Learning","Language Models","Machine Learning","Natural Language Processing","Neural Networks"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1807/89798","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Zemel, Richard","Salakhutdinov, Ruslan"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science"]},{"key":"dc:creator","label":"Author","values":["Kiros, Jamie Ryan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-06"]},{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2018-07-18T19:03:36Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2018-07-18T19:03:36Z"]},{"key":"dc:date.issued","label":"Date","values":["2018-06"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Vision","Deep Learning","Language Models","Machine Learning","Natural Language Processing","Neural Networks"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1807/89798"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In this thesis we introduce conditional neural language models based on log-bilinear and recurrent neural networks with applications to multimodal learning and natural language understanding. We first introduce a LSTM encoder for learning visual-semantic embeddings for ranking the relevance of text to images in a joint embedding space. Next we introduce three log-bilinear models for generating image descriptions that integrate both additive and multiplicative interactions. Beyond image conditioning, we describe a multiplicative conditional neural language model for learning distributed representations of attributes and meta data. Our model allows for contextual word relatedness comparisons through decompositions of a word embedding tensor. Finally we show how we can abstract the skip-gram model for learning word representations to a conditional recurrent neural language model for unsupervised learning of sentence representations. We introduce a family of models called contextual encoder-decoders and demonstrate how our models can be used to induce generic sentence representations as well as unaligned generation of short stories conditioned on images. This thesis closes by highlighting several open areas of future work."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Ph.D."]},{"key":"dc:title","label":"Title","values":["Conditional Neural Language Models for Multimodal Learning and Natural Language Understanding"]}]}],"canonical_facts":{"dc:contributor.advisor":["Zemel, Richard","Salakhutdinov, Ruslan"],"dc:contributor.department":["Computer Science"],"dc:creator":["Kiros, Jamie Ryan"],"dc:date":["2018-06"],"dc:date.accessioned":["2018-07-18T19:03:36Z"],"dc:date.available":["2018-07-18T19:03:36Z"],"dc:date.issued":["2018-06"],"dc:description.abstract":["In this thesis we introduce conditional neural language models based on log-bilinear and recurrent neural networks with applications to multimodal learning and natural language understanding. We first introduce a LSTM encoder for learning visual-semantic embeddings for ranking the relevance of text to images in a joint embedding space. Next we introduce three log-bilinear models for generating image descriptions that integrate both additive and multiplicative interactions. Beyond image conditioning, we describe a multiplicative conditional neural language model for learning distributed representations of attributes and meta data. Our model allows for contextual word relatedness comparisons through decompositions of a word embedding tensor. Finally we show how we can abstract the skip-gram model for learning word representations to a conditional recurrent neural language model for unsupervised learning of sentence representations. We introduce a family of models called contextual encoder-decoders and demonstrate how our models can be used to induce generic sentence representations as well as unaligned generation of short stories conditioned on images. This thesis closes by highlighting several open areas of future work."],"dc:description.degree":["Ph.D."],"dc:identifier.uri":["http://hdl.handle.net/1807/89798"],"dc:subject":["Computer Vision","Deep Learning","Language Models","Machine Learning","Natural Language Processing","Neural Networks"],"dc:title":["Conditional Neural Language Models for Multimodal Learning and Natural Language Understanding"],"dc:type":["Thesis"]},"updated_at":"2026-07-27T21:28:02Z"}