Abstract
dc:description.abstractThis work introduces methods for learning distributed, vector representations of cooking recipes. The individual components of a recipe -- the images, instructions, and ingredients -- are first treated individually. These representations are learned from a large, multi-modal dataset collected -- and publicly released -- as part of this work. Their representations are then embedded in a joint vector space using a novel neural network model. Experiments on cross-modal retrieval and vector space arithmetic demonstrate the utility and generalizability of both the per-component and joint embeddings.
Degree
thesis:*- Department dc:contributor.department
- Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science.
- Grantor dc:publisher
- Massachusetts Institute of Technology
- Year dc:date.issued
- 2017
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Hynes, Nick (Nick I.)
- Advisor dc:contributor.advisor
-
- Antonio Torralba.
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- MIT theses are protected by copyright. They may be viewed, downloaded, or printed from this source but further reproduction or distribution in any format is prohibited without written permission.
- Licence dc:rights.uri
- Language dc:language.iso
- eng
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/1721.1/113147
- OAI identifier oai:identifier
- oai:dspace.mit.edu:1721.1/113147