{"id":{"repo_id":"gatech","oai_identifier":"oai:repository.gatech.edu:1853/72039"},"canonical_url":"https://search.dev.ndltd.org/etd/gatech/oai:repository.gatech.edu:1853/72039","repository":{"repo_id":"gatech","name":"Georgia Tech","base_url":"https://repository.gatech.edu/server/oai/request"},"display":{"title":"Zero-shot object-goal navigation using multimodal goal embeddings","abstract":"My thesis presents a scalable approach for learning open-world object-goal navigation (ObjectNav) - the task of asking a virtual robot (agent) to find any instance of an object in an unexplored environment (e.g., \"find a sink\"). The approach is entirely zero-shot - i.e., it does not require ObjectNav rewards or demonstrations of any kind. Instead, we train on the image-goal navigation (ImageNav) task, in which agents find the location where a picture (i.e., goal image) was captured. Specifically, we encode goal images into a multimodal, semantic embedding space to enable training semantic-goal navigation (SemanticNav) agents at scale in unannotated 3D environments (e.g., HM3D). After training, SemanticNav agents can be instructed to find objects described in free-form natural language (e.g., \"sink,\" \"bathroom sink,\" etc.) by projecting language goals into the same multimodal, semantic embedding space. As a result, our approach enables openworld ObjectNav. We extensively evaluate our agents on three ObjectNav datasets (Gibson, HM3D, and MP3D) and observe absolute improvements in success of 4.2% - 20.0% over existing zero-shot methods. For reference, these gains are similar or better than the 5% improvement in success between the Habitat 2020 and 2021 ObjectNav challenge winners. In an open-world setting, we discover that our agents can generalize to compound instructions with a room explicitly mentioned (e.g., \"Find a kitchen sink\") and when the target room can be inferred (e.g., \"Find a sink and a stove\").","abstract_html":"My thesis presents a scalable approach for learning open-world object-goal navigation (ObjectNav) - the task of asking a virtual robot (agent) to find any instance of an object in an unexplored environment (e.g., &quot;find a sink&quot;). The approach is entirely zero-shot - i.e., it does not require ObjectNav rewards or demonstrations of any kind. Instead, we train on the image-goal navigation (ImageNav) task, in which agents find the location where a picture (i.e., goal image) was captured. Specifically, we encode goal images into a multimodal, semantic embedding space to enable training semantic-goal navigation (SemanticNav) agents at scale in unannotated 3D environments (e.g., HM3D). After training, SemanticNav agents can be instructed to find objects described in free-form natural language (e.g., &quot;sink,&quot; &quot;bathroom sink,&quot; etc.) by projecting language goals into the same multimodal, semantic embedding space. As a result, our approach enables openworld ObjectNav. We extensively evaluate our agents on three ObjectNav datasets (Gibson, HM3D, and MP3D) and observe absolute improvements in success of 4.2% - 20.0% over existing zero-shot methods. For reference, these gains are similar or better than the 5% improvement in success between the Habitat 2020 and 2021 ObjectNav challenge winners. In an open-world setting, we discover that our agents can generalize to compound instructions with a room explicitly mentioned (e.g., &quot;Find a kitchen sink&quot;) and when the target room can be inferred (e.g., &quot;Find a sink and a stove&quot;).","abstract_has_math":false,"creators":["Aggarwal, Gunjan"],"institution":"Georgia Institute of Technology","degree_name":null,"degree_level":"Masters","degree_discipline":null,"degree_department":"Computer Science","school":null,"contributors":[],"advisors":["Batra, Dhruv"],"committee_chairs":[],"committee_members":["Hoffmann, Judy","Parikh, Devi"],"year":2023,"date_issued":"2023-05-01","date_published":"2023-05-01","updated_at":"2026-07-27T19:49:46Z","subjects":["Embodied AI","multi-modal","navigation","zero-shot"],"languages":["en_US"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1853/72039","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Batra, Dhruv"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Hoffmann, Judy","Parikh, Devi"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science"]},{"key":"dc:creator","label":"Author","values":["Aggarwal, Gunjan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2023-05-18T17:54:18Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2023-05-18T17:54:18Z"]},{"key":"dc:date.issued","label":"Date","values":["2023-05-01"]},{"key":"dc:publisher","label":"Institution","values":["Georgia Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Text"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Embodied AI","multi-modal","navigation","zero-shot"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1853/72039"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["My thesis presents a scalable approach for learning open-world object-goal navigation (ObjectNav) - the task of asking a virtual robot (agent) to find any instance of an object in an unexplored environment (e.g., \"find a sink\"). The approach is entirely zero-shot - i.e., it does not require ObjectNav rewards or demonstrations of any kind. Instead, we train on the image-goal navigation (ImageNav) task, in which agents find the location where a picture (i.e., goal image) was captured. Specifically, we encode goal images into a multimodal, semantic embedding space to enable training semantic-goal navigation (SemanticNav) agents at scale in unannotated 3D environments (e.g., HM3D). After training, SemanticNav agents can be instructed to find objects described in free-form natural language (e.g., \"sink,\" \"bathroom sink,\" etc.) by projecting language goals into the same multimodal, semantic embedding space. As a result, our approach enables openworld ObjectNav. We extensively evaluate our agents on three ObjectNav datasets (Gibson, HM3D, and MP3D) and observe absolute improvements in success of 4.2% - 20.0% over existing zero-shot methods. For reference, these gains are similar or better than the 5% improvement in success between the Habitat 2020 and 2021 ObjectNav challenge winners. In an open-world setting, we discover that our agents can generalize to compound instructions with a room explicitly mentioned (e.g., \"Find a kitchen sink\") and when the target room can be inferred (e.g., \"Find a sink and a stove\")."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["M.S."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Zero-shot object-goal navigation using multimodal goal embeddings"]}]}],"canonical_facts":{"dc:contributor.advisor":["Batra, Dhruv"],"dc:contributor.committeemember":["Hoffmann, Judy","Parikh, Devi"],"dc:contributor.department":["Computer Science"],"dc:creator":["Aggarwal, Gunjan"],"dc:date.accessioned":["2023-05-18T17:54:18Z"],"dc:date.available":["2023-05-18T17:54:18Z"],"dc:date.issued":["2023-05-01"],"dc:description.abstract":["My thesis presents a scalable approach for learning open-world object-goal navigation (ObjectNav) - the task of asking a virtual robot (agent) to find any instance of an object in an unexplored environment (e.g., \"find a sink\"). The approach is entirely zero-shot - i.e., it does not require ObjectNav rewards or demonstrations of any kind. Instead, we train on the image-goal navigation (ImageNav) task, in which agents find the location where a picture (i.e., goal image) was captured. Specifically, we encode goal images into a multimodal, semantic embedding space to enable training semantic-goal navigation (SemanticNav) agents at scale in unannotated 3D environments (e.g., HM3D). After training, SemanticNav agents can be instructed to find objects described in free-form natural language (e.g., \"sink,\" \"bathroom sink,\" etc.) by projecting language goals into the same multimodal, semantic embedding space. As a result, our approach enables openworld ObjectNav. We extensively evaluate our agents on three ObjectNav datasets (Gibson, HM3D, and MP3D) and observe absolute improvements in success of 4.2% - 20.0% over existing zero-shot methods. For reference, these gains are similar or better than the 5% improvement in success between the Habitat 2020 and 2021 ObjectNav challenge winners. In an open-world setting, we discover that our agents can generalize to compound instructions with a room explicitly mentioned (e.g., \"Find a kitchen sink\") and when the target room can be inferred (e.g., \"Find a sink and a stove\")."],"dc:description.degree":["M.S."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/1853/72039"],"dc:language.iso":["en_US"],"dc:publisher":["Georgia Institute of Technology"],"dc:subject":["Embodied AI","multi-modal","navigation","zero-shot"],"dc:title":["Zero-shot object-goal navigation using multimodal goal embeddings"],"dc:type":["Text"],"thesis:degree_level":["Masters"]},"updated_at":"2026-07-27T19:49:46Z"}