{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/101574"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/101574","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Visual questioning agents","abstract":"Curious questioning or the ability to inquire about surrounding environment or additional context, is an important step towards building agents which go beyond learning from a static knowledge base. The ability to request feedback is the first step in building intelligent agents which can incorporate this feedback to enhance learning. Visual questioning tasks help model this human skill of “curiosity.” In this thesis, we focus on two relevant vision based questioning tasks – visual question generation and visual dialog. We propose novel approaches and evaluation metrics for these tasks. For visual question generation, we combined language models with variational autoencoders to enhance diversity in text generations. We also suggest diversity metrics to quantify these improvements. For visual dialog, we introduce a reformulated dataset to enable training of questioning agents in a dialog setup. We also introduce simpler and more effective baselines for the task. Our combined results in visual question generation and visual dialog contribute to establishing visual questioning as an important next step for computer vision, and more generally, for artificial intelligence.","abstract_html":"Curious questioning or the ability to inquire about surrounding environment or additional context, is an important step towards building agents which go beyond learning from a static knowledge base. The ability to request feedback is the first step in building intelligent agents which can incorporate this feedback to enhance learning. Visual questioning tasks help model this human skill of “curiosity.” In this thesis, we focus on two relevant vision based questioning tasks – visual question generation and visual dialog. We propose novel approaches and evaluation metrics for these tasks. For visual question generation, we combined language models with variational autoencoders to enhance diversity in text generations. We also suggest diversity metrics to quantify these improvements. For visual dialog, we introduce a reformulated dataset to enable training of questioning agents in a dialog setup. We also introduce simpler and more effective baselines for the task. Our combined results in visual question generation and visual dialog contribute to establishing visual questioning as an important next step for computer vision, and more generally, for artificial intelligence.","abstract_has_math":false,"creators":["Jain, Unnat"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Lazebnik, Svetlana","Schwing, Alexander Gerhard"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-09-27T16:17:50Z","date_published":"2018-09-27T16:17:50Z","updated_at":"2026-07-22T22:24:40Z","subjects":["Visual Question Generation, Visual Dialog, Variational Autoencoders, Language and Vision, Computer Vision"],"languages":["en"],"rights":["Copyright 2018 Unnat Jain"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/101574","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lazebnik, Svetlana","Schwing, Alexander Gerhard"]},{"key":"dc:creator","label":"Author","values":["Jain, Unnat"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-09-27T16:17:50Z","2018-07-16","2018-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Visual Question Generation, Visual Dialog, Variational Autoencoders, Language and Vision, Computer Vision"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Unnat Jain"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/101574"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Curious questioning or the ability to inquire about surrounding environment or additional context, is an important step towards building agents which go beyond learning from a static knowledge base. The ability to request feedback is the first step in building intelligent agents which can incorporate this feedback to enhance learning. Visual questioning tasks help model this human skill of “curiosity.” In this thesis, we focus on two relevant vision based questioning tasks – visual question generation and visual dialog. We propose novel approaches and evaluation metrics for these tasks. For visual question generation, we combined language models with variational autoencoders to enhance diversity in text generations. We also suggest diversity metrics to quantify these improvements. For visual dialog, we introduce a reformulated dataset to enable training of questioning agents in a dialog setup. We also introduce simpler and more effective baselines for the task. Our combined results in visual question generation and visual dialog contribute to establishing visual questioning as an important next step for computer vision, and more generally, for artificial intelligence.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2018-09-27 without embargo terms","The student, Unnat Jain, accepted the attached license on 2018-07-12 at 16:30.","The student, Unnat Jain, submitted this Thesis for approval on 2018-07-12 at 16:30.","This Thesis was approved for publication on 2018-07-16 at 09:02.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12856 on 2018-09-27 at 10:48:08","Made available in DSpace on 2018-09-27T16:17:50Z (GMT). No. of bitstreams: 2 JAIN-THESIS-2018.pdf: 4941312 bytes, checksum: 738c261f3846c2b230af34f426838f5c (MD5) LICENSE.txt: 4207 bytes, checksum: b6dbd54489d2a7d13dfe4d72d1b2865e (MD5) Previous issue date: 2018-07-16"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Visual questioning agents"]}]}],"canonical_facts":{"dc:contributor":["Lazebnik, Svetlana","Schwing, Alexander Gerhard"],"dc:creator":["Jain, Unnat"],"dc:date":["2018-09-27T16:17:50Z","2018-07-16","2018-08"],"dc:description":["Curious questioning or the ability to inquire about surrounding environment or additional context, is an important step towards building agents which go beyond learning from a static knowledge base. The ability to request feedback is the first step in building intelligent agents which can incorporate this feedback to enhance learning. Visual questioning tasks help model this human skill of “curiosity.” In this thesis, we focus on two relevant vision based questioning tasks – visual question generation and visual dialog. We propose novel approaches and evaluation metrics for these tasks. For visual question generation, we combined language models with variational autoencoders to enhance diversity in text generations. We also suggest diversity metrics to quantify these improvements. For visual dialog, we introduce a reformulated dataset to enable training of questioning agents in a dialog setup. We also introduce simpler and more effective baselines for the task. Our combined results in visual question generation and visual dialog contribute to establishing visual questioning as an important next step for computer vision, and more generally, for artificial intelligence.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2018-09-27 without embargo terms","The student, Unnat Jain, accepted the attached license on 2018-07-12 at 16:30.","The student, Unnat Jain, submitted this Thesis for approval on 2018-07-12 at 16:30.","This Thesis was approved for publication on 2018-07-16 at 09:02.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12856 on 2018-09-27 at 10:48:08","Made available in DSpace on 2018-09-27T16:17:50Z (GMT). No. of bitstreams: 2 JAIN-THESIS-2018.pdf: 4941312 bytes, checksum: 738c261f3846c2b230af34f426838f5c (MD5) LICENSE.txt: 4207 bytes, checksum: b6dbd54489d2a7d13dfe4d72d1b2865e (MD5) Previous issue date: 2018-07-16"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/101574"],"dc:language":["en"],"dc:rights":["Copyright 2018 Unnat Jain"],"dc:subject":["Visual Question Generation, Visual Dialog, Variational Autoencoders, Language and Vision, Computer Vision"],"dc:title":["Visual questioning agents"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:40Z"}