Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 12 of 12 for “"VQA"”.
-
Visual Question Answering in the Medical Domain
… to have a reliable Visual Question Answering (VQA) system which can provide a "second opinion" on medical cases. However, most of the VQA systems that work today cater to real-world problems and are not specifically tailored for handling medical images. Moreover, the VQA system for medical …
-
Role of Premises in Visual Question Answering
… premises can help Visual Question Answering (VQA) models respond more intelligently to irrelevant or previously unseen questions. When presented with a question that is irrelevant to an image, state-of-the-art VQA models will still answer based purely on learned language biases, resulting in …
-
Assessing the Benefits of Virginia Tech Agricultural Programs: Studies in Feeder Cattle Certification and Small Grains Breeding
… No significant effect on price from VQA certification is found. However, enterprise budgets indicate that VQA cattle allow higher farm profits due to their lower sale weight, which allows for faster turnover and lower prices. The second paper studies the benefits to producers from …
-
Leveraging Multimodal Perspectives to Learn Common Sense for Vision and Language Tasks
… active learning for Visual Question Answering (VQA) where a model iteratively grows its knowledge through querying informative questions about images for answers. Drawing analogies from human learning, we explore cramming (entropy), curiosity-driven (expected model change), and goal-driven …
-
Visual datasets for artificial intelligence agents
… procedurally generated new data to test VQA agents and other visual Al models on. The first tool is Spatial IQ Generative Dataset (SIQGD). This tool generates images based on the Raven's Progressive Matrices spatial IQ examination metric. The second tool is a collection of 3D models along …
-
Towards Interpretable Vision Systems
… semantic matching and Visual Question Answering (VQA). In VQA, we take binary questions on abstract scenes as the first stage, then we extend to all question types on real images. In both cases, we take attention as an important intermediate output. By explicitly forcing the systems to attend …
-
Augmenting Multi-modal Question Answering Systems with Retrieval Methods
… visually-grounded questions, we introduce RA-VQA (Retrieval Augmented Visual Question Answering), a framework tailored for knowledge-based visual question answering (KB-VQA). We demonstrate the efficacy of joint training for retriever and generator models in maximising performance. Secondly, …
-
Data Augmentation with Seq2Seq Models
… search. We evaluate the results on the standard VQA validation set. Our approach results in a significantly expanded training dataset and vocabulary size, but has slightly worse performance when tested on the validation split. Although not as fruitful as we had hoped, our work highlights …
-
Exploiting multimodality and structure in world representations
… By extending popular visual question answering (VQA) paradigms, I also designed several models that were evaluated on the novel dataset. This produced initial performance estimates for environment understanding, through the lens of a more challenging VQA downstream task. The second work presents …
-
Fact-based visual question answering using knowledge graph embeddings
… language. Fact-based Visual Question Answering (FVQA), a challenging variant of VQA, requires a QA-system to mimic this human ability. It must include facts from a diverse knowledge graph (KG) in its reasoning process to produce an answer. Large KGs, especially common-sense KGs, are known to be …
-
Unifying cross-modal concepts in vision and language
Enabling computers to demonstrate a proficient understanding of the physical world is an exceedingly challenging task that necessitates the ability to perceive, through vision or other senses, and communicate through natural language. Key to this endeavor is the representation of concepts present …
-
Learning visual tasks with selective attention
Knowing where to look in an image can significantly improve performance in computer vision tasks by eliminating irrelevant information from the rest of the input image, and by breaking down complex scenes into simpler and more familiar sub-components. We show that a framework for identifying …