Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 35 for “"Vision and Language"”.
-
Representations from vision and language
Replicating a human-level understanding of the physical world in computers is a monumental task. Achieving this requires building representations of concepts that manifest themselves visually, linguistically or through other senses. Furthermore concepts do not exist in isolation but are related to …
-
Vision and Language: Information Integration and Transformation
Images and text are large-scale data sources for deep learning models to mimic how humans perceive and understand multimodal information. Combining images and text can construct better feature representations by involving complementary information from different sources. For example, text …
-
Generative and Multimodal Learning for Vision and Language
L'abstract è presente nell'allegato / the abstract is in the attachment
-
Cognitive Map Generation for Vision and Language Navigation
<p>Visual-Language Navigation (VLN) presents significant challenges for autonomous agents, such as robots and virtual assistants, particularly in complex, dynamic environments where the seamless integration of visual perception and natural language understanding is critical. Traditional VLN systems …
-
Unifying cross-modal concepts in vision and language
… computers to demonstrate a proficient understanding of the physical world is an exceedingly challenging task that necessitates the ability to perceive, through vision or other senses, and communicate through natural language. Key to this endeavor is the representation of concepts present in …
-
Connecting vision and language via image retrieval and captioning
Many real-world problems involve jointly understanding vision and language, e.g. image/video captioning, multi-modal image retrieval, visual question answering. In this thesis, we consider several problems in cross-modal learning from vision and language. First, the problem of composed query image …
-
Representation Learning for Grounding Vision and Language in Hierarchical Robot Planning
… to enhance representation learning for grounding vision and language in hierarchical robot planning, aiming to improve the full autonomy and robustness of robot systems for indoor household activities. Autonomous robots interact with local operational environments by following the standard …
-
Measuring the robustness and generalization properties of unified vision and language models
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-09-01 without embargo terms
-
Leveraging Multimodal Perspectives to Learn Common Sense for Vision and Language Tasks
Learning and reasoning with common sense is a challenging problem in Artificial Intelligence (AI). Humans have the remarkable ability to interpret images and text from different perspectives in multiple modalities, and to use large amounts of commonsense knowledge while performing visual or textual …
-
Connecting Deep Learning Models to the Human Brain
… models, particularly models that integrate vision and language with human brain processing. These models have shown remarkable advancements in tasks such as object recognition, scene classification, and language processing, achieving near-human accuracy in some cases. This raises intriguing …
-
TOWARDS EFFICIENT TRANSFORMER SCALING
… parameter efficiency using parameter-sharing and Mixture-of-Experts, achieving strong results in both vision and language tasks. Second, we investigate transformer configurations, finding that token-level training benefits from deeper, narrower models, while sequence-level tasks face scaling …
-
Communication-efficient personalization in federated learning for edge devices
… constraints of heterogeneity, limited resources, and diverse data modalities. It develops algorithms that make collaborative training more efficient, robust, and personalized without centralizing data. First, we introduce a dataset-aware dynamic pruning strategy coupled with gradient control to …
-
Spatial Accelerator Generation and Optimization for Tensor Applications
Modern foundation models and generative AI applications require multiple input modalities (both vision and language), which increases the demand for flexible accelerator architecture. Existing frameworks suffer from the trade-off between design flexibility and productivity of RTL generation: either …
-
Understanding The Effects of Incorporating Scientific Knowledge on Neural Network Outputs and Loss Landscapes
… success on several mainstream problems in vision and language modeling, they are still challenged by their lack of interpretable decision-making that is consistent with scientific knowledge, limiting their applicability for scientific discovery applications. Recently, a new field of machine …
-
COMPOSITIONAL GENERALIZATION IN INSTRUCTION FOLLOWING TASKS
Understanding instructions expressed in natural language is a fundamental task in artificialintelligence. A key feature of natural language that humans use while giving and following instructions is compositionality: the capacity to understand and produce a potentially infinite number of novel …
-
Understanding and Improving Representational Robustness of Machine Learning Models
… amount of attention from both academia and the public. In this thesis, we will do a systematic study on the understanding and improvement of several machine learning models, including smoothed models and generic representation networks. Specifically, we put our focus on studying …
-
Travelling in Language and Image: Representing Alterity in French Romantic Fiction, Travel Narrative, Journalism, and Art, 1830--1870
… the importance of theorizing about ""pictures"" and their relationship to ""words,"" W. J. T. Mitchell states that ""the tensions between visual and verbal representations are inseparable from struggles in cultural politics and political culture"" (3). The present study examines the implications …
-
Learning 3D Robotics Perception using Inductive Priors
… the potential to ingest a large amount of data and be really good at performing digital tasks such as text-to-image generation, machine-human conversation, and image recognition. This thesis covers the topic of learning with structured inductive bias and priors to design approaches and …
-
Learning Transferable Representations
… of this thesis is to propose causality as a language for problems of distribution shift. First, we consider domain generalisation, where no data from the test distribution are observed during training. What assumptions can be made regarding the relation between train and test distributions …
-
Multimodal Representation Learning for Agentic AI Systems
… transform the scientific process, from ideation and experimentation to peer review. Many researchers posit that emerging generalist AI “agents” will soon no longer be mere tools, but equal partners in scientific exploration. In this work, we contribute to this evolving landscape through …
Page 1 of 2