Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 35 for “"Vision and Language"”.

  1. Representations from vision and language

    Replicating a human-level understanding of the physical world in computers is a monumental task. Achieving this requires building representations of concepts that manifest themselves visually, linguistically or through other senses. Furthermore concepts do not exist in isolation but are related to …

    uiuc Repository record for Representations from vision and language (opens in a new tab)

  2. Vision and Language: Information Integration and Transformation

    Images and text are large-scale data sources for deep learning models to mimic how humans perceive and understand multimodal information. Combining images and text can construct better feature representations by involving complementary information from different sources. For example, text …

    rice Repository record for Vision and Language: Information Integration and Transformation (opens in a new tab)

  3. Generative and Multimodal Learning for Vision and Language

    L'abstract è presente nell'allegato / the abstract is in the attachment

    poli-torino Repository record for Generative and Multimodal Learning for Vision and Language (opens in a new tab)

  4. Cognitive Map Generation for Vision and Language Navigation

    <p>Visual-Language Navigation (VLN) presents significant challenges for autonomous agents, such as robots and virtual assistants, particularly in complex, dynamic environments where the seamless integration of visual perception and natural language understanding is critical. Traditional VLN systems …

    cuny Repository record for Cognitive Map Generation for Vision and Language Navigation (opens in a new tab)

  5. Unifying cross-modal concepts in vision and language

    … computers to demonstrate a proficient understanding of the physical world is an exceedingly challenging task that necessitates the ability to perceive, through vision or other senses, and communicate through natural language. Key to this endeavor is the representation of concepts present in …

    uiuc Repository record for Unifying cross-modal concepts in vision and language (opens in a new tab)

  6. Connecting vision and language via image retrieval and captioning

    Many real-world problems involve jointly understanding vision and language, e.g. image/video captioning, multi-modal image retrieval, visual question answering. In this thesis, we consider several problems in cross-modal learning from vision and language. First, the problem of composed query image …

    manitoba Repository record for Connecting vision and language via image retrieval and captioning (opens in a new tab)

  7. Representation Learning for Grounding Vision and Language in Hierarchical Robot Planning

    … to enhance representation learning for grounding vision and language in hierarchical robot planning, aiming to improve the full autonomy and robustness of robot systems for indoor household activities. Autonomous robots interact with local operational environments by following the standard …

    gatech Repository record for Representation Learning for Grounding Vision and Language in Hierarchical Robot Planning (opens in a new tab)

  8. Measuring the robustness and generalization properties of unified vision and language models

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-09-01 without embargo terms

    uiuc Repository record for Measuring the robustness and generalization properties of unified vision and language models (opens in a new tab)

  9. Leveraging Multimodal Perspectives to Learn Common Sense for Vision and Language Tasks

    Learning and reasoning with common sense is a challenging problem in Artificial Intelligence (AI). Humans have the remarkable ability to interpret images and text from different perspectives in multiple modalities, and to use large amounts of commonsense knowledge while performing visual or textual …

    vt Repository record for Leveraging Multimodal Perspectives to Learn Common Sense for Vision and Language Tasks (opens in a new tab)

  10. Connecting Deep Learning Models to the Human Brain

    … models, particularly models that integrate vision and language with human brain processing. These models have shown remarkable advancements in tasks such as object recognition, scene classification, and language processing, achieving near-human accuracy in some cases. This raises intriguing …

    mit Repository record for Connecting Deep Learning Models to the Human Brain (opens in a new tab)

  11. TOWARDS EFFICIENT TRANSFORMER SCALING

    … parameter efficiency using parameter-sharing and Mixture-of-Experts, achieving strong results in both vision and language tasks. Second, we investigate transformer configurations, finding that token-level training benefits from deeper, narrower models, while sequence-level tasks face scaling …

    nus Repository record for TOWARDS EFFICIENT TRANSFORMER SCALING (opens in a new tab)

  12. Communication-efficient personalization in federated learning for edge devices

    … constraints of heterogeneity, limited resources, and diverse data modalities. It develops algorithms that make collaborative training more efficient, robust, and personalized without centralizing data. First, we introduce a dataset-aware dynamic pruning strategy coupled with gradient control to …

    iastate Repository record for Communication-efficient personalization in federated learning for edge devices (opens in a new tab)

  13. Spatial Accelerator Generation and Optimization for Tensor Applications

    Modern foundation models and generative AI applications require multiple input modalities (both vision and language), which increases the demand for flexible accelerator architecture. Existing frameworks suffer from the trade-off between design flexibility and productivity of RTL generation: either …

    mit Repository record for Spatial Accelerator Generation and Optimization for Tensor Applications (opens in a new tab)

  14. Understanding The Effects of Incorporating Scientific Knowledge on Neural Network Outputs and Loss Landscapes

    … success on several mainstream problems in vision and language modeling, they are still challenged by their lack of interpretable decision-making that is consistent with scientific knowledge, limiting their applicability for scientific discovery applications. Recently, a new field of machine …

    vt Repository record for Understanding The Effects of Incorporating Scientific Knowledge on Neural Network Outputs and Loss Landscapes (opens in a new tab)

  15. COMPOSITIONAL GENERALIZATION IN INSTRUCTION FOLLOWING TASKS

    Understanding instructions expressed in natural language is a fundamental task in artificialintelligence. A key feature of natural language that humans use while giving and following instructions is compositionality: the capacity to understand and produce a potentially infinite number of novel …

    penn Repository record for COMPOSITIONAL GENERALIZATION IN INSTRUCTION FOLLOWING TASKS (opens in a new tab)

  16. Understanding and Improving Representational Robustness of Machine Learning Models

    … amount of attention from both academia and the public. In this thesis, we will do a systematic study on the understanding and improvement of several machine learning models, including smoothed models and generic representation networks. Specifically, we put our focus on studying …

    mit Repository record for Understanding and Improving Representational Robustness of Machine Learning Models (opens in a new tab)

  17. Travelling in Language and Image: Representing Alterity in French Romantic Fiction, Travel Narrative, Journalism, and Art, 1830--1870

    … the importance of theorizing about ""pictures"" and their relationship to ""words,"" W. J. T. Mitchell states that ""the tensions between visual and verbal representations are inseparable from struggles in cultural politics and political culture"" (3). The present study examines the implications …

    uiuc Repository record for Travelling in Language and Image: Representing Alterity in French Romantic Fiction, Travel Narrative, Journalism, and Art, 1830--1870 (opens in a new tab)

  18. Learning 3D Robotics Perception using Inductive Priors

    … the potential to ingest a large amount of data and be really good at performing digital tasks such as text-to-image generation, machine-human conversation, and image recognition. This thesis covers the topic of learning with structured inductive bias and priors to design approaches and

    gatech Repository record for Learning 3D Robotics Perception using Inductive Priors (opens in a new tab)

  19. Learning Transferable Representations

    … of this thesis is to propose causality as a language for problems of distribution shift. First, we consider domain generalisation, where no data from the test distribution are observed during training. What assumptions can be made regarding the relation between train and test distributions …

    cambridge Repository record for Learning Transferable Representations (opens in a new tab)

  20. Multimodal Representation Learning for Agentic AI Systems

    … transform the scientific process, from ideation and experimentation to peer review. Many researchers posit that emerging generalist AI “agents” will soon no longer be mere tools, but equal partners in scientific exploration. In this work, we contribute to this evolving landscape through …

    mit Repository record for Multimodal Representation Learning for Agentic AI Systems (opens in a new tab)

Page 1 of 2