Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 16 of 16 for “"Multimodal Fusion"”.

  1. Multimodal Fusion With Applications to Audio -Visual Speech Recognition

    … and in multichannel biometrics defy a universal fusion method for both applications. For audio-visual speech modeling, we propose a novel sensory fusion method based on the coupled hidden Markov models (CHMMs). The CHMM framework allows the fusion of two temporally coupled information sources to …

    uiuc Repository record for Multimodal Fusion With Applications to Audio -Visual Speech Recognition (opens in a new tab)

  2. A Multi-head Attention Approach with Complementary Multimodal Fusion for Vehicle Detection

    … the development of an improved version of the Multimodal Vehicle Detection Network (MVDNet), distinguished by the integration of a multi-head attention layer. This key enhancement significantly refines the network's capability to process and integrate multimodal sensor data, an aspect that …

    iupui Repository record for A Multi-head Attention Approach with Complementary Multimodal Fusion for Vehicle Detection (opens in a new tab)

  3. Information Fusion for Robust Audio -Visual Speech Recognition

    … and effortlessly perform sensory information fusion. The most important human-to-human communication tool is speech. Automatic speech recognition by machines has been able to achieve very high recognition accuracy for large vocabulary sets and speaker-independent tasks. However, in some …

    uiuc Repository record for Information Fusion for Robust Audio -Visual Speech Recognition (opens in a new tab)

  4. Intuitive Audio Interaction and Control in Multi-Source Environments

    … audio control. Future work should refine multimodal fusion, improve segmentation accuracy, and enhance accessibility to create systems that dynamically respond to users’ natural behaviors—reducing cognitive strain and enabling more fluid, user-centric auditory experiences.

    mit Repository record for Intuitive Audio Interaction and Control in Multi-Source Environments (opens in a new tab)

  5. Healthcare Agents: Large Language Models in Health Prediction and Decision-Making

    … in healthcare AI: (1) leveraging LLMs for multimodal health prediction from wearable sensor data and (2) developing collaborative AI framework for medical decision-making. We first introduce a Health-LLM framework that performs multimodal fusion of temporal physiological signals from …

    mit Repository record for Healthcare Agents: Large Language Models in Health Prediction and Decision-Making (opens in a new tab)

  6. Advancing Chart Question Answering with Robust Chart Component Recognition

    … Co-Attention (QDCAt), which facilitates multimodal fusion by incorporating question information into a deformable offset network and enhancing visual representation from ChartFormer through a deformable co-attention block.

    vt Repository record for Advancing Chart Question Answering with Robust Chart Component Recognition (opens in a new tab)

  7. Grounded SCAN Human: A Benchmark for Zero-Shot Generalizations

    … replacing the encoder, and one with early multimodal fusion of the sentence encoding with the visual embedding. We also test a multimodal transformer similar to VilBERT, which is the state of the art on the original gSCAN splits. We find that the models are somewhat robust to varying …

    mit Repository record for Grounded SCAN Human: A Benchmark for Zero-Shot Generalizations (opens in a new tab)

  8. Multimodal Non-Contact Sensing of Neonatal Vital Signs Using Radar and Video

    … This thesis establishes the foundation for a multimodal device designed for noncontact monitoring of neonates in the Neonatal Intensive Care Unit (NICU) that integrates a video camera and a radar. The device is used to estimate vital signs such as respiratory rate (RR), using both unimodal …

    mit Repository record for Multimodal Non-Contact Sensing of Neonatal Vital Signs Using Radar and Video (opens in a new tab)

  9. Wearable Gut and Brain Interfaces for Valence Detection and Modulation

    … activity (EDA) and respiration rate multimodal fusion model. I also present Serosa, an novel electrogastrography (EGG) GBCI which non-invasively records indices of gastric neurons that can be correlated with emotional states and provide a new affect detection modality. This thesis …

    mit Repository record for Wearable Gut and Brain Interfaces for Valence Detection and Modulation (opens in a new tab)

  10. Fuse and Adapt: Investigating the Use of Pre-Trained Self-Supervising Learning Models in Limited Data NLU problems

    … pre-trained SSL models in the two main areas of multimodal fusion and domain adaptation. Under these two main topics, I explored four research question that introduces novel fusion and adaptation techniques.

    auckland-ms Repository record for Fuse and Adapt: Investigating the Use of Pre-Trained Self-Supervising Learning Models in Limited Data NLU problems (opens in a new tab)

  11. Machine learning to model health with multimodal mobile sensor data

    … distributions of behaviors, lack of labels, and multimodality. This dissertation addresses these challenges by developing new models that leverage multi-task learning for accurate forecasting, multimodal fusion for improved population subtyping, and self-supervision for learning generalized …

    cambridge Repository record for Machine learning to model health with multimodal mobile sensor data (opens in a new tab)

  12. Learning Representations for Limited and Heterogeneous Medical Data

    … and integrate the corresponding metadata as a multimodal resource to introduce inductive biases. We find that the representations learned by the developed approach yield better downstream task performance, such as ultrasound image quality classification and organ segmentation, compared with the …

    mit Repository record for Learning Representations for Limited and Heterogeneous Medical Data (opens in a new tab)

  13. Text Prompt-Driven Medical Image Segmentation

    … make the direct transfer of natural image multimodal techniques to the medical domain suboptimal. Therefore, developing domain-specific vision-language segmentation frameworks tailored to medical images is of urgent importance. This thesis proposes three multimodal segmentation methods …

    unsw Repository record for Text Prompt-Driven Medical Image Segmentation (opens in a new tab)

  14. Multi-modal Multi-Level Neuroimaging Fusion with Modality-Aware Mask-Guided Attention and Deep Canonical Correlation Analysis to Improve Dementia Risk Prediction

    … planning. This thesis presents a novel multimodal deep learning framework that integrates T1-weighted MRI and Amyloid PET imaging to improve the diagnosis and stratification of AD. The proposed architecture leverages a two-stage pipeline involving modality-specific feature extraction …

    vt Repository record for Multi-modal Multi-Level Neuroimaging Fusion with Modality-Aware Mask-Guided Attention and Deep Canonical Correlation Analysis to Improve Dementia Risk Prediction (opens in a new tab)

  15. Advanced boundary-enhanced instance segmentation and spatial-temporal transformer models for automated schizophrenic investigation

    lethbridge