University of Ontario Institute of Technology
Hybrid ConVIRT - enhancing medical image-text representation learning of vision language models
Abstract
dc:description.abstractChallenges in medical imaging are being addressed through advancements in image-text representation learning, as demonstrated by Hybrid-ConVIRT, which builds on contrastive learning frameworks such as ConVIRT and MedCLIP. These medical contrastive learning models trained on domain-specific datasets, have tackled issues related to the costly, expert-annotated datasets typically required in traditional medical imaging. These models serve as foundational, general purpose models that, once trained on vast amounts of data, can be adapted for various downstream tasks. Hybrid-ConVIRT further improves upon this by incorporating image transformations during training and employing CXR-BERT-specialized as the text encoder, specifically tailored for chest Xray reports. To optimize feature alignment, it introduces a Fused Similarity Matrix (FSM), which serves as a more refined target, leveraging the learned capabilities of Med-CLIP and ConVIRT. Evaluation across four datasets (COVID, Tuberculosis, Pneumonia/TB Mix, and CheXpert) using linear probe and zero-shot classification demonstrates Hybrid-ConVIRT’s enhanced discriminative power and feature robustness, particularly on COVID and CheXpert datasets. Although MedCLIP shows a slight advantage on the Tuberculosis dataset, this is likely due to its larger training set. The Hybrid model’s performance supports the potential of combining the learned representations from both preceding models.
Degree
thesis:*- Name thesis:degree_name
- Master of Science (MSc)
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Ontario Institute of Technology
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Sharma, Abhinav
- Advisor dc:contributor.advisor
-
- Ebrahimi, Mehran
Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/10155/1904
- OAI identifier oai:identifier
- oai:ontariotechu.scholaris.ca:10155/1904