Back to results

University of Ontario Institute of Technology

Hybrid ConVIRT - enhancing medical image-text representation learning of vision language models

Abstract

dc:description.abstract

Challenges in medical imaging are being addressed through advancements in image-text representation learning, as demonstrated by Hybrid-ConVIRT, which builds on contrastive learning frameworks such as ConVIRT and MedCLIP. These medical contrastive learning models trained on domain-specific datasets, have tackled issues related to the costly, expert-annotated datasets typically required in traditional medical imaging. These models serve as foundational, general purpose models that, once trained on vast amounts of data, can be adapted for various downstream tasks. Hybrid-ConVIRT further improves upon this by incorporating image transformations during training and employing CXR-BERT-specialized as the text encoder, specifically tailored for chest Xray reports. To optimize feature alignment, it introduces a Fused Similarity Matrix (FSM), which serves as a more refined target, leveraging the learned capabilities of Med-CLIP and ConVIRT. Evaluation across four datasets (COVID, Tuberculosis, Pneumonia/TB Mix, and CheXpert) using linear probe and zero-shot classification demonstrates Hybrid-ConVIRT’s enhanced discriminative power and feature robustness, particularly on COVID and CheXpert datasets. Although MedCLIP shows a slight advantage on the Tuberculosis dataset, this is likely due to its larger training set. The Hybrid model’s performance supports the potential of combining the learned representations from both preceding models.

Degree

thesis:*
Name thesis:degree_name
Master of Science (MSc)
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Ontario Institute of Technology
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Sharma, Abhinav
Advisor dc:contributor.advisor
  • Ebrahimi, Mehran

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/10155/1904
OAI identifier oai:identifier
oai:ontariotechu.scholaris.ca:10155/1904

Chain of custody

source
Harvested from
Ontario Institute of Technology
Base URL
ontariotechu.scholaris.ca/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Sharma, Abhinav. Hybrid ConVIRT - enhancing medical image-text representation learning of vision language models. University of Ontario Institute of Technology, 2024. https://hdl.handle.net/10155/1904