{"id":{"repo_id":"uoit","oai_identifier":"oai:ontariotechu.scholaris.ca:10155/1904"},"canonical_url":"https://search.dev.ndltd.org/etd/uoit/oai:ontariotechu.scholaris.ca:10155/1904","repository":{"repo_id":"uoit","name":"Ontario Institute of Technology","base_url":"https://ontariotechu.scholaris.ca/server/oai/request"},"display":{"title":"Hybrid ConVIRT - enhancing medical image-text representation learning of vision language models","abstract":"Challenges in medical imaging are being addressed through advancements in image-text representation learning, as demonstrated by Hybrid-ConVIRT, which builds on contrastive learning frameworks such as ConVIRT and MedCLIP. These medical contrastive learning models trained on domain-specific datasets, have tackled issues related to the costly, expert-annotated datasets typically required in traditional medical imaging. These models serve as foundational, general purpose models that, once trained on vast amounts of data, can be adapted for various downstream tasks. Hybrid-ConVIRT further improves upon this by incorporating image transformations during training and employing CXR-BERT-specialized as the text encoder, specifically tailored for chest Xray reports. To optimize feature alignment, it introduces a Fused Similarity Matrix (FSM), which serves as a more refined target, leveraging the learned capabilities of Med-CLIP and ConVIRT. Evaluation across four datasets (COVID, Tuberculosis, Pneumonia/TB Mix, and CheXpert) using linear probe and zero-shot classification demonstrates Hybrid-ConVIRT’s enhanced discriminative power and feature robustness, particularly on COVID and CheXpert datasets. Although MedCLIP shows a slight advantage on the Tuberculosis dataset, this is likely due to its larger training set. The Hybrid model’s performance supports the potential of combining the learned representations from both preceding models.","abstract_html":"Challenges in medical imaging are being addressed through advancements in image-text representation learning, as demonstrated by Hybrid-ConVIRT, which builds on contrastive learning frameworks such as ConVIRT and MedCLIP. These medical contrastive learning models trained on domain-specific datasets, have tackled issues related to the costly, expert-annotated datasets typically required in traditional medical imaging. These models serve as foundational, general purpose models that, once trained on vast amounts of data, can be adapted for various downstream tasks. Hybrid-ConVIRT further improves upon this by incorporating image transformations during training and employing CXR-BERT-specialized as the text encoder, specifically tailored for chest Xray reports. To optimize feature alignment, it introduces a Fused Similarity Matrix (FSM), which serves as a more refined target, leveraging the learned capabilities of Med-CLIP and ConVIRT. Evaluation across four datasets (COVID, Tuberculosis, Pneumonia/TB Mix, and CheXpert) using linear probe and zero-shot classification demonstrates Hybrid-ConVIRT’s enhanced discriminative power and feature robustness, particularly on COVID and CheXpert datasets. Although MedCLIP shows a slight advantage on the Tuberculosis dataset, this is likely due to its larger training set. The Hybrid model’s performance supports the potential of combining the learned representations from both preceding models.","abstract_has_math":false,"creators":["Sharma, Abhinav"],"institution":"University of Ontario Institute of Technology","degree_name":"Master of Science (MSc)","degree_level":null,"degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Ebrahimi, Mehran"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-12-01","date_published":"2024-12-01","updated_at":"2026-07-24T05:35:39Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10155/1904","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Ebrahimi, Mehran"]},{"key":"dc:creator","label":"Author","values":["Sharma, Abhinav"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-03-18T19:41:13Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-03-18T19:41:13Z"]},{"key":"dc:date.issued","label":"Date","values":["2024-12-01"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MSc)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Ontario Institute of Technology"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10155/1904"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Challenges in medical imaging are being addressed through advancements in image-text representation learning, as demonstrated by Hybrid-ConVIRT, which builds on contrastive learning frameworks such as ConVIRT and MedCLIP. These medical contrastive learning models trained on domain-specific datasets, have tackled issues related to the costly, expert-annotated datasets typically required in traditional medical imaging. These models serve as foundational, general purpose models that, once trained on vast amounts of data, can be adapted for various downstream tasks. Hybrid-ConVIRT further improves upon this by incorporating image transformations during training and employing CXR-BERT-specialized as the text encoder, specifically tailored for chest Xray reports. To optimize feature alignment, it introduces a Fused Similarity Matrix (FSM), which serves as a more refined target, leveraging the learned capabilities of Med-CLIP and ConVIRT. Evaluation across four datasets (COVID, Tuberculosis, Pneumonia/TB Mix, and CheXpert) using linear probe and zero-shot classification demonstrates Hybrid-ConVIRT’s enhanced discriminative power and feature robustness, particularly on COVID and CheXpert datasets. Although MedCLIP shows a slight advantage on the Tuberculosis dataset, this is likely due to its larger training set. The Hybrid model’s performance supports the potential of combining the learned representations from both preceding models."]},{"key":"dc:title","label":"Title","values":["Hybrid ConVIRT - enhancing medical image-text representation learning of vision language models"]}]}],"canonical_facts":{"dc:contributor.advisor":["Ebrahimi, Mehran"],"dc:creator":["Sharma, Abhinav"],"dc:date.accessioned":["2025-03-18T19:41:13Z"],"dc:date.available":["2025-03-18T19:41:13Z"],"dc:date.issued":["2024-12-01"],"dc:description.abstract":["Challenges in medical imaging are being addressed through advancements in image-text representation learning, as demonstrated by Hybrid-ConVIRT, which builds on contrastive learning frameworks such as ConVIRT and MedCLIP. These medical contrastive learning models trained on domain-specific datasets, have tackled issues related to the costly, expert-annotated datasets typically required in traditional medical imaging. These models serve as foundational, general purpose models that, once trained on vast amounts of data, can be adapted for various downstream tasks. Hybrid-ConVIRT further improves upon this by incorporating image transformations during training and employing CXR-BERT-specialized as the text encoder, specifically tailored for chest Xray reports. To optimize feature alignment, it introduces a Fused Similarity Matrix (FSM), which serves as a more refined target, leveraging the learned capabilities of Med-CLIP and ConVIRT. Evaluation across four datasets (COVID, Tuberculosis, Pneumonia/TB Mix, and CheXpert) using linear probe and zero-shot classification demonstrates Hybrid-ConVIRT’s enhanced discriminative power and feature robustness, particularly on COVID and CheXpert datasets. Although MedCLIP shows a slight advantage on the Tuberculosis dataset, this is likely due to its larger training set. The Hybrid model’s performance supports the potential of combining the learned representations from both preceding models."],"dc:identifier.uri":["https://hdl.handle.net/10155/1904"],"dc:language.iso":["en"],"dc:title":["Hybrid ConVIRT - enhancing medical image-text representation learning of vision language models"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_name":["Master of Science (MSc)"],"thesis:institution_name":["University of Ontario Institute of Technology"]},"updated_at":"2026-07-24T05:35:39Z"}