{"id":{"repo_id":"middlesex","oai_identifier":"oai:repository.mdx.ac.uk:116z34"},"canonical_url":"https://search.dev.ndltd.org/etd/middlesex/oai:repository.mdx.ac.uk:116z34","repository":{"repo_id":"middlesex","name":"Middlesex University","base_url":"https://repository.mdx.ac.uk/oai2"},"display":{"title":"Latent diffusion for generative visual attribution in medical image diagnostics","abstract":"Visual attribution in medical imaging seeks to make evident the diagnostically-relevant components of a medical image, in contrast to the more common detection of diseased tissue deployed in conventional machine vision pipelines (due to the inherent learning nature of these latter models, they are typically not easily interpretable/explainable to clinicians). State-of-the-art techniques in visual attribution generally consist of different variants of deep neural networks, implemented as classifiers, or segmenters. However, they have not thus far included an explicit linguistic component. We here present a novel generative visual attribution technique, one that leverages latent diffusion models in combination with domain-specific large language models, in order to generate normal counterparts of abnormal images. The discrepancy between the two hence gives rise to a mapping indicating the diagnostically-relevant image components. To achieve this, we deploy image priors in conjunction with appropriate conditioning mechanisms in order to control the image generative process, including natural language text prompts acquired from medical science and applied radiology. We perform experiments and quantitatively evaluate our results on the COVID-19 Radiography Database containing labelled chest X-rays with differing pathologies via the Frechet Inception Distance (FID), Structural Similarity (SSIM) and Multi Scale Structural Similarity Metric (MS-SSIM) metrics obtained between real and generated images. The resulting system also exhibits a range of latent capabilities including super-resolution and zero-shot localized disease induction, which are evaluated with real examples from the cheXpert dataset.","abstract_html":"Visual attribution in medical imaging seeks to make evident the diagnostically-relevant components of a medical image, in contrast to the more common detection of diseased tissue deployed in conventional machine vision pipelines (due to the inherent learning nature of these latter models, they are typically not easily interpretable/explainable to clinicians). State-of-the-art techniques in visual attribution generally consist of different variants of deep neural networks, implemented as classifiers, or segmenters. However, they have not thus far included an explicit linguistic component. We here present a novel generative visual attribution technique, one that leverages latent diffusion models in combination with domain-specific large language models, in order to generate normal counterparts of abnormal images. The discrepancy between the two hence gives rise to a mapping indicating the diagnostically-relevant image components. To achieve this, we deploy image priors in conjunction with appropriate conditioning mechanisms in order to control the image generative process, including natural language text prompts acquired from medical science and applied radiology. We perform experiments and quantitatively evaluate our results on the COVID-19 Radiography Database containing labelled chest X-rays with differing pathologies via the Frechet Inception Distance (FID), Structural Similarity (SSIM) and Multi Scale Structural Similarity Metric (MS-SSIM) metrics obtained between real and generated images. The resulting system also exhibits a range of latent capabilities including super-resolution and zero-shot localized disease induction, which are evaluated with real examples from the cheXpert dataset.","abstract_has_math":false,"creators":["Siddiqui, A."],"institution":"Middlesex University","degree_name":null,"degree_level":"Masters thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023","date_published":"2023","updated_at":"2026-07-24T03:03:41Z","subjects":["Visual Attribution","Explainable AI","Diffusion models","Medical imaging"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:repository.mdx.ac.uk:116z34"],"render_values":[{"text":"oai:repository.mdx.ac.uk:116z34","href":null,"code":true}]}]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Siddiqui, A."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023"]},{"key":"dc:date.issued","label":"Date","values":["2023"]},{"key":"dc:publisher","label":"Institution","values":["Middlesex University Research Repository"]},{"key":"dc:publisher.department","label":"Dc Publisher Department","values":["Computer Science","Science and Technology"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["Middlesex University"]},{"key":"dc:relation","label":"Dc Relation","values":["https://repository.mdx.ac.uk/item/116z34"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://repository.mdx.ac.uk/item/116z34"]},{"key":"dc:type","label":"Dc Type","values":["Thesis or dissertation"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Masters thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Visual Attribution","Explainable AI","Diffusion models","Medical imaging"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:repository.mdx.ac.uk:116z34"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://repository.mdx.ac.uk/download/a973cc0435076b4bdcfef3484209c06d8e600e4ee42a3e0f553178751d563106/13442968/AASiddiqui%20thesis%20EMBARGO.pdf"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Visual attribution in medical imaging seeks to make evident the diagnostically-relevant components of a medical image, in contrast to the more common detection of diseased tissue deployed in conventional machine vision pipelines (due to the inherent learning nature of these latter models, they are typically not easily interpretable/explainable to clinicians). State-of-the-art techniques in visual attribution generally consist of different variants of deep neural networks, implemented as classifiers, or segmenters. However, they have not thus far included an explicit linguistic component. We here present a novel generative visual attribution technique, one that leverages latent diffusion models in combination with domain-specific large language models, in order to generate normal counterparts of abnormal images. The discrepancy between the two hence gives rise to a mapping indicating the diagnostically-relevant image components. To achieve this, we deploy image priors in conjunction with appropriate conditioning mechanisms in order to control the image generative process, including natural language text prompts acquired from medical science and applied radiology. We perform experiments and quantitatively evaluate our results on the COVID-19 Radiography Database containing labelled chest X-rays with differing pathologies via the Frechet Inception Distance (FID), Structural Similarity (SSIM) and Multi Scale Structural Similarity Metric (MS-SSIM) metrics obtained between real and generated images. The resulting system also exhibits a range of latent capabilities including super-resolution and zero-shot localized disease induction, which are evaluated with real examples from the cheXpert dataset."]},{"key":"dc:description.abstract","label":"Abstract","values":["Visual attribution in medical imaging seeks to make evident the diagnostically-relevant components of a medical image, in contrast to the more common detection of diseased tissue deployed in conventional machine vision pipelines (due to the inherent learning nature of these latter models, they are typically not easily interpretable/explainable to clinicians). State-of-the-art techniques in visual attribution generally consist of different variants of deep neural networks, implemented as classifiers, or segmenters. However, they have not thus far included an explicit linguistic component. We here present a novel generative visual attribution technique, one that leverages latent diffusion models in combination with domain-specific large language models, in order to generate normal counterparts of abnormal images. The discrepancy between the two hence gives rise to a mapping indicating the diagnostically-relevant image components. To achieve this, we deploy image priors in conjunction with appropriate conditioning mechanisms in order to control the image generative process, including natural language text prompts acquired from medical science and applied radiology. We perform experiments and quantitatively evaluate our results on the COVID-19 Radiography Database containing labelled chest X-rays with differing pathologies via the Frechet Inception Distance (FID), Structural Similarity (SSIM) and Multi Scale Structural Similarity Metric (MS-SSIM) metrics obtained between real and generated images. The resulting system also exhibits a range of latent capabilities including super-resolution and zero-shot localized disease induction, which are evaluated with real examples from the cheXpert dataset."]},{"key":"dc:title","label":"Title","values":["Latent diffusion for generative visual attribution in medical image diagnostics"]}]}],"canonical_facts":{"dc:creator":["Siddiqui, A."],"dc:date":["2023"],"dc:date.issued":["2023"],"dc:description":["Visual attribution in medical imaging seeks to make evident the diagnostically-relevant components of a medical image, in contrast to the more common detection of diseased tissue deployed in conventional machine vision pipelines (due to the inherent learning nature of these latter models, they are typically not easily interpretable/explainable to clinicians). State-of-the-art techniques in visual attribution generally consist of different variants of deep neural networks, implemented as classifiers, or segmenters. However, they have not thus far included an explicit linguistic component. We here present a novel generative visual attribution technique, one that leverages latent diffusion models in combination with domain-specific large language models, in order to generate normal counterparts of abnormal images. The discrepancy between the two hence gives rise to a mapping indicating the diagnostically-relevant image components. To achieve this, we deploy image priors in conjunction with appropriate conditioning mechanisms in order to control the image generative process, including natural language text prompts acquired from medical science and applied radiology. We perform experiments and quantitatively evaluate our results on the COVID-19 Radiography Database containing labelled chest X-rays with differing pathologies via the Frechet Inception Distance (FID), Structural Similarity (SSIM) and Multi Scale Structural Similarity Metric (MS-SSIM) metrics obtained between real and generated images. The resulting system also exhibits a range of latent capabilities including super-resolution and zero-shot localized disease induction, which are evaluated with real examples from the cheXpert dataset."],"dc:description.abstract":["Visual attribution in medical imaging seeks to make evident the diagnostically-relevant components of a medical image, in contrast to the more common detection of diseased tissue deployed in conventional machine vision pipelines (due to the inherent learning nature of these latter models, they are typically not easily interpretable/explainable to clinicians). State-of-the-art techniques in visual attribution generally consist of different variants of deep neural networks, implemented as classifiers, or segmenters. However, they have not thus far included an explicit linguistic component. We here present a novel generative visual attribution technique, one that leverages latent diffusion models in combination with domain-specific large language models, in order to generate normal counterparts of abnormal images. The discrepancy between the two hence gives rise to a mapping indicating the diagnostically-relevant image components. To achieve this, we deploy image priors in conjunction with appropriate conditioning mechanisms in order to control the image generative process, including natural language text prompts acquired from medical science and applied radiology. We perform experiments and quantitatively evaluate our results on the COVID-19 Radiography Database containing labelled chest X-rays with differing pathologies via the Frechet Inception Distance (FID), Structural Similarity (SSIM) and Multi Scale Structural Similarity Metric (MS-SSIM) metrics obtained between real and generated images. The resulting system also exhibits a range of latent capabilities including super-resolution and zero-shot localized disease induction, which are evaluated with real examples from the cheXpert dataset."],"dc:identifier":["oai:repository.mdx.ac.uk:116z34"],"dc:identifier.uri":["https://repository.mdx.ac.uk/download/a973cc0435076b4bdcfef3484209c06d8e600e4ee42a3e0f553178751d563106/13442968/AASiddiqui%20thesis%20EMBARGO.pdf"],"dc:publisher":["Middlesex University Research Repository"],"dc:publisher.department":["Computer Science","Science and Technology"],"dc:publisher.institution":["Middlesex University"],"dc:relation":["https://repository.mdx.ac.uk/item/116z34"],"dc:relation.isreferencedby":["https://repository.mdx.ac.uk/item/116z34"],"dc:subject":["Visual Attribution","Explainable AI","Diffusion models","Medical imaging"],"dc:title":["Latent diffusion for generative visual attribution in medical image diagnostics"],"dc:type":["Thesis or dissertation"],"dc:type.qualificationlevel":["Masters thesis"]},"updated_at":"2026-07-24T03:03:41Z"}