Back to results

UNSW, Sydney

Text Prompt-Driven Medical Image Segmentation

Abstract

dc:description

Medical image segmentation plays a crucial role in accurate diagnosis, treatment planning, and surgical navigation by precisely identifying pathological regions. However, traditional segmentation methods typically rely on dense pixel-level annotations and heavy computational resources, which pose significant challenges for real-world clinical applications. Vision-language segmentation has emerged as a promising solution by introducing rich semantic knowledge from textual descriptions as auxiliary supervision. This integration reduces dependence on manual annotation and computational demand. Although this paradigm has shown impressive success in natural image tasks, its application in medical imaging remains limited due to the unique complexity of anatomical structures, fine-grained semantics, and modality discrepancies. These factors make the direct transfer of natural image multimodal techniques to the medical domain suboptimal. Therefore, developing domain-specific vision-language segmentation frameworks tailored to medical images is of urgent importance. This thesis proposes three multimodal segmentation methods targeting fully supervised and weakly supervised settings, with each method addressing specific bottlenecks. Our main contribution lies in the use of text-driven guidance to solve the respective challenges: computational inefficiency in full supervision and insufficient supervision information in weak supervision settings. The first study addresses the problem of computational cost and training complexity in fully supervised medical segmentation. We develop a lightweight, text-guided framework that fine-tunes a small number of parameters within a CLIP-based structure and incorporates MedSAM2 for mask generation. This approach eliminates the need for heavy architectural components such as large attention modules. Experimental results demonstrate a 70% reduction in trainable parameters without sacrificing segmentation performance, offering a deployable solution for clinical environments. The second study focuses on the challenge of supervisory information deficiency in weakly supervised learning. We propose a pixel-level classification method that computes semantic similarity between image patches and image-level textual labels. This method enables the extraction of discriminative knowledge from complex morphological descriptions and significantly improves mask quality by addressing the typical limitation of CAM-based methods, which tend to highlight only the most discriminative regions. Our segmentation results achieve superior mIoU performance, outperforming the second-best approach by 1.75% on LUAD-HistoSeg and 1.58% on BCSS-WSSS. Building upon this, the third study introduces a novel two-stage weakly supervised framework. In the first stage, coarse pseudo masks are generated using image-level labels and enhanced by a channel reactivation mechanism to highlight previously ignored regions. In the second stage, refined segmentation is achieved through more complex long-text prompts and a Fourier-based multimodal fusion module. This progressive training strategy gradually injects semantic knowledge, improving both boundary accuracy and contextual understanding. Our final segmentation results significantly refine the pseudo masks, achieving a 2.32% improvement on LUAD-HistoSeg and a 3.15% improvement on BCSS-WSSS compared with the pseudo-mask baseline. Meanwhile, a comparative analysis with the second study shows that the two-stage approach achieves an improvement of 2.33% on LUAD-HistoSeg and 3.62% on BCSS-WSSS, validating the effectiveness of the proposed framework. This research presents a unified exploration of vision-language segmentation across different supervision levels, demonstrating that textual guidance can effectively alleviate the limitations of both fully supervised and weakly supervised methods. By aligning methodological design with the structural and semantic characteristics of medical data, this work extends the frontier of multimodal medical image segmentation and offers practical and scalable frameworks for real-world applications.

Degree

thesis:*
Grantor dc:publisher
UNSW, Sydney
Year dc:date
2026

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Gao, Sicong

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • open access
  • CC BY 4.0
  • free_to_read
Language dc:language
en

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:unsworks.library.unsw.edu.au:1959.4/107240

Chain of custody

source
Harvested from
University of New South Wales
Base URL
unsworks.unsw.edu.au/oai/provider
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Gao, Sicong. Text Prompt-Driven Medical Image Segmentation. UNSW, Sydney, 2026. http://hdl.handle.net/1959.4/107240