{"id":{"repo_id":"usm","oai_identifier":"oai:aquila.usm.edu:masters_theses-2290"},"canonical_url":"https://search.dev.ndltd.org/etd/usm/oai:aquila.usm.edu:masters_theses-2290","repository":{"repo_id":"usm","name":"University of Southern Mississippi","base_url":"https://aquila.usm.edu/do/oai/"},"display":{"title":"LLM-as-Judge for Reliable Crack Segmentation in Edge-Based Structural Inspection","abstract":"<p>Crack detection and segmentation are fundamental tasks in structural health monitoring, enabling early identification of damage in critical infrastructure such as roads, bridges, and buildings. While deep learning models—particularly convolutional encoder--decoder architectures—have achieved high accuracy in pixel-level crack segmentation, their deployment in real-world environments introduces significant reliability challenges. In practical scenarios, especially in UAV-based inspection, visual conditions such as illumination variation, motion blur, low resolution, and occlusions can severely degrade segmentation performance. Moreover, the absence of ground-truth annotations during deployment makes conventional evaluation metrics, such as Intersection-over-Union and Dice score, inapplicable, creating a critical gap between model performance and operational trust.</p> <p>This thesis reframes crack segmentation from a purely accuracy-driven task to a reliability-centered problem. First, it provides a comprehensive analysis of crack segmentation methods, highlighting limitations in thin-structure preservation, generalization, and robustness under real-world conditions. Building on these insights, the thesis introduces a novel semantic monitoring framework based on the LLM-as-Judge paradigm. In this framework, a lightweight crack segmentation model operates onboard a UAV, while a multimodal large language model evaluates segmentation outputs using visual reasoning, producing a quality score, confidence estimate, and explanatory feedback without requiring ground-truth annotations.</p> <p>To ensure trustworthiness, a rigorous evaluation methodology is proposed, defining repeatability and sensitivity as key reliability criteria. Extensive experiments under controlled perturbations demonstrate that the proposed framework achieves stable, consistent, and perceptually meaningful evaluations. This work establishes a new direction for crack segmentation by enabling reliable, interpretable, and deployment-ready assessment in safety-critical environments.</p>","abstract_html":"&lt;p&gt;Crack detection and segmentation are fundamental tasks in structural health monitoring, enabling early identification of damage in critical infrastructure such as roads, bridges, and buildings. While deep learning models—particularly convolutional encoder--decoder architectures—have achieved high accuracy in pixel-level crack segmentation, their deployment in real-world environments introduces significant reliability challenges. In practical scenarios, especially in UAV-based inspection, visual conditions such as illumination variation, motion blur, low resolution, and occlusions can severely degrade segmentation performance. Moreover, the absence of ground-truth annotations during deployment makes conventional evaluation metrics, such as Intersection-over-Union and Dice score, inapplicable, creating a critical gap between model performance and operational trust.&lt;/p&gt; &lt;p&gt;This thesis reframes crack segmentation from a purely accuracy-driven task to a reliability-centered problem. First, it provides a comprehensive analysis of crack segmentation methods, highlighting limitations in thin-structure preservation, generalization, and robustness under real-world conditions. Building on these insights, the thesis introduces a novel semantic monitoring framework based on the LLM-as-Judge paradigm. In this framework, a lightweight crack segmentation model operates onboard a UAV, while a multimodal large language model evaluates segmentation outputs using visual reasoning, producing a quality score, confidence estimate, and explanatory feedback without requiring ground-truth annotations.&lt;/p&gt; &lt;p&gt;To ensure trustworthiness, a rigorous evaluation methodology is proposed, defining repeatability and sensitivity as key reliability criteria. Extensive experiments under controlled perturbations demonstrate that the proposed framework achieves stable, consistent, and perceptually meaningful evaluations. This work establishes a new direction for crack segmentation by enabling reliable, interpretable, and deployment-ready assessment in safety-critical environments.&lt;/p&gt;","abstract_has_math":false,"creators":["Hasan, Murad"],"institution":null,"degree_name":"Master of Science (MS)","degree_level":"Masters Thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Dr. Rabab Abdelfattah","Dr. Sarah Lee","Dr. Ahmed Sherif"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-05-01T07:00:00Z","date_published":"2026-05-01T07:00:00Z","updated_at":"2026-07-24T05:45:55Z","subjects":["Crack Segmentation","Structural Health Monitoring","UAV-Based Inspection","LLM-as-Judge","Ground-Truth-Free Evaluation","Infrastructure Inspection","Real-World Vision System Reliability","Artificial Intelligence and Robotics","Data Science"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://aquila.usm.edu/masters_theses/1196","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dr. Rabab Abdelfattah","Dr. Sarah Lee","Dr. Ahmed Sherif"]},{"key":"dc:creator","label":"Author","values":["Hasan, Murad"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2027-05-31T07:00:00Z"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Crack Segmentation","Structural Health Monitoring","UAV-Based Inspection","LLM-as-Judge","Ground-Truth-Free Evaluation","Infrastructure Inspection","Real-World Vision System Reliability","Artificial Intelligence and Robotics","Data Science"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://aquila.usm.edu/masters_theses/1196"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Crack detection and segmentation are fundamental tasks in structural health monitoring, enabling early identification of damage in critical infrastructure such as roads, bridges, and buildings. While deep learning models—particularly convolutional encoder--decoder architectures—have achieved high accuracy in pixel-level crack segmentation, their deployment in real-world environments introduces significant reliability challenges. In practical scenarios, especially in UAV-based inspection, visual conditions such as illumination variation, motion blur, low resolution, and occlusions can severely degrade segmentation performance. Moreover, the absence of ground-truth annotations during deployment makes conventional evaluation metrics, such as Intersection-over-Union and Dice score, inapplicable, creating a critical gap between model performance and operational trust.</p> <p>This thesis reframes crack segmentation from a purely accuracy-driven task to a reliability-centered problem. First, it provides a comprehensive analysis of crack segmentation methods, highlighting limitations in thin-structure preservation, generalization, and robustness under real-world conditions. Building on these insights, the thesis introduces a novel semantic monitoring framework based on the LLM-as-Judge paradigm. In this framework, a lightweight crack segmentation model operates onboard a UAV, while a multimodal large language model evaluates segmentation outputs using visual reasoning, producing a quality score, confidence estimate, and explanatory feedback without requiring ground-truth annotations.</p> <p>To ensure trustworthiness, a rigorous evaluation methodology is proposed, defining repeatability and sensitivity as key reliability criteria. Extensive experiments under controlled perturbations demonstrate that the proposed framework achieves stable, consistent, and perceptually meaningful evaluations. This work establishes a new direction for crack segmentation by enabling reliable, interpretable, and deployment-ready assessment in safety-critical environments.</p>"]},{"key":"dc:title","label":"Title","values":["LLM-as-Judge for Reliable Crack Segmentation in Edge-Based Structural Inspection"]}]}],"canonical_facts":{"dc:contributor":["Dr. Rabab Abdelfattah","Dr. Sarah Lee","Dr. Ahmed Sherif"],"dc:creator":["Hasan, Murad"],"dc:date.available":["2027-05-31T07:00:00Z"],"dc:description.abstract":["<p>Crack detection and segmentation are fundamental tasks in structural health monitoring, enabling early identification of damage in critical infrastructure such as roads, bridges, and buildings. While deep learning models—particularly convolutional encoder--decoder architectures—have achieved high accuracy in pixel-level crack segmentation, their deployment in real-world environments introduces significant reliability challenges. In practical scenarios, especially in UAV-based inspection, visual conditions such as illumination variation, motion blur, low resolution, and occlusions can severely degrade segmentation performance. Moreover, the absence of ground-truth annotations during deployment makes conventional evaluation metrics, such as Intersection-over-Union and Dice score, inapplicable, creating a critical gap between model performance and operational trust.</p> <p>This thesis reframes crack segmentation from a purely accuracy-driven task to a reliability-centered problem. First, it provides a comprehensive analysis of crack segmentation methods, highlighting limitations in thin-structure preservation, generalization, and robustness under real-world conditions. Building on these insights, the thesis introduces a novel semantic monitoring framework based on the LLM-as-Judge paradigm. In this framework, a lightweight crack segmentation model operates onboard a UAV, while a multimodal large language model evaluates segmentation outputs using visual reasoning, producing a quality score, confidence estimate, and explanatory feedback without requiring ground-truth annotations.</p> <p>To ensure trustworthiness, a rigorous evaluation methodology is proposed, defining repeatability and sensitivity as key reliability criteria. Extensive experiments under controlled perturbations demonstrate that the proposed framework achieves stable, consistent, and perceptually meaningful evaluations. This work establishes a new direction for crack segmentation by enabling reliable, interpretable, and deployment-ready assessment in safety-critical environments.</p>"],"dc:identifier":["https://aquila.usm.edu/masters_theses/1196"],"dc:subject":["Crack Segmentation","Structural Health Monitoring","UAV-Based Inspection","LLM-as-Judge","Ground-Truth-Free Evaluation","Infrastructure Inspection","Real-World Vision System Reliability","Artificial Intelligence and Robotics","Data Science"],"dc:title":["LLM-as-Judge for Reliable Crack Segmentation in Edge-Based Structural Inspection"],"thesis:degree_level":["Masters Thesis"],"thesis:degree_name":["Master of Science (MS)"]},"updated_at":"2026-07-24T05:45:55Z"}