{"id":{"repo_id":"stellenbosch","oai_identifier":"oai:scholar.sun.ac.za:10019.1/136200"},"canonical_url":"https://search.dev.ndltd.org/etd/stellenbosch/oai:scholar.sun.ac.za:10019.1/136200","repository":{"repo_id":"stellenbosch","name":"Stellenbosch University","base_url":"https://scholar.sun.ac.za/server/oai/request"},"display":{"title":"Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing","abstract":"This study investigates the integration of Vision Language Model (VLM)s, Retrieval-Augmented Generation (RAG), and structured prompt-engineering strategies to enhance surface defect classification and Root Cause Analysis (RCA) in manufacturing. The research examines whether generative and multimodal Artificial Intelligence (AI) systems can deliver accurate, explainable, and cost efficient quality control for manufactured components. The first study assessed GPT-4o’s capability for surface defect detection and classification using two datasets: the MVTec AD screw subset and a custom #3DBenchy dataset developed by the author. Results showed that image-only prompting consistently outperformed multimodal (image + text) inputs, achieving up to 94.4% defect identification and 82.6% classification accuracy on the #3DBenchy dataset, and 90.5% defect identification with 76% classification accuracy on the MVTec AD screw subset. Small-sample fine-tuning (5–30 images per class) further improved classification accuracy across both datasets. The second study extended the framework to RCA using RAG and structured prompting, incorporating domain-specific manufacturing knowledge to improve diagnostic reasoning. Three GPT-5 model variants (gpt-5-nano, gpt-5-mini, and gpt-5) were evaluated across reasoning effort levels and prompting strategies (ReAct and ReAct + 5 Whys). Larger models achieved higher diagnostic accuracy, whereas smaller variants—particularly gpt-5-nano at low reasoning effort—offered optimal cost-efficiency. Inconsistent RAG performance underscored the need for curated retrieval contexts. The final study introduced a gated ReAct-style RCA pipeline applying sequential reasoning across the Man, Method, and Machine categories of the 6M framework. Incorporating multimodal data such as G-code parameters, operator logs, temperature time-series, and bed-mesh images enabled near-perfect parameter-level diagnosis. Token usage scaled predictably with reasoning depth, validating the trade-off between interpretability and computational efficiency. Overall, the findings confirm that VLMs, when combined with structured prompting and domain-specific retrieval, can serve as reliable tools for automated defect diagnosis and causal reasoning—laying the groundwork for scalable, explainable, and data-driven quality control systems in advanced manufacturing.","abstract_html":"This study investigates the integration of Vision Language Model (VLM)s, Retrieval-Augmented Generation (RAG), and structured prompt-engineering strategies to enhance surface defect classification and Root Cause Analysis (RCA) in manufacturing. The research examines whether generative and multimodal Artificial Intelligence (AI) systems can deliver accurate, explainable, and cost efficient quality control for manufactured components. The first study assessed GPT-4o’s capability for surface defect detection and classification using two datasets: the MVTec AD screw subset and a custom #3DBenchy dataset developed by the author. Results showed that image-only prompting consistently outperformed multimodal (image + text) inputs, achieving up to 94.4% defect identification and 82.6% classification accuracy on the #3DBenchy dataset, and 90.5% defect identification with 76% classification accuracy on the MVTec AD screw subset. Small-sample fine-tuning (5–30 images per class) further improved classification accuracy across both datasets. The second study extended the framework to RCA using RAG and structured prompting, incorporating domain-specific manufacturing knowledge to improve diagnostic reasoning. Three GPT-5 model variants (gpt-5-nano, gpt-5-mini, and gpt-5) were evaluated across reasoning effort levels and prompting strategies (ReAct and ReAct + 5 Whys). Larger models achieved higher diagnostic accuracy, whereas smaller variants—particularly gpt-5-nano at low reasoning effort—offered optimal cost-efficiency. Inconsistent RAG performance underscored the need for curated retrieval contexts. The final study introduced a gated ReAct-style RCA pipeline applying sequential reasoning across the Man, Method, and Machine categories of the 6M framework. Incorporating multimodal data such as G-code parameters, operator logs, temperature time-series, and bed-mesh images enabled near-perfect parameter-level diagnosis. Token usage scaled predictably with reasoning depth, validating the trade-off between interpretability and computational efficiency. Overall, the findings confirm that VLMs, when combined with structured prompting and domain-specific retrieval, can serve as reliable tools for automated defect diagnosis and causal reasoning—laying the groundwork for scalable, explainable, and data-driven quality control systems in advanced manufacturing.","abstract_has_math":false,"creators":["Mocke, Johannes Jacobus"],"institution":"Stellenbosch : Stellenbosch University","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Louw, Louis","Lucke, D."],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-03","date_published":"2026-03","updated_at":"2026-07-24T04:40:12Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://scholar.sun.ac.za/handle/10019.1/136200","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Louw, Louis","Lucke, D."]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Stellenbosch University. Faculty of Engineering. Dept. of Industrial Engineering."]},{"key":"dc:creator","label":"Author","values":["Mocke, Johannes Jacobus"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-04-24T13:45:36Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-04-24T13:45:36Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-03"]},{"key":"dc:publisher","label":"Institution","values":["Stellenbosch : Stellenbosch University"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholar.sun.ac.za/handle/10019.1/136200"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (MEng)--Stellenbosch University, 2026.","Mocke, J. J. 2026. Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing. Unpublished masters thesis. Stellenbosch: Stellenbosch University [online]. Available: https://scholar.sun.ac.za/items/a1e4dce0-c0aa-4c69-b21e-3e30890ee295"]},{"key":"dc:description.abstract","label":"Abstract","values":["This study investigates the integration of Vision Language Model (VLM)s, Retrieval-Augmented Generation (RAG), and structured prompt-engineering strategies to enhance surface defect classification and Root Cause Analysis (RCA) in manufacturing. The research examines whether generative and multimodal Artificial Intelligence (AI) systems can deliver accurate, explainable, and cost efficient quality control for manufactured components. The first study assessed GPT-4o’s capability for surface defect detection and classification using two datasets: the MVTec AD screw subset and a custom #3DBenchy dataset developed by the author. Results showed that image-only prompting consistently outperformed multimodal (image + text) inputs, achieving up to 94.4% defect identification and 82.6% classification accuracy on the #3DBenchy dataset, and 90.5% defect identification with 76% classification accuracy on the MVTec AD screw subset. Small-sample fine-tuning (5–30 images per class) further improved classification accuracy across both datasets. The second study extended the framework to RCA using RAG and structured prompting, incorporating domain-specific manufacturing knowledge to improve diagnostic reasoning. Three GPT-5 model variants (gpt-5-nano, gpt-5-mini, and gpt-5) were evaluated across reasoning effort levels and prompting strategies (ReAct and ReAct + 5 Whys). Larger models achieved higher diagnostic accuracy, whereas smaller variants—particularly gpt-5-nano at low reasoning effort—offered optimal cost-efficiency. Inconsistent RAG performance underscored the need for curated retrieval contexts. The final study introduced a gated ReAct-style RCA pipeline applying sequential reasoning across the Man, Method, and Machine categories of the 6M framework. Incorporating multimodal data such as G-code parameters, operator logs, temperature time-series, and bed-mesh images enabled near-perfect parameter-level diagnosis. Token usage scaled predictably with reasoning depth, validating the trade-off between interpretability and computational efficiency. Overall, the findings confirm that VLMs, when combined with structured prompting and domain-specific retrieval, can serve as reliable tools for automated defect diagnosis and causal reasoning—laying the groundwork for scalable, explainable, and data-driven quality control systems in advanced manufacturing."]},{"key":"dc:title","label":"Title","values":["Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing"]}]}],"canonical_facts":{"dc:contributor.advisor":["Louw, Louis","Lucke, D."],"dc:contributor.other":["Stellenbosch University. Faculty of Engineering. Dept. of Industrial Engineering."],"dc:creator":["Mocke, Johannes Jacobus"],"dc:date.accessioned":["2026-04-24T13:45:36Z"],"dc:date.available":["2026-04-24T13:45:36Z"],"dc:date.issued":["2026-03"],"dc:description":["Thesis (MEng)--Stellenbosch University, 2026.","Mocke, J. J. 2026. Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing. Unpublished masters thesis. Stellenbosch: Stellenbosch University [online]. Available: https://scholar.sun.ac.za/items/a1e4dce0-c0aa-4c69-b21e-3e30890ee295"],"dc:description.abstract":["This study investigates the integration of Vision Language Model (VLM)s, Retrieval-Augmented Generation (RAG), and structured prompt-engineering strategies to enhance surface defect classification and Root Cause Analysis (RCA) in manufacturing. The research examines whether generative and multimodal Artificial Intelligence (AI) systems can deliver accurate, explainable, and cost efficient quality control for manufactured components. The first study assessed GPT-4o’s capability for surface defect detection and classification using two datasets: the MVTec AD screw subset and a custom #3DBenchy dataset developed by the author. Results showed that image-only prompting consistently outperformed multimodal (image + text) inputs, achieving up to 94.4% defect identification and 82.6% classification accuracy on the #3DBenchy dataset, and 90.5% defect identification with 76% classification accuracy on the MVTec AD screw subset. Small-sample fine-tuning (5–30 images per class) further improved classification accuracy across both datasets. The second study extended the framework to RCA using RAG and structured prompting, incorporating domain-specific manufacturing knowledge to improve diagnostic reasoning. Three GPT-5 model variants (gpt-5-nano, gpt-5-mini, and gpt-5) were evaluated across reasoning effort levels and prompting strategies (ReAct and ReAct + 5 Whys). Larger models achieved higher diagnostic accuracy, whereas smaller variants—particularly gpt-5-nano at low reasoning effort—offered optimal cost-efficiency. Inconsistent RAG performance underscored the need for curated retrieval contexts. The final study introduced a gated ReAct-style RCA pipeline applying sequential reasoning across the Man, Method, and Machine categories of the 6M framework. Incorporating multimodal data such as G-code parameters, operator logs, temperature time-series, and bed-mesh images enabled near-perfect parameter-level diagnosis. Token usage scaled predictably with reasoning depth, validating the trade-off between interpretability and computational efficiency. Overall, the findings confirm that VLMs, when combined with structured prompting and domain-specific retrieval, can serve as reliable tools for automated defect diagnosis and causal reasoning—laying the groundwork for scalable, explainable, and data-driven quality control systems in advanced manufacturing."],"dc:identifier.uri":["https://scholar.sun.ac.za/handle/10019.1/136200"],"dc:language.iso":["en"],"dc:publisher":["Stellenbosch : Stellenbosch University"],"dc:title":["Leveraging Retrieval-Augmented Generation, Prompt Engineering, and Vision Language Models for Surface Defect Classification and Root Cause Analysis in Manufacturing"],"dc:type":["Thesis"]},"updated_at":"2026-07-24T04:40:12Z"}