Back to results

Virginia Tech

Statistical Methods for Performance Evaluation of Machine Learning and Artificial Intelligence Models

Abstract

dc:description.abstractgeneral

This dissertation explores strategies to improve the reliability and effectiveness of artificial intelligence (AI) and machine learning (ML) in practical data analysis tasks. As AI tech- nologies become increasingly capable and widely adopted in real-world applications—from detecting defect severity in solar panels to automatically generating analytical code—they often encounter challenges in complex scenarios, such as imbalanced datasets with rare out- comes. This research focuses on developing and refining tools and methodologies that enhance decision-making and enable more accurate evaluation of AI models, particularly under such challenging conditions. The first project focuses on using EL images to detect defects in solar panels. It compares several machine learning and deep learning models to see how well they identify severity of defectiveness. The results provide useful guidance for choosing the right prediction methods and evaluation tools in solar panel research. Building upon insights from the first project, the second project tackles a significant chal- lenge: although machine learning and deep learning models generally perform well, they struggle to accurately detect less frequent defect classes, such as "mildly defective" and "mod- erately defective" solar panels. To overcome this issue, we introduce customized loss functions alongwithmini-batchstratifiedsampling, aimingtoimprovepredictionaccuracyfortheserare defect classes. The proposed methods are evaluated using both a simulated dataset derived from Fashion MNIST—which mirrors the class proportion of the EL image dataset—and real EL image datasets, utilizing VGG-19 and ResNet-50 architectures. To ensure reliability, the analysis is repeated 50 times on the simulated dataset and 30 times on the EL image dataset. The third project examines how well AI tools—specifically LLMs like ChatGPT and Llama—can generate SAS code for automated statistical analysis. Although the code often appears correct, these tools sometimes fall short in handling more complex tasks or producing code that runs properly. This project evaluates the quality of the AI-generated code based on human expert assessment, focusing on code quality, correctness, executability and output. To support this evaluation, the last part of this dissertation introduces a new open-source dataset called StatLLM. This dataset provides examples of statistical tasks, code written by AI, and expert ratings of the results. StatLLM helps researchers and developers understand where AI tools perform well, wheretheyneedimprovementwhenitcomestowritingstatistical code. In summary, this dissertation advances our ability to evaluate and improve AI tools in data science. It helps ensure these technologies are not only powerful but also trustworthy and practical in solving real-world problems.

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy
Level thesis:degree_level
doctoral
Discipline thesis:degree_discipline
Statistics
Department dc:contributor.department
Statistics
Grantor dc:publisher
Virginia Tech
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Song, Xinyi
Chair dc:contributor.committeechair
  • Hong, Yili
Committee members dc:contributor.committeemember
  • Freeman, Laura June
  • Deng, Xinwei
  • Xing, Xin

Subjects

dc:subject × 6

Rights

dc:rights
Statement dc:rights
  • In Copyright
Language dc:language.iso
en

Identifiers

dc:identifier.*
Dc Identifier Other
vt_gsexam:44183
OAI identifier oai:identifier
oai:vtechworks.lib.vt.edu:10919/135034

Chain of custody

source
Harvested from
Virginia Tech
Base URL
vtechworks.lib.vt.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Song, Xinyi. Statistical Methods for Performance Evaluation of Machine Learning and Artificial Intelligence Models. doctoral thesis, Virginia Tech, 2025. https://hdl.handle.net/10919/135034