Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 3 of 3 for “"Evaluation of LLMs"”.

  1. An Empirical Evaluation of LLMs for the Assessment of Subjective Qualities

    Large Language Models (LLMs) have achieved remarkable success in natural language processing tasks and are increasingly being used for language generation. Significant advancements in this field have unlocked capabilities that enable their adoption in sophisticated roles, including acting as …

    mit Repository record for An Empirical Evaluation of LLMs for the Assessment of Subjective Qualities (opens in a new tab)

  2. SWE-Bench+: Enhanced Coding Benchmark for LLMs

    Large Language Models (LLMs) in Software Engineering (SE) can offer valuable assistance for coding tasks. To facilitate a rigorous evaluation of LLMs in practical coding contexts, Carlos et al. introduced the SWE-bench dataset, which comprises 2,294 real-world GitHub issues. Several impressive …

    york Repository record for SWE-Bench+: Enhanced Coding Benchmark for LLMs (opens in a new tab)

  3. Statistical Methods for Performance Evaluation of Machine Learning and Artificial Intelligence Models

    … to improve the reliability and effectiveness of artificial intelligence (AI) and machine learning (ML) in practical data analysis tasks. As AI tech- nologies become increasingly capable and widely adopted in real-world applications—from detecting defect severity in solar panels to …

    vt Repository record for Statistical Methods for Performance Evaluation of Machine Learning and Artificial Intelligence Models (opens in a new tab)