Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 6 of 6 for “"Evaluation Benchmark"”.

  1. Domain Adaptation of LLMs for Materials Science: Dataset Curation, Fine-Tuning, and Evaluation Benchmark

    … of domain-specific natural language datasets and evaluation benchmarks. To overcome this challenge, the thesis introduces a curated instruction-tuning dataset composed of diverse question-answer (QA) pairs drawn from various materials science sources such as textbooks, property databases, and …

    gatech Repository record for Domain Adaptation of LLMs for Materials Science: Dataset Curation, Fine-Tuning, and Evaluation Benchmark (opens in a new tab)

  2. An fMRI dataset of 1,102 natural videos for visual event understanding

    … we demonstrate its utility as a model evaluation benchmark to bridge the gap between visual neuroscience and video-based computer vision research.

    mit Repository record for An fMRI dataset of 1,102 natural videos for visual event understanding (opens in a new tab)

  3. Modeling Empathic Similarity in Personal Narratives

    … similar stories, as well as the first evaluation benchmark on this task. We operationalize empathic similarity in personal stories using insights from social psychology and narratology, and introduce EmpathicStories, a crowdsourced dataset of emotional personal experiences annotated …

    mit Repository record for Modeling Empathic Similarity in Personal Narratives (opens in a new tab)

  4. The Illinois Retrieval Benchmark : A scalable framework for characterizing retrieval-augmented generation via automated fact-checking

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms

    uiuc Repository record for The Illinois Retrieval Benchmark : A scalable framework for characterizing retrieval-augmented generation via automated fact-checking (opens in a new tab)

  5. Benchmarking Reinforcement Learning and Off Policy Evaluation for Medical Decision Making

    … patient care scenarios, no standardized benchmark has yet been developed. To address this gap, this thesis introduces Episodes of Care (EpiCare), a benchmark designed to mimic the challenges associated with applying RL to longitudinal healthcare settings. I leverage this benchmark to test …

    rockefeller Repository record for Benchmarking Reinforcement Learning and Off Policy Evaluation for Medical Decision Making (opens in a new tab)