Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 8 of 8 for “"LLM as a judge"”.
-
Domain Adaptation of LLMs for Materials Science: Dataset Curation, Fine-Tuning, and Evaluation Benchmark
… of general-purpose large language models (LLMs) and the specialized needs of materials science. While LLMs have shown transformative potential in many fields, their application in materials science remains limited due to the lack of domain-specific natural language datasets and evaluation …
-
Evaluating Differences in GPT-4 Treatment by Gender in Healthcare Applications
LLMs already permeate medical settings, supporting patient messaging, medical scribing, and chatbots. While prior work has examined bias in medical LLMs, few studies focus on realistic use cases or analyze the source of the bias. To assess whether medical LLMs exhibit differential performance by …
-
Toward Trustworthy Power Line Perception from Aerial Imagery: Segmentation and Reliability Monitoring
… the safety and reliability of modern infrastructure, yet traditional methods remain costly, slow, and potentially hazardous. With the increasing use of unmanned aerial vehicles (UAVs), automated vision systems have emerged as a promising alternative for large-scale aerial inspection. A key …
-
From Form to Meaning: Interlingual Sense-Alignment of Offensive Language with LLMs
… μια μεθοδολογία «γλωσσικού μοντέλου ως κριτή» (LLM-as-a-judge) για την ευθυγράμμιση προσβλητικών σημασιών μεταξύ της Αραβικής, της Βουλγαρικής, της Νέας Ελληνικής, της Γαλλικής και της Ιταλικής, μέσω τριάδων λήμματος–ορισμού–παραδείγματος στις πρωτότυπες γλώσσες. Σε περίπου 2,87 εκατομμύρια …
-
Agentic Framework for Domain-specific RAG Evaluation
Large Language Models (LLMs) have revolutionized generation and understanding of textual information and have wide-ranging applications. While the capabilities of LLMs are immediately impressive to any user, extended use quickly reveals problems of inaccurate text generation associated with these …
-
DART-RAG: Dynamic Agentic Replanning for Trustworthy Retrieval-Augmented Generation in Scientific QA
… για τη θεμελίωση των Μεγάλων Γλωσσικών Μοντέλων (LLMs) σε εξωτερική γνώση, αλλά τα τρέχοντα πλαίσια συστημάτων με ένα ή πολλαπλούς πράκτορες δυσκολεύονται με επιστημονικά ερωτήματα που απαιτούν συλλογιστική πολλαπλών βημάτων και λεπτομερή απόδοση πηγών. Τέτοια συστήματα συχνά βασίζονται σε …
-
Towards Gender-Inclusive Machine Translation
… systematically encode and perpetuate gender bias. When translating into grammatical gender languages such as Italian, Spanish, and German, these systems default to masculine forms for gender-ambiguous referents, reinforce stereotypical associations between gender and social roles, and fail to …
-
Adversarial Attacks on Natural Language and Speech Processing Models
… designed to capture patterns in data. The increasing availability of computational resources in recent years has facilitated the training and deployment of deep learning models across a wide range of applications in computer vision, natural language processing, and speech processing. Despite …