{"id":{"repo_id":"gatech","oai_identifier":"oai:repository.gatech.edu:1853/81505"},"canonical_url":"https://search.dev.ndltd.org/etd/gatech/oai:repository.gatech.edu:1853/81505","repository":{"repo_id":"gatech","name":"Georgia Tech","base_url":"https://repository.gatech.edu/server/oai/request"},"display":{"title":"Domain Adaptation of LLMs for Materials Science: Dataset Curation, Fine-Tuning, and Evaluation Benchmark","abstract":"This thesis addresses the gap between the capabilities of general-purpose large language models (LLMs) and the specialized needs of materials science. While LLMs have shown transformative potential in many fields, their application in materials science remains limited due to the lack of domain-specific natural language datasets and evaluation benchmarks. To overcome this challenge, the thesis introduces a curated instruction-tuning dataset composed of diverse question-answer (QA) pairs drawn from various materials science sources such as textbooks, property databases, and expert forums. This dataset was used to fine-tune the LLaMA-3-8B-Instruct model, to create a materials science domain expert. However, the experiments revealed both fine-tuning datasets and evaluation limitations. These findings led to a refined, higher-quality second pipeline focused on building an evaluation benchmark rather than a full training corpus. The final benchmark includes 2,100 QA pairs across a wide range of materials science topics, with a subset validated by human experts. The thesis adopts an LLM-as-judge framework to evaluate models on this benchmark, where GPT-4o is used to compare model answers against gold responses using a rubric. We are conducting a human expert annotation to show that GPT-4o, as a judge, is reliable. Ultimately, this work contributes a complete pipeline—from dataset construction to evaluation—for developing and benchmarking domain-specific LLMs in materials science. It lays the groundwork for future efforts toward using LLMs in scientific research.","abstract_html":"This thesis addresses the gap between the capabilities of general-purpose large language models (LLMs) and the specialized needs of materials science. While LLMs have shown transformative potential in many fields, their application in materials science remains limited due to the lack of domain-specific natural language datasets and evaluation benchmarks. To overcome this challenge, the thesis introduces a curated instruction-tuning dataset composed of diverse question-answer (QA) pairs drawn from various materials science sources such as textbooks, property databases, and expert forums. This dataset was used to fine-tune the LLaMA-3-8B-Instruct model, to create a materials science domain expert. However, the experiments revealed both fine-tuning datasets and evaluation limitations. These findings led to a refined, higher-quality second pipeline focused on building an evaluation benchmark rather than a full training corpus. The final benchmark includes 2,100 QA pairs across a wide range of materials science topics, with a subset validated by human experts. The thesis adopts an LLM-as-judge framework to evaluate models on this benchmark, where GPT-4o is used to compare model answers against gold responses using a rubric. We are conducting a human expert annotation to show that GPT-4o, as a judge, is reliable. Ultimately, this work contributes a complete pipeline—from dataset construction to evaluation—for developing and benchmarking domain-specific LLMs in materials science. It lays the groundwork for future efforts toward using LLMs in scientific research.","abstract_has_math":false,"creators":["Shen, Shiyao"],"institution":"Georgia Institute of Technology","degree_name":null,"degree_level":"Masters","degree_discipline":null,"degree_department":"Computer Science","school":null,"contributors":[],"advisors":["Zhang, Chao"],"committee_chairs":[],"committee_members":["Fung, Victor","Ramprasad, Rampi"],"year":2025,"date_issued":"2025-04-23","date_published":"2025-04-23","updated_at":"2026-07-27T19:49:22Z","subjects":["Benchmark Curation","LLM as a judge"],"languages":["en_US"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1853/81505","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Zhang, Chao"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Fung, Victor","Ramprasad, Rampi"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science"]},{"key":"dc:creator","label":"Author","values":["Shen, Shiyao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-05-21T20:38:57Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-05-21T20:38:57Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-04-23"]},{"key":"dc:publisher","label":"Institution","values":["Georgia Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Text"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Benchmark Curation","LLM as a judge"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1853/81505"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This thesis addresses the gap between the capabilities of general-purpose large language models (LLMs) and the specialized needs of materials science. While LLMs have shown transformative potential in many fields, their application in materials science remains limited due to the lack of domain-specific natural language datasets and evaluation benchmarks. To overcome this challenge, the thesis introduces a curated instruction-tuning dataset composed of diverse question-answer (QA) pairs drawn from various materials science sources such as textbooks, property databases, and expert forums. This dataset was used to fine-tune the LLaMA-3-8B-Instruct model, to create a materials science domain expert. However, the experiments revealed both fine-tuning datasets and evaluation limitations. These findings led to a refined, higher-quality second pipeline focused on building an evaluation benchmark rather than a full training corpus. The final benchmark includes 2,100 QA pairs across a wide range of materials science topics, with a subset validated by human experts. The thesis adopts an LLM-as-judge framework to evaluate models on this benchmark, where GPT-4o is used to compare model answers against gold responses using a rubric. We are conducting a human expert annotation to show that GPT-4o, as a judge, is reliable. Ultimately, this work contributes a complete pipeline—from dataset construction to evaluation—for developing and benchmarking domain-specific LLMs in materials science. It lays the groundwork for future efforts toward using LLMs in scientific research."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["M.S."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Domain Adaptation of LLMs for Materials Science: Dataset Curation, Fine-Tuning, and Evaluation Benchmark"]}]}],"canonical_facts":{"dc:contributor.advisor":["Zhang, Chao"],"dc:contributor.committeemember":["Fung, Victor","Ramprasad, Rampi"],"dc:contributor.department":["Computer Science"],"dc:creator":["Shen, Shiyao"],"dc:date.accessioned":["2026-05-21T20:38:57Z"],"dc:date.available":["2026-05-21T20:38:57Z"],"dc:date.issued":["2025-04-23"],"dc:description.abstract":["This thesis addresses the gap between the capabilities of general-purpose large language models (LLMs) and the specialized needs of materials science. While LLMs have shown transformative potential in many fields, their application in materials science remains limited due to the lack of domain-specific natural language datasets and evaluation benchmarks. To overcome this challenge, the thesis introduces a curated instruction-tuning dataset composed of diverse question-answer (QA) pairs drawn from various materials science sources such as textbooks, property databases, and expert forums. This dataset was used to fine-tune the LLaMA-3-8B-Instruct model, to create a materials science domain expert. However, the experiments revealed both fine-tuning datasets and evaluation limitations. These findings led to a refined, higher-quality second pipeline focused on building an evaluation benchmark rather than a full training corpus. The final benchmark includes 2,100 QA pairs across a wide range of materials science topics, with a subset validated by human experts. The thesis adopts an LLM-as-judge framework to evaluate models on this benchmark, where GPT-4o is used to compare model answers against gold responses using a rubric. We are conducting a human expert annotation to show that GPT-4o, as a judge, is reliable. Ultimately, this work contributes a complete pipeline—from dataset construction to evaluation—for developing and benchmarking domain-specific LLMs in materials science. It lays the groundwork for future efforts toward using LLMs in scientific research."],"dc:description.degree":["M.S."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/1853/81505"],"dc:language.iso":["en_US"],"dc:publisher":["Georgia Institute of Technology"],"dc:subject":["Benchmark Curation","LLM as a judge"],"dc:title":["Domain Adaptation of LLMs for Materials Science: Dataset Curation, Fine-Tuning, and Evaluation Benchmark"],"dc:type":["Text"],"thesis:degree_level":["Masters"]},"updated_at":"2026-07-27T19:49:22Z"}