{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132689"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132689","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Enhancing knowledge distillation in large language models via domain adaptation","abstract":"Domain-Adaptive Pre-Training (DAPT) is widely used to improve Large Language Models on specialized domains, yet its interaction with knowledge distillation (KD) remains poorly understood. In particular, intermediate DAPT checkpoints are rarely analyzed, and the evolution of teacher uncertainty across such checkpoints has not been systematically studied. This thesis develops a unified framework to examine how DAPT reshapes teacher confidence and how these shifts influence KD performance, downstream performance and calibration. Using LLaMA-2-7B teachers adapted for 2,000, 5,000, 7,500, and 10,000 DAPT steps, together with Sheared-LLaMA-1.3B students distilled under four KD variants, we evaluate two biomedical QA benchmarks: PubMedQA and BioASQ. We analyze teacher entropy, entropy–performance correlations, and student Expected Calibration Error (ECE) across checkpoints. Our findings reveal three key insights: (1) teacher entropy shifts moderately with deeper DAPT but redistributes most strongly over semantically informative tokens; (2) moderate entropy reduction yields the strongest KD gains for abstractive reasoning tasks such as PubMedQA, whereas extractive QA tasks benefit more from heavily domain-adapted teachers whose predictions are sharper and more concentrated, and (3) student calibration closely tracks teacher entropy, with sharper teachers generally producing better-calibrated models, though excessively low entropy can introduce calibration trade-offs.","abstract_html":"Domain-Adaptive Pre-Training (DAPT) is widely used to improve Large Language Models on specialized domains, yet its interaction with knowledge distillation (KD) remains poorly understood. In particular, intermediate DAPT checkpoints are rarely analyzed, and the evolution of teacher uncertainty across such checkpoints has not been systematically studied. This thesis develops a unified framework to examine how DAPT reshapes teacher confidence and how these shifts influence KD performance, downstream performance and calibration. Using LLaMA-2-7B teachers adapted for 2,000, 5,000, 7,500, and 10,000 DAPT steps, together with Sheared-LLaMA-1.3B students distilled under four KD variants, we evaluate two biomedical QA benchmarks: PubMedQA and BioASQ. We analyze teacher entropy, entropy–performance correlations, and student Expected Calibration Error (ECE) across checkpoints. Our findings reveal three key insights: (1) teacher entropy shifts moderately with deeper DAPT but redistributes most strongly over semantically informative tokens; (2) moderate entropy reduction yields the strongest KD gains for abstractive reasoning tasks such as PubMedQA, whereas extractive QA tasks benefit more from heavily domain-adapted teachers whose predictions are sharper and more concentrated, and (3) student calibration closely tracks teacher entropy, with sharper teachers generally producing better-calibrated models, though excessively low entropy can introduce calibration trade-offs.","abstract_has_math":false,"creators":["Zhang, Xitong (Jacqueline)"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Bioinformatics","degree_department":null,"school":null,"contributors":["He, Jingrui","Ma, Jiaqi"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Knowledge Distillation","Large Language Models","Deep Learning","Machine Learning","Artificial Intelligence"],"languages":["en"],"rights":["Copyright 2025 Xitong (Jacqueline) Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132689","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["He, Jingrui","Ma, Jiaqi"]},{"key":"dc:creator","label":"Author","values":["Zhang, Xitong (Jacqueline)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-12-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Bioinformatics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Knowledge Distillation","Large Language Models","Deep Learning","Machine Learning","Artificial Intelligence"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Xitong (Jacqueline) Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132689"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Domain-Adaptive Pre-Training (DAPT) is widely used to improve Large Language Models on specialized domains, yet its interaction with knowledge distillation (KD) remains poorly understood. In particular, intermediate DAPT checkpoints are rarely analyzed, and the evolution of teacher uncertainty across such checkpoints has not been systematically studied. This thesis develops a unified framework to examine how DAPT reshapes teacher confidence and how these shifts influence KD performance, downstream performance and calibration. Using LLaMA-2-7B teachers adapted for 2,000, 5,000, 7,500, and 10,000 DAPT steps, together with Sheared-LLaMA-1.3B students distilled under four KD variants, we evaluate two biomedical QA benchmarks: PubMedQA and BioASQ. We analyze teacher entropy, entropy–performance correlations, and student Expected Calibration Error (ECE) across checkpoints. Our findings reveal three key insights: (1) teacher entropy shifts moderately with deeper DAPT but redistributes most strongly over semantically informative tokens; (2) moderate entropy reduction yields the strongest KD gains for abstractive reasoning tasks such as PubMedQA, whereas extractive QA tasks benefit more from heavily domain-adapted teachers whose predictions are sharper and more concentrated, and (3) student calibration closely tracks teacher entropy, with sharper teachers generally producing better-calibrated models, though excessively low entropy can introduce calibration trade-offs.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-12-01","The student, Xitong (Jacqueline) Zhang, accepted the attached license on 2025-12-04 at 12:24.","The student, Xitong (Jacqueline) Zhang, submitted this Thesis for approval on 2025-12-04 at 12:35.","This Thesis was approved for publication on 2025-12-08 at 16:07.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23060 on 2026-02-19 at 18:46:43"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Enhancing knowledge distillation in large language models via domain adaptation"]}]}],"canonical_facts":{"dc:contributor":["He, Jingrui","Ma, Jiaqi"],"dc:creator":["Zhang, Xitong (Jacqueline)"],"dc:date":["2025-12","2025-12-08"],"dc:description":["Domain-Adaptive Pre-Training (DAPT) is widely used to improve Large Language Models on specialized domains, yet its interaction with knowledge distillation (KD) remains poorly understood. In particular, intermediate DAPT checkpoints are rarely analyzed, and the evolution of teacher uncertainty across such checkpoints has not been systematically studied. This thesis develops a unified framework to examine how DAPT reshapes teacher confidence and how these shifts influence KD performance, downstream performance and calibration. Using LLaMA-2-7B teachers adapted for 2,000, 5,000, 7,500, and 10,000 DAPT steps, together with Sheared-LLaMA-1.3B students distilled under four KD variants, we evaluate two biomedical QA benchmarks: PubMedQA and BioASQ. We analyze teacher entropy, entropy–performance correlations, and student Expected Calibration Error (ECE) across checkpoints. Our findings reveal three key insights: (1) teacher entropy shifts moderately with deeper DAPT but redistributes most strongly over semantically informative tokens; (2) moderate entropy reduction yields the strongest KD gains for abstractive reasoning tasks such as PubMedQA, whereas extractive QA tasks benefit more from heavily domain-adapted teachers whose predictions are sharper and more concentrated, and (3) student calibration closely tracks teacher entropy, with sharper teachers generally producing better-calibrated models, though excessively low entropy can introduce calibration trade-offs.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-12-01","The student, Xitong (Jacqueline) Zhang, accepted the attached license on 2025-12-04 at 12:24.","The student, Xitong (Jacqueline) Zhang, submitted this Thesis for approval on 2025-12-04 at 12:35.","This Thesis was approved for publication on 2025-12-08 at 16:07.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23060 on 2026-02-19 at 18:46:43"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132689"],"dc:language":["en"],"dc:rights":["Copyright 2025 Xitong (Jacqueline) Zhang"],"dc:subject":["Knowledge Distillation","Large Language Models","Deep Learning","Machine Learning","Artificial Intelligence"],"dc:title":["Enhancing knowledge distillation in large language models via domain adaptation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Bioinformatics"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}