Back to results

University of Illinois Urbana-Champaign

Enhancing knowledge distillation in large language models via domain adaptation

Abstract

dc:description

Domain-Adaptive Pre-Training (DAPT) is widely used to improve Large Language Models on specialized domains, yet its interaction with knowledge distillation (KD) remains poorly understood. In particular, intermediate DAPT checkpoints are rarely analyzed, and the evolution of teacher uncertainty across such checkpoints has not been systematically studied. This thesis develops a unified framework to examine how DAPT reshapes teacher confidence and how these shifts influence KD performance, downstream performance and calibration. Using LLaMA-2-7B teachers adapted for 2,000, 5,000, 7,500, and 10,000 DAPT steps, together with Sheared-LLaMA-1.3B students distilled under four KD variants, we evaluate two biomedical QA benchmarks: PubMedQA and BioASQ. We analyze teacher entropy, entropy–performance correlations, and student Expected Calibration Error (ECE) across checkpoints. Our findings reveal three key insights: (1) teacher entropy shifts moderately with deeper DAPT but redistributes most strongly over semantically informative tokens; (2) moderate entropy reduction yields the strongest KD gains for abstractive reasoning tasks such as PubMedQA, whereas extractive QA tasks benefit more from heavily domain-adapted teachers whose predictions are sharper and more concentrated, and (3) student calibration closely tracks teacher entropy, with sharper teachers generally producing better-calibrated models, though excessively low entropy can introduce calibration trade-offs.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Bioinformatics
Grantor
University of Illinois Urbana-Champaign
Year dc:date
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zhang, Xitong (Jacqueline)
Contributors dc:contributor
  • He, Jingrui
  • Ma, Jiaqi

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • Copyright 2025 Xitong (Jacqueline) Zhang
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/132689
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/132689

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Zhang, Xitong (Jacqueline). Enhancing knowledge distillation in large language models via domain adaptation. Thesis thesis, University of Illinois Urbana-Champaign, 2025. https://hdl.handle.net/2142/132689