A Statistical Perspective on Knowledge Distillation: Foundations, Classical Methods, and Large Language Model Extensions
Knowledge distillation (KD) has emerged as a vital paradigm for transferring the capabilities of high-capacity models to efficient student counterparts, addressing critical challenges in computational cost, deployment constraints, and privacy-sensitive settings. Although KD is widely used in practice, it is often viewed primarily as an engineering technique, with a unified statistical perspective remaining less developed. This review bridges that gap by presenting a unified Bayesian formulation that formulates teacher predictions as prior information. This provides a principled interpretation of how teacher information is incorporated into student learning and establishes a rigorous connection to uncertainty quantification. We demonstrate how this foundational lens reconciles classical distillation with modern extensions in generative and foundation-model systems, showing that contemporary developments remain rooted in these same statistical principles. By synthesizing theory with emerging methodologies and diverse applications, this review provides a conceptual road map and identifies critical open problems for the future of the field.
Authors
- Huimin Cheng (ORCID: https://orcid.org/0000-0001-5150-0011)
- Wenxuan Zhong (ORCID: https://orcid.org/0000-0001-9006-622X)
- Luyang Fang (ORCID: https://orcid.org/0009-0003-2465-6864)
- Haoran Lu
- Jiazhang Cai (ORCID: https://orcid.org/0009-0002-0267-9726)
- Tao Wang
- Ping Ma
Institutions
- Boston University (US)
- University of Georgia (US)
- University of Georgia Press (US)
Publication Details
- Journal
- Annual Review of Statistics and Its Application
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1146/annurev-statistics-043025-102547
- Primary Topic
- Intelligent Tutoring Systems and Adaptive Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00