KL Divergence VS MSE for Knowledge Distillation
Unknown
MSE loss outperforms KL divergence in knowledge distillation by enabling direct logit matching, with sequential distillation and small-tau KL improving noise robustness.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
MSE loss outperforms KL divergence in knowledge distillation by enabling direct logit matching, with sequential distillation and small-tau KL improving noise robustness.
Unknown
An open-source family of heterogeneous reasoning models (Nano, Super, Ultra) with dynamic reasoning toggle, trained via NAS, distillation, and RL.
Unknown
Gemma 2 improves open language models by interleaving local-global attention and using knowledge distillation, achieving competitive performance with much larger models.
Unknown
Relational knowledge distillation extends traditional distillation by transferring structural relationships between data points from teacher to student.
Xiaohan Xu, Ming Li, Chongyang Tao, et al.
This survey systematically reviews knowledge distillation techniques for transferring capabilities from large proprietary LLMs to smaller models.
Unknown
This paper investigates the factors influencing the efficacy of knowledge distillation, a technique for compressing large models into smaller ones.
Unknown
This paper integrates knowledge distillation with self-supervised contrastive learning to improve student model performance without requiring labeled data.
Unknown
Proposes similarity-preserving knowledge distillation that uses pairwise similarity matrices as supervisory signals to train student networks.
Amir M. Mansourian, Rozhan Ahmadi, Masoud Ghafouri, et al.
A comprehensive survey of knowledge distillation methods, covering recent techniques and categorizing approaches in the field.
Unknown
Proposes logit standardization as a pre-processing step to improve vanilla knowledge distillation performance.
Unknown
This paper proposes Teacher Assistant Knowledge Distillation (TAKD) to improve knowledge transfer by introducing an intermediate assistant model between teacher and student.
Unknown
A comprehensive survey of knowledge distillation advancements, covering foundational techniques, relation-based methods, and novel approaches.