dorsal/arxiv
View SchemaFollowing the Teacher's Footsteps: Scheduled Checkpoint Distillation for Domain-Specific LLMs
| Authors | Cheng Feng, Chaoliang Zhong, Jun Sun, Yusuke Oishi |
|---|---|
| Categories | |
| ArXiv ID | 2601.10114vv1 |
| URL | https://arxiv.org/abs/2601.10114 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Large language models (LLMs) are challenging to deploy for domain-specific tasks due to their massive scale. While distilling a fine-tuned LLM into a smaller student model is a promising alternative, the capacity gap between teacher and student often leads to suboptimal performance. This raises a key question: when and how can a student model match or even surpass its teacher on domain-specific tasks? In this work, we propose a novel theoretical insight: a student can outperform its teacher if its advantage on a Student-Favored Subdomain (SFS) outweighs its deficit on the Teacher-Favored Subdomain (TFS). Guided by this insight, we propose Scheduled Checkpoint Distillation (SCD), which reduces the TFS deficit by emulating the teacher's convergence process during supervised fine-tuning (SFT) on the domain task, and a sample-wise Adaptive Weighting (AW) mechanism to preserve student strengths on SFS. Experiments across diverse domain tasks--including QA, NER, and text classification in multiple languages--show that our method consistently outperforms existing distillation approaches, allowing the student model to match or even exceed the performance of its fine-tuned teacher.
{
"annotation_id": "053ff0c9-3d24-43c4-8547-6acdd87212d2",
"date_created": "2026-02-17T05:53:23.557000Z",
"date_modified": "2026-02-17T05:53:23.557000Z",
"file_hash": "ed72e31a64457e3eb5004f711987ec8e5f9c2578a447a9dcab2f8da9322a33e4",
"private": false,
"record": {
"abstract": "Large language models (LLMs) are challenging to deploy for domain-specific tasks due to their massive scale. While distilling a fine-tuned LLM into a smaller student model is a promising alternative, the capacity gap between teacher and student often leads to suboptimal performance. This raises a key question: when and how can a student model match or even surpass its teacher on domain-specific tasks? In this work, we propose a novel theoretical insight: a student can outperform its teacher if its advantage on a Student-Favored Subdomain (SFS) outweighs its deficit on the Teacher-Favored Subdomain (TFS). Guided by this insight, we propose Scheduled Checkpoint Distillation (SCD), which reduces the TFS deficit by emulating the teacher\u0027s convergence process during supervised fine-tuning (SFT) on the domain task, and a sample-wise Adaptive Weighting (AW) mechanism to preserve student strengths on SFS. Experiments across diverse domain tasks--including QA, NER, and text classification in multiple languages--show that our method consistently outperforms existing distillation approaches, allowing the student model to match or even exceed the performance of its fine-tuned teacher.",
"arxiv_id": "2601.10114",
"authors": [
"Cheng Feng",
"Chaoliang Zhong",
"Jun Sun",
"Yusuke Oishi"
],
"categories": [
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Following the Teacher\u0027s Footsteps: Scheduled Checkpoint Distillation for Domain-Specific LLMs",
"url": "https://arxiv.org/abs/2601.10114",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "981279e9-c16f-4574-b2ec-86f697a02943",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}