dorsal/arxiv
View SchemaSkill-Aware Data Selection and Fine-Tuning for Data-Efficient Reasoning Distillation
| Authors | Lechen Zhang, Yunxiang Zhang, Wei Hu, Lu Wang |
|---|---|
| Categories | |
| ArXiv ID | 2601.10109vv1 |
| URL | https://arxiv.org/abs/2601.10109 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Large reasoning models such as DeepSeek-R1 and their distilled variants achieve strong performance on complex reasoning tasks. Yet, distilling these models often demands large-scale data for supervised fine-tuning (SFT), motivating the pursuit of data-efficient training methods. To address this, we propose a skill-centric distillation framework that efficiently transfers reasoning ability to weaker models with two components: (1) Skill-based data selection, which prioritizes examples targeting the student model's weaker skills, and (2) Skill-aware fine-tuning, which encourages explicit skill decomposition during problem solving. With only 1,000 training examples selected from a 100K teacher-generated corpus, our method surpasses random SFT baselines by +1.6% on Qwen3-4B and +1.4% on Qwen3-8B across five mathematical reasoning benchmarks. Further analysis confirms that these gains concentrate on skills emphasized during training, highlighting the effectiveness of skill-centric training for efficient reasoning distillation.
{
"annotation_id": "687402b8-1ed6-4535-acee-c1c42c080029",
"date_created": "2026-02-17T05:53:23.803000Z",
"date_modified": "2026-02-17T05:53:23.803000Z",
"file_hash": "9bab84e5884ed1824557ddeef65b53ff8a29b819ee186168b5543f58aa1c7b10",
"private": false,
"record": {
"abstract": "Large reasoning models such as DeepSeek-R1 and their distilled variants achieve strong performance on complex reasoning tasks. Yet, distilling these models often demands large-scale data for supervised fine-tuning (SFT), motivating the pursuit of data-efficient training methods. To address this, we propose a skill-centric distillation framework that efficiently transfers reasoning ability to weaker models with two components: (1) Skill-based data selection, which prioritizes examples targeting the student model\u0027s weaker skills, and (2) Skill-aware fine-tuning, which encourages explicit skill decomposition during problem solving. With only 1,000 training examples selected from a 100K teacher-generated corpus, our method surpasses random SFT baselines by +1.6% on Qwen3-4B and +1.4% on Qwen3-8B across five mathematical reasoning benchmarks. Further analysis confirms that these gains concentrate on skills emphasized during training, highlighting the effectiveness of skill-centric training for efficient reasoning distillation.",
"arxiv_id": "2601.10109",
"authors": [
"Lechen Zhang",
"Yunxiang Zhang",
"Wei Hu",
"Lu Wang"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Skill-Aware Data Selection and Fine-Tuning for Data-Efficient Reasoning Distillation",
"url": "https://arxiv.org/abs/2601.10109",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "a52f9cbb-6059-4eb8-b187-145095bbd35e",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}