dorsal/arxiv
View SchemaClosing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
| Authors | Xin Gao, Xiaoyang Wang, Yun Zhu, Mengzhang Cai, Conghui He, Lijun Wu |
|---|---|
| Categories | |
| ArXiv ID | 2601.09733vv1 |
| URL | https://arxiv.org/abs/2601.09733 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
The construction of Supervised Fine-Tuning (SFT) datasets is a critical yet under-theorized stage in the post-training of Large Language Models (LLMs), as prevalent practices often rely on heuristic aggregation without a systematic understanding of how individual samples contribute to model performance. In this report, we propose a paradigm shift from ad-hoc curation to a closed-loop dataset engineering framework using OpenDataArena (ODA), which leverages value-anchored rankings and multi-dimensional analysis to transform value benchmarking into feedback signals guiding dataset construction. We instantiate this methodology through two new datasets: \textbf{ODA-Math-460k}, a specialized mathematics reasoning dataset that utilizes a novel two-stage difficulty-aware pipeline to achieve State-of-the-Art (SOTA) results on benchmarks such as AIME and HMMT, and \textbf{ODA-Mixture (100k \& 500k)}, a series of multi-domain instruction datasets built via an ``Anchor-and-Patch'' strategy that outperforms significantly larger open-source baselines. Our empirical results demonstrate that ODA-driven datasets significantly improve both domain-specific reasoning and general utility while achieving superior data efficiency, validating a transition toward data-centric AI where transparent evaluation serves as the primary engine for engineering high-quality training data.
{
"annotation_id": "22256fd1-3136-4428-9b9c-83ddd2d83cd0",
"date_created": "2026-02-17T05:53:23.006000Z",
"date_modified": "2026-02-17T05:53:23.006000Z",
"file_hash": "401361544d7722552b7c4c9d03398040a1b18549b246501f16454ad52e0f52e6",
"private": false,
"record": {
"abstract": "The construction of Supervised Fine-Tuning (SFT) datasets is a critical yet under-theorized stage in the post-training of Large Language Models (LLMs), as prevalent practices often rely on heuristic aggregation without a systematic understanding of how individual samples contribute to model performance. In this report, we propose a paradigm shift from ad-hoc curation to a closed-loop dataset engineering framework using OpenDataArena (ODA), which leverages value-anchored rankings and multi-dimensional analysis to transform value benchmarking into feedback signals guiding dataset construction. We instantiate this methodology through two new datasets: \\textbf{ODA-Math-460k}, a specialized mathematics reasoning dataset that utilizes a novel two-stage difficulty-aware pipeline to achieve State-of-the-Art (SOTA) results on benchmarks such as AIME and HMMT, and \\textbf{ODA-Mixture (100k \\\u0026 500k)}, a series of multi-domain instruction datasets built via an ``Anchor-and-Patch\u0027\u0027 strategy that outperforms significantly larger open-source baselines. Our empirical results demonstrate that ODA-driven datasets significantly improve both domain-specific reasoning and general utility while achieving superior data efficiency, validating a transition toward data-centric AI where transparent evaluation serves as the primary engine for engineering high-quality training data.",
"arxiv_id": "2601.09733",
"authors": [
"Xin Gao",
"Xiaoyang Wang",
"Yun Zhu",
"Mengzhang Cai",
"Conghui He",
"Lijun Wu"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets",
"url": "https://arxiv.org/abs/2601.09733",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "71617ab3-85fb-4487-a4ee-d8592df306be",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}