dorsal/arxiv
View SchemaDistributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models
| Authors | Shaoning Sun, Mingzhu Cai, Huang He, Bingjin Chen, Siqi Bao, Yujiu Yang, Hua Wu, Haifeng Wang |
|---|---|
| Categories | |
| ArXiv ID | 2601.06911vv1 |
| URL | https://arxiv.org/abs/2601.06911 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Language model families exhibit striking disparity in their capacity to benefit from reinforcement learning: under identical training, models like Qwen achieve substantial gains, while others like Llama yield limited improvements. Complementing data-centric approaches, we reveal that this disparity reflects a hidden structural property: \textbf{distributional clarity} in probability space. Through a three-stage analysis-from phenomenon to mechanism to interpretation-we uncover that RL-friendly models exhibit intra-class compactness and inter-class separation in their probability assignments to correct vs. incorrect responses. We quantify this clarity using the \textbf{Silhouette Coefficient} ($S$) and demonstrate that (1) high $S$ correlates strongly with RL performance; (2) low $S$ is associated with severe logic errors and reasoning instability. To confirm this property, we introduce a Silhouette-Aware Reweighting strategy that prioritizes low-$S$ samples during training. Experiments across six mathematical benchmarks show consistent improvements across all model families, with gains up to 5.9 points on AIME24. Our work establishes distributional clarity as a fundamental, trainable property underlying RL-Friendliness.
{
"annotation_id": "06b46980-0819-4469-97c0-3351fc54fea2",
"date_created": "2026-02-17T05:53:08.501000Z",
"date_modified": "2026-02-17T05:53:08.501000Z",
"file_hash": "5faf590fa9d00dab4976d07800aa8a44a29183b3cafb81ccc586c7a2deda1fb8",
"private": false,
"record": {
"abstract": "Language model families exhibit striking disparity in their capacity to benefit from reinforcement learning: under identical training, models like Qwen achieve substantial gains, while others like Llama yield limited improvements. Complementing data-centric approaches, we reveal that this disparity reflects a hidden structural property: \\textbf{distributional clarity} in probability space. Through a three-stage analysis-from phenomenon to mechanism to interpretation-we uncover that RL-friendly models exhibit intra-class compactness and inter-class separation in their probability assignments to correct vs. incorrect responses. We quantify this clarity using the \\textbf{Silhouette Coefficient} ($S$) and demonstrate that (1) high $S$ correlates strongly with RL performance; (2) low $S$ is associated with severe logic errors and reasoning instability. To confirm this property, we introduce a Silhouette-Aware Reweighting strategy that prioritizes low-$S$ samples during training. Experiments across six mathematical benchmarks show consistent improvements across all model families, with gains up to 5.9 points on AIME24. Our work establishes distributional clarity as a fundamental, trainable property underlying RL-Friendliness.",
"arxiv_id": "2601.06911",
"authors": [
"Shaoning Sun",
"Mingzhu Cai",
"Huang He",
"Bingjin Chen",
"Siqi Bao",
"Yujiu Yang",
"Hua Wu",
"Haifeng Wang"
],
"categories": [
"cs.CL",
"cs.AI",
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models",
"url": "https://arxiv.org/abs/2601.06911",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "814069a7-fb81-4721-afaa-a53f0f03c2ab",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}