dorsal/arxiv
View SchemaEvidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
| Authors | Xin Guan, Zijian Li, Shen Huang, Pengjun Xie, Jingren Zhou, Jiuxin Cao |
|---|---|
| Categories | |
| ArXiv ID | 2601.10306vv1 |
| URL | https://arxiv.org/abs/2601.10306 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
While Reinforcement Learning (RL) has advanced LLM reasoning, applying it to long-context scenarios is hindered by sparsity of outcome rewards. This limitation fails to penalize ungrounded "lucky guesses," leaving the critical process of needle-in-a-haystack evidence retrieval largely unsupervised. To address this, we propose EAPO (Evidence-Augmented Policy Optimization). We first establish the Evidence-Augmented Reasoning paradigm, validating via Tree-Structured Evidence Sampling that precise evidence extraction is the decisive bottleneck for long-context reasoning. Guided by this insight, EAPO introduces a specialized RL algorithm where a reward model computes a Group-Relative Evidence Reward, providing dense process supervision to explicitly improve evidence quality. To sustain accurate supervision throughout training, we further incorporate an Adaptive Reward-Policy Co-Evolution mechanism. This mechanism iteratively refines the reward model using outcome-consistent rollouts, sharpening its discriminative capability to ensure precise process guidance. Comprehensive evaluations across eight benchmarks demonstrate that EAPO significantly enhances long-context reasoning performance compared to SOTA baselines.
{
"annotation_id": "54d7bea6-6b95-4c2a-80aa-00d1ff7d782c",
"date_created": "2026-02-17T05:53:23.626000Z",
"date_modified": "2026-02-17T05:53:23.626000Z",
"file_hash": "d15c334f30617aa233456bac4872d8671621bd8d50cc361bcbafeebac3115ad3",
"private": false,
"record": {
"abstract": "While Reinforcement Learning (RL) has advanced LLM reasoning, applying it to long-context scenarios is hindered by sparsity of outcome rewards. This limitation fails to penalize ungrounded \"lucky guesses,\" leaving the critical process of needle-in-a-haystack evidence retrieval largely unsupervised. To address this, we propose EAPO (Evidence-Augmented Policy Optimization). We first establish the Evidence-Augmented Reasoning paradigm, validating via Tree-Structured Evidence Sampling that precise evidence extraction is the decisive bottleneck for long-context reasoning. Guided by this insight, EAPO introduces a specialized RL algorithm where a reward model computes a Group-Relative Evidence Reward, providing dense process supervision to explicitly improve evidence quality. To sustain accurate supervision throughout training, we further incorporate an Adaptive Reward-Policy Co-Evolution mechanism. This mechanism iteratively refines the reward model using outcome-consistent rollouts, sharpening its discriminative capability to ensure precise process guidance. Comprehensive evaluations across eight benchmarks demonstrate that EAPO significantly enhances long-context reasoning performance compared to SOTA baselines.",
"arxiv_id": "2601.10306",
"authors": [
"Xin Guan",
"Zijian Li",
"Shen Huang",
"Pengjun Xie",
"Jingren Zhou",
"Jiuxin Cao"
],
"categories": [
"cs.AI",
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning",
"url": "https://arxiv.org/abs/2601.10306",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "9aa1bd49-9915-4f18-9a25-7db022378c2a",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}