dorsal/arxiv
View SchemaOUTLINEFORGE: Hierarchical Reinforcement Learning with Explicit States for Scientific Writing
| Authors | Yilin Bao, Ziyao He, Zayden Yang |
|---|---|
| Categories | |
| ArXiv ID | 2601.09858vv1 |
| URL | https://arxiv.org/abs/2601.09858 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Scientific paper generation requires document-level planning and factual grounding, but current large language models, despite their strong local fluency, often fail in global structure, input coverage, and citation consistency. We present a reinforcement learning framework that casts scientific outline construction as a long-horizon planning problem over hierarchical document structures. Our approach models edit evolving outlines through structured actions, enabling the system to incrementally build a complete scientific manuscript. To support effective and stabilize learning,we introduce a two-stage optimization procedure consisting of (i) backward outline reconstruction from partial plans to enforce global structural consistency, and (ii) forward value-guided reinforcement learning with rewards explicitly modeling scientific correctness, discourse coherence, and citation fidelity. In addition, We further introduce a benchmark for scientific paper generation that evaluates document planning, input utilization, reference faithfulness, outline organization, and content-level factual accuracy. Our results show consistent improvements over strong neural and LLM baselines, particularly in long-range structural coherence and citation reliability.
{
"annotation_id": "d86e8b1f-7f78-4700-97d7-611b31b729f3",
"date_created": "2026-02-17T05:53:23.735000Z",
"date_modified": "2026-02-17T05:53:23.735000Z",
"file_hash": "b4195e6619aa402afccd475fec9da88fc8f20d6218bbabad46807ab17341b86f",
"private": false,
"record": {
"abstract": "Scientific paper generation requires document-level planning and factual grounding, but current large language models, despite their strong local fluency, often fail in global structure, input coverage, and citation consistency. We present a reinforcement learning framework that casts scientific outline construction as a long-horizon planning problem over hierarchical document structures. Our approach models edit evolving outlines through structured actions, enabling the system to incrementally build a complete scientific manuscript. To support effective and stabilize learning,we introduce a two-stage optimization procedure consisting of (i) backward outline reconstruction from partial plans to enforce global structural consistency, and (ii) forward value-guided reinforcement learning with rewards explicitly modeling scientific correctness, discourse coherence, and citation fidelity. In addition, We further introduce a benchmark for scientific paper generation that evaluates document planning, input utilization, reference faithfulness, outline organization, and content-level factual accuracy. Our results show consistent improvements over strong neural and LLM baselines, particularly in long-range structural coherence and citation reliability.",
"arxiv_id": "2601.09858",
"authors": [
"Yilin Bao",
"Ziyao He",
"Zayden Yang"
],
"categories": [
"cs.CL",
"cs.AI",
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "OUTLINEFORGE: Hierarchical Reinforcement Learning with Explicit States for Scientific Writing",
"url": "https://arxiv.org/abs/2601.09858",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "62814d3a-166d-4976-8969-9e9087a10927",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}