dorsal/arxiv
View SchemaFailure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation
| Authors | Huanyu Li, Kun Lei, Sheng Zang, Kaizhe Hu, Yongyuan Liang, Bo An, Xiaoli Li, Huazhe Xu |
|---|---|
| Categories | |
| ArXiv ID | 2601.07821vv1 |
| URL | https://arxiv.org/abs/2601.07821 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Post-training algorithms based on deep reinforcement learning can push the limits of robotic models for specific objectives, such as generalizability, accuracy, and robustness. However, Intervention-requiring Failures (IR Failures) (e.g., a robot spilling water or breaking fragile glass) during real-world exploration happen inevitably, hindering the practical deployment of such a paradigm. To tackle this, we introduce Failure-Aware Offline-to-Online Reinforcement Learning (FARL), a new paradigm minimizing failures during real-world reinforcement learning. We create FailureBench, a benchmark that incorporates common failure scenarios requiring human intervention, and propose an algorithm that integrates a world-model-based safety critic and a recovery policy trained offline to prevent failures during online exploration. Extensive simulation and real-world experiments demonstrate the effectiveness of FARL in significantly reducing IR Failures while improving performance and generalization during online reinforcement learning post-training. FARL reduces IR Failures by 73.1% while elevating performance by 11.3% on average during real-world RL post-training. Videos and code are available at https://failure-aware-rl.github.io.
{
"annotation_id": "c2e1318f-b06f-4dd5-90cb-8c88fe7213db",
"date_created": "2026-02-17T05:53:11.947000Z",
"date_modified": "2026-02-17T05:53:11.947000Z",
"file_hash": "9583954f351a6634f139c55103994ba87aa9e22f7d6dc7234856936579ce1c1b",
"private": false,
"record": {
"abstract": "Post-training algorithms based on deep reinforcement learning can push the limits of robotic models for specific objectives, such as generalizability, accuracy, and robustness. However, Intervention-requiring Failures (IR Failures) (e.g., a robot spilling water or breaking fragile glass) during real-world exploration happen inevitably, hindering the practical deployment of such a paradigm. To tackle this, we introduce Failure-Aware Offline-to-Online Reinforcement Learning (FARL), a new paradigm minimizing failures during real-world reinforcement learning. We create FailureBench, a benchmark that incorporates common failure scenarios requiring human intervention, and propose an algorithm that integrates a world-model-based safety critic and a recovery policy trained offline to prevent failures during online exploration. Extensive simulation and real-world experiments demonstrate the effectiveness of FARL in significantly reducing IR Failures while improving performance and generalization during online reinforcement learning post-training. FARL reduces IR Failures by 73.1% while elevating performance by 11.3% on average during real-world RL post-training. Videos and code are available at https://failure-aware-rl.github.io.",
"arxiv_id": "2601.07821",
"authors": [
"Huanyu Li",
"Kun Lei",
"Sheng Zang",
"Kaizhe Hu",
"Yongyuan Liang",
"Bo An",
"Xiaoli Li",
"Huazhe Xu"
],
"categories": [
"cs.RO",
"cs.AI",
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation",
"url": "https://arxiv.org/abs/2601.07821",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "a3b6275f-594e-436c-9fe3-ec5509adbc52",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}