dorsal/arxiv
View SchemaReward-Preserving Attacks For Robust Reinforcement Learning
| Authors | Lucas Schott, Elies Gherbi, Hatem Hajri, Sylvain Lamprier |
|---|---|
| Categories | |
| ArXiv ID | 2601.07118vv2 |
| URL | https://arxiv.org/abs/2601.07118 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Adversarial robustness in RL is difficult because perturbations affect entire trajectories: strong attacks can break learning, while weak attacks yield little robustness, and the appropriate strength varies by state. We propose $\alpha$-reward-preserving attacks, which adapt the strength of the adversary so that an $\alpha$ fraction of the nominal-to-worst-case return gap remains achievable at each state. In deep RL, we use a gradient-based attack direction and learn a state-dependent magnitude $\eta\le \eta_{\mathcal B}$ selected via a critic $Q^{\pi}_\alpha((s,a),\eta)$ trained off-policy over diverse radii. This adaptive tuning calibrates attack strength and, with intermediate $\alpha$, improves robustness across radii while preserving nominal performance, outperforming fixed- and random-radius baselines.
{
"annotation_id": "94827e15-b024-4380-bb4e-706e3c62dbc2",
"date_created": "2026-02-17T05:53:11.278000Z",
"date_modified": "2026-02-17T05:53:11.278000Z",
"file_hash": "7d55fac4d3ffaf2ad85bb447ab2f9e76960d5a82a4cbf628d4e4de0ad1df4fe8",
"private": false,
"record": {
"abstract": "Adversarial robustness in RL is difficult because perturbations affect entire trajectories: strong attacks can break learning, while weak attacks yield little robustness, and the appropriate strength varies by state. We propose $\\alpha$-reward-preserving attacks, which adapt the strength of the adversary so that an $\\alpha$ fraction of the nominal-to-worst-case return gap remains achievable at each state. In deep RL, we use a gradient-based attack direction and learn a state-dependent magnitude $\\eta\\le \\eta_{\\mathcal B}$ selected via a critic $Q^{\\pi}_\\alpha((s,a),\\eta)$ trained off-policy over diverse radii. This adaptive tuning calibrates attack strength and, with intermediate $\\alpha$, improves robustness across radii while preserving nominal performance, outperforming fixed- and random-radius baselines.",
"arxiv_id": "2601.07118",
"authors": [
"Lucas Schott",
"Elies Gherbi",
"Hatem Hajri",
"Sylvain Lamprier"
],
"categories": [
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Reward-Preserving Attacks For Robust Reinforcement Learning",
"url": "https://arxiv.org/abs/2601.07118",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "494b46b6-89b1-4bbe-9059-0b50ce1c5022",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}