dorsal/arxiv
View SchemaReward-Preserving Attacks For Robust Reinforcement Learning
| Authors | Lucas Schott, Elies Gherbi, Hatem Hajri, Sylvain Lamprier |
|---|---|
| Categories | |
| ArXiv ID | 2601.07118vv1 |
| URL | https://arxiv.org/abs/2601.07118 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Adversarial robustness in RL is difficult because perturbations affect entire trajectories: strong attacks can break learning, while weak attacks yield little robustness, and the appropriate strength varies by state. We propose $\alpha$-reward-preserving attacks, which adapt the strength of the adversary so that an $\alpha$ fraction of the nominal-to-worst-case return gap remains achievable at each state. In deep RL, we use a gradient-based attack direction and learn a state-dependent magnitude $\eta\le \eta_{\mathcal B}$ selected via a critic $Q^{\pi}_\alpha((s,a),\eta)$ trained off-policy over diverse radii. This adaptive tuning calibrates attack strength and, with intermediate $\alpha$, improves robustness across radii while preserving nominal performance, outperforming fixed- and random-radius baselines.
{
"annotation_id": "dd3aeb6b-7cff-4717-a3c1-5d2740c63772",
"date_created": "2026-02-17T05:53:12.428000Z",
"date_modified": "2026-02-17T05:53:12.428000Z",
"file_hash": "79aa4ccbf9a81f922da8374a5f4056eb264bb0281c657c47c895ca4a874d2d13",
"private": false,
"record": {
"abstract": "Adversarial robustness in RL is difficult because perturbations affect entire trajectories: strong attacks can break learning, while weak attacks yield little robustness, and the appropriate strength varies by state. We propose $\\alpha$-reward-preserving attacks, which adapt the strength of the adversary so that an $\\alpha$ fraction of the nominal-to-worst-case return gap remains achievable at each state. In deep RL, we use a gradient-based attack direction and learn a state-dependent magnitude $\\eta\\le \\eta_{\\mathcal B}$ selected via a critic $Q^{\\pi}_\\alpha((s,a),\\eta)$ trained off-policy over diverse radii. This adaptive tuning calibrates attack strength and, with intermediate $\\alpha$, improves robustness across radii while preserving nominal performance, outperforming fixed- and random-radius baselines.",
"arxiv_id": "2601.07118",
"authors": [
"Lucas Schott",
"Elies Gherbi",
"Hatem Hajri",
"Sylvain Lamprier"
],
"categories": [
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Reward-Preserving Attacks For Robust Reinforcement Learning",
"url": "https://arxiv.org/abs/2601.07118",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e3c3aee7-82ae-48b6-bc45-d4189439bb57",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}