dorsal/arxiv
View SchemaOffline Meta-Reinforcement Learning with Flow-Based Task Inference and Adaptive Correction of Feature Overgeneralization
| Authors | Min Wang, Xin Li, Mingzhong Wang, Hasnaa Bennis |
|---|---|
| Categories | |
| ArXiv ID | 2601.07164vv1 |
| URL | https://arxiv.org/abs/2601.07164 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Offline meta-reinforcement learning (OMRL) combines the strengths of learning from diverse datasets in offline RL with the adaptability to new tasks of meta-RL, promising safe and efficient knowledge acquisition by RL agents. However, OMRL still suffers extrapolation errors due to out-of-distribution (OOD) actions, compromised by broad task distributions and Markov Decision Process (MDP) ambiguity in meta-RL setups. Existing research indicates that the generalization of the $Q$ network affects the extrapolation error in offline RL. This paper investigates this relationship by decomposing the $Q$ value into feature and weight components, observing that while decomposition enhances adaptability and convergence in the case of high-quality data, it often leads to policy degeneration or collapse in complex tasks. We observe that decomposed $Q$ values introduce a large estimation bias when the feature encounters OOD samples, a phenomenon we term ''feature overgeneralization''. To address this issue, we propose FLORA, which identifies OOD samples by modeling feature distributions and estimating their uncertainties. FLORA integrates a return feedback mechanism to adaptively adjust feature components. Furthermore, to learn precise task representations, FLORA explicitly models the complex task distribution using a chain of invertible transformations. We theoretically and empirically demonstrate that FLORA achieves rapid adaptation and meta-policy improvement compared to baselines across various environments.
{
"annotation_id": "cd248e95-038f-4542-93e9-72fa573344af",
"date_created": "2026-02-17T05:53:12.439000Z",
"date_modified": "2026-02-17T05:53:12.439000Z",
"file_hash": "9c017018c63cce6a2888373e301ac008e134394041f5055fa602a593ff3658e8",
"private": false,
"record": {
"abstract": "Offline meta-reinforcement learning (OMRL) combines the strengths of learning from diverse datasets in offline RL with the adaptability to new tasks of meta-RL, promising safe and efficient knowledge acquisition by RL agents. However, OMRL still suffers extrapolation errors due to out-of-distribution (OOD) actions, compromised by broad task distributions and Markov Decision Process (MDP) ambiguity in meta-RL setups. Existing research indicates that the generalization of the $Q$ network affects the extrapolation error in offline RL. This paper investigates this relationship by decomposing the $Q$ value into feature and weight components, observing that while decomposition enhances adaptability and convergence in the case of high-quality data, it often leads to policy degeneration or collapse in complex tasks. We observe that decomposed $Q$ values introduce a large estimation bias when the feature encounters OOD samples, a phenomenon we term \u0027\u0027feature overgeneralization\u0027\u0027. To address this issue, we propose FLORA, which identifies OOD samples by modeling feature distributions and estimating their uncertainties. FLORA integrates a return feedback mechanism to adaptively adjust feature components. Furthermore, to learn precise task representations, FLORA explicitly models the complex task distribution using a chain of invertible transformations. We theoretically and empirically demonstrate that FLORA achieves rapid adaptation and meta-policy improvement compared to baselines across various environments.",
"arxiv_id": "2601.07164",
"authors": [
"Min Wang",
"Xin Li",
"Mingzhong Wang",
"Hasnaa Bennis"
],
"categories": [
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Offline Meta-Reinforcement Learning with Flow-Based Task Inference and Adaptive Correction of Feature Overgeneralization",
"url": "https://arxiv.org/abs/2601.07164",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "c9a7e17f-6cda-45f6-bc56-f641995696db",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}