dorsal/arxiv
View SchemaRSA-Bench: Benchmarking Audio Large Models in Real-World Acoustic Scenarios
| Authors | Yibo Zhang, Liang Lin, Kaiwen Luo, Shilinlu Yan, Jin Wang, Yaoqi Guo, Yitian Chen, Yalan Qin, Zhenhong Zhou, Kun Wang, Li Sun |
|---|---|
| Categories | |
| ArXiv ID | 2601.10384vv1 |
| URL | https://arxiv.org/abs/2601.10384 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
While Audio Large Models (ALMs) have achieved remarkable proficiency, their robustness remains brittle in real-world deployment. Existing evaluations largely rely on synthetic Gaussian noise or simplistic single-source interference, failing to capture the intricate, multi-layered acoustic dynamics -- or ``Acoustic Ecology'' -- that characterize authentic physical environments. To bridge this ecological gap, we introduce \textbf{RSA-Bench}, a comprehensive robustness benchmark designed to stress-test ALLMs through high-fidelity auditory scene simulations. Unlike traditional methods, we construct evaluation samples by naturally superimposing diverse environmental soundscapes -- spanning \textit{Pasture}, \textit{Extreme Weather}, \textit{Classroom}, and \textit{Outdoors} -- onto clean speech signals across a spectrum of interference intensities. By evaluating models on six core tasks ranging from fundamental perception to complex reasoning, our study unveils three macro-level insights: \textbf{(I) The Perception-Cognition Gap:} Models maintain relative resilience in low-level recognition but suffer a \textbf{functional collapse} in high-order reasoning tasks under stress; \textbf{(II) Scenario Sensitivity:} ``Vocal-like'' interference (e.g., background laughter) proves significantly more destructive than mechanical noise, challenging the model's auditory attention mechanisms; and \textbf{(III) The Denoising Paradox:} Standard speech enhancement often exacerbates performance degradation, as ALLMs prove highly sensitive to the semantic distortions introduced by denoising artifacts.
{
"annotation_id": "df3835e0-56a9-449d-81d3-510368e7ad98",
"date_created": "2026-02-17T05:53:24.384000Z",
"date_modified": "2026-02-17T05:53:24.384000Z",
"file_hash": "9563281bdc676ab43e72d08c1700002519a31453aa420228b3d236a7a749d6e9",
"private": false,
"record": {
"abstract": "While Audio Large Models (ALMs) have achieved remarkable proficiency, their robustness remains brittle in real-world deployment. Existing evaluations largely rely on synthetic Gaussian noise or simplistic single-source interference, failing to capture the intricate, multi-layered acoustic dynamics -- or ``Acoustic Ecology\u0027\u0027 -- that characterize authentic physical environments. To bridge this ecological gap, we introduce \\textbf{RSA-Bench}, a comprehensive robustness benchmark designed to stress-test ALLMs through high-fidelity auditory scene simulations. Unlike traditional methods, we construct evaluation samples by naturally superimposing diverse environmental soundscapes -- spanning \\textit{Pasture}, \\textit{Extreme Weather}, \\textit{Classroom}, and \\textit{Outdoors} -- onto clean speech signals across a spectrum of interference intensities. By evaluating models on six core tasks ranging from fundamental perception to complex reasoning, our study unveils three macro-level insights: \\textbf{(I) The Perception-Cognition Gap:} Models maintain relative resilience in low-level recognition but suffer a \\textbf{functional collapse} in high-order reasoning tasks under stress; \\textbf{(II) Scenario Sensitivity:} ``Vocal-like\u0027\u0027 interference (e.g., background laughter) proves significantly more destructive than mechanical noise, challenging the model\u0027s auditory attention mechanisms; and \\textbf{(III) The Denoising Paradox:} Standard speech enhancement often exacerbates performance degradation, as ALLMs prove highly sensitive to the semantic distortions introduced by denoising artifacts.",
"arxiv_id": "2601.10384",
"authors": [
"Yibo Zhang",
"Liang Lin",
"Kaiwen Luo",
"Shilinlu Yan",
"Jin Wang",
"Yaoqi Guo",
"Yitian Chen",
"Yalan Qin",
"Zhenhong Zhou",
"Kun Wang",
"Li Sun"
],
"categories": [
"cs.SD"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "RSA-Bench: Benchmarking Audio Large Models in Real-World Acoustic Scenarios",
"url": "https://arxiv.org/abs/2601.10384",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "5b31dad8-551c-42c0-ac65-2fc32d68f94b",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}