dorsal/arxiv
View SchemaVideo-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
| Authors | Rui Zhu, Xin Shen, Shuchen Wu, Chenxi Miao, Xin Yu, Yang Li, Weikang Li, Deguo Xia, Jizhou Huang |
|---|---|
| Categories | |
| ArXiv ID | 2601.09430vv1 |
| URL | https://arxiv.org/abs/2601.09430 |
| License | http://creativecommons.org/licenses/by-nc-nd/4.0/ |
Abstract
Spatial reasoning has emerged as a critical capability for Multimodal Large Language Models (MLLMs), drawing increasing attention and rapid advancement. However, existing benchmarks primarily focus on single-step perception-to-judgment tasks, leaving scenarios requiring complex visual-spatial logical chains significantly underexplored. To bridge this gap, we introduce Video-MSR, the first benchmark specifically designed to evaluate Multi-hop Spatial Reasoning (MSR) in dynamic video scenarios. Video-MSR systematically probes MSR capabilities through four distinct tasks: Constrained Localization, Chain-based Reference Retrieval, Route Planning, and Counterfactual Physical Deduction. Our benchmark comprises 3,052 high-quality video instances with 4,993 question-answer pairs, constructed via a scalable, visually-grounded pipeline combining advanced model generation with rigorous human verification. Through a comprehensive evaluation of 20 state-of-the-art MLLMs, we uncover significant limitations, revealing that while models demonstrate proficiency in surface-level perception, they exhibit distinct performance drops in MSR tasks, frequently suffering from spatial disorientation and hallucination during multi-step deductions. To mitigate these shortcomings and empower models with stronger MSR capabilities, we further curate MSR-9K, a specialized instruction-tuning dataset, and fine-tune Qwen-VL, achieving a +7.82% absolute improvement on Video-MSR. Our results underscore the efficacy of multi-hop spatial instruction data and establish Video-MSR as a vital foundation for future research. The code and data will be available at https://github.com/ruiz-nju/Video-MSR.
{
"annotation_id": "28f1e892-3ef9-45ae-bec8-1943aa304376",
"date_created": "2026-02-17T05:53:20.337000Z",
"date_modified": "2026-02-17T05:53:20.337000Z",
"file_hash": "1d13e87270a5fa01818d94281160a56d44a3b799ebb5a46bb49de0ccf4582e93",
"private": false,
"record": {
"abstract": "Spatial reasoning has emerged as a critical capability for Multimodal Large Language Models (MLLMs), drawing increasing attention and rapid advancement. However, existing benchmarks primarily focus on single-step perception-to-judgment tasks, leaving scenarios requiring complex visual-spatial logical chains significantly underexplored. To bridge this gap, we introduce Video-MSR, the first benchmark specifically designed to evaluate Multi-hop Spatial Reasoning (MSR) in dynamic video scenarios. Video-MSR systematically probes MSR capabilities through four distinct tasks: Constrained Localization, Chain-based Reference Retrieval, Route Planning, and Counterfactual Physical Deduction. Our benchmark comprises 3,052 high-quality video instances with 4,993 question-answer pairs, constructed via a scalable, visually-grounded pipeline combining advanced model generation with rigorous human verification. Through a comprehensive evaluation of 20 state-of-the-art MLLMs, we uncover significant limitations, revealing that while models demonstrate proficiency in surface-level perception, they exhibit distinct performance drops in MSR tasks, frequently suffering from spatial disorientation and hallucination during multi-step deductions. To mitigate these shortcomings and empower models with stronger MSR capabilities, we further curate MSR-9K, a specialized instruction-tuning dataset, and fine-tune Qwen-VL, achieving a +7.82% absolute improvement on Video-MSR. Our results underscore the efficacy of multi-hop spatial instruction data and establish Video-MSR as a vital foundation for future research. The code and data will be available at https://github.com/ruiz-nju/Video-MSR.",
"arxiv_id": "2601.09430",
"authors": [
"Rui Zhu",
"Xin Shen",
"Shuchen Wu",
"Chenxi Miao",
"Xin Yu",
"Yang Li",
"Weikang Li",
"Deguo Xia",
"Jizhou Huang"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by-nc-nd/4.0/",
"title": "Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs",
"url": "https://arxiv.org/abs/2601.09430",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "4575602c-e611-4578-bce6-599a41f971f1",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}