dorsal/arxiv
View SchemaVideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
| Authors | Jiapeng Shi, Junke Wang, Zuyao You, Bo He, Zuxuan Wu |
|---|---|
| Categories | |
| ArXiv ID | 2601.07290vv1 |
| URL | https://arxiv.org/abs/2601.07290 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
This paper presents VideoLoom, a unified Video Large Language Model (Video LLM) for joint spatial-temporal understanding. To facilitate the development of fine-grained spatial and temporal localization capabilities, we curate LoomData-8.7k, a human-centric video dataset with temporally grounded and spatially localized captions. With this, VideoLoom achieves state-of-the-art or highly competitive performance across a variety of spatial and temporal benchmarks (e.g., 63.1 J&F on ReVOS for referring video object segmentation, and 48.3 R1@0.7 on Charades-STA for temporal grounding). In addition, we introduce LoomBench, a novel benchmark consisting of temporal, spatial, and compositional video-question pairs, enabling a comprehensive evaluation of Video LLMs from diverse aspects. Collectively, these contributions offer a universal and effective suite for joint spatial-temporal video understanding, setting a new standard in multimodal intelligence.
{
"annotation_id": "07a426f3-1cbe-402c-8267-3016d700fb87",
"date_created": "2026-02-17T05:53:12.656000Z",
"date_modified": "2026-02-17T05:53:12.656000Z",
"file_hash": "630b3ce9d0ddbc05ff395d60f9bf62c11045d9055be1dff2e0df0916c03f7ca3",
"private": false,
"record": {
"abstract": "This paper presents VideoLoom, a unified Video Large Language Model (Video LLM) for joint spatial-temporal understanding. To facilitate the development of fine-grained spatial and temporal localization capabilities, we curate LoomData-8.7k, a human-centric video dataset with temporally grounded and spatially localized captions. With this, VideoLoom achieves state-of-the-art or highly competitive performance across a variety of spatial and temporal benchmarks (e.g., 63.1 J\u0026F on ReVOS for referring video object segmentation, and 48.3 R1@0.7 on Charades-STA for temporal grounding). In addition, we introduce LoomBench, a novel benchmark consisting of temporal, spatial, and compositional video-question pairs, enabling a comprehensive evaluation of Video LLMs from diverse aspects. Collectively, these contributions offer a universal and effective suite for joint spatial-temporal video understanding, setting a new standard in multimodal intelligence.",
"arxiv_id": "2601.07290",
"authors": [
"Jiapeng Shi",
"Junke Wang",
"Zuyao You",
"Bo He",
"Zuxuan Wu"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding",
"url": "https://arxiv.org/abs/2601.07290",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "f724c117-1c7f-48cb-b077-27ca54cb98ef",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}