dorsal/arxiv
View Schema3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence
| Authors | Hao Tang, Ting Huang, Zeyu Zhang |
|---|---|
| Categories | |
| ArXiv ID | 2601.06496vv1 |
| URL | https://arxiv.org/abs/2601.06496 |
| License | http://creativecommons.org/licenses/by-nc-sa/4.0/ |
Abstract
Spatial intelligence refers to the ability to perceive, reason about, and describe objects and their relationships within three-dimensional environments, forming a foundation for embodied perception and scene understanding. 3D captioning aims to describe 3D scenes in natural language; however, it remains challenging due to the sparsity and irregularity of point clouds and, more critically, the weak grounding and limited out-of-distribution (OOD) generalization of existing captioners across drastically different environments, including indoor and outdoor 3D scenes. To address this challenge, we propose 3D CoCa v2, a generalizable 3D captioning framework that unifies contrastive vision-language learning with 3D caption generation and further improves robustness via test-time search (TTS) without updating the captioner parameters. 3D CoCa v2 builds on a frozen CLIP-based semantic prior, a spatially-aware 3D scene encoder for geometry, and a multimodal decoder jointly optimized with contrastive and captioning objectives, avoiding external detectors or handcrafted proposals. At inference, TTS produces diverse caption candidates and performs reward-guided selection using a compact scene summary. Experiments show improvements over 3D CoCa of +1.50 CIDEr@0.5IoU on ScanRefer and +1.61 CIDEr@0.5IoU on Nr3D, and +3.8 CIDEr@0.25 in zero-shot OOD evaluation on TOD3Cap. Code will be released at https://github.com/AIGeeksGroup/3DCoCav2.
{
"annotation_id": "6e911bf2-1497-4837-8469-30bb7c1755cd",
"date_created": "2026-02-17T05:53:08.816000Z",
"date_modified": "2026-02-17T05:53:08.816000Z",
"file_hash": "412a4d0c5290a4181a91ac4df1a6b78b7f88428ce3ab2119d5f9384ccc38db88",
"private": false,
"record": {
"abstract": "Spatial intelligence refers to the ability to perceive, reason about, and describe objects and their relationships within three-dimensional environments, forming a foundation for embodied perception and scene understanding. 3D captioning aims to describe 3D scenes in natural language; however, it remains challenging due to the sparsity and irregularity of point clouds and, more critically, the weak grounding and limited out-of-distribution (OOD) generalization of existing captioners across drastically different environments, including indoor and outdoor 3D scenes. To address this challenge, we propose 3D CoCa v2, a generalizable 3D captioning framework that unifies contrastive vision-language learning with 3D caption generation and further improves robustness via test-time search (TTS) without updating the captioner parameters. 3D CoCa v2 builds on a frozen CLIP-based semantic prior, a spatially-aware 3D scene encoder for geometry, and a multimodal decoder jointly optimized with contrastive and captioning objectives, avoiding external detectors or handcrafted proposals. At inference, TTS produces diverse caption candidates and performs reward-guided selection using a compact scene summary. Experiments show improvements over 3D CoCa of +1.50 CIDEr@0.5IoU on ScanRefer and +1.61 CIDEr@0.5IoU on Nr3D, and +3.8 CIDEr@0.25 in zero-shot OOD evaluation on TOD3Cap. Code will be released at https://github.com/AIGeeksGroup/3DCoCav2.",
"arxiv_id": "2601.06496",
"authors": [
"Hao Tang",
"Ting Huang",
"Zeyu Zhang"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by-nc-sa/4.0/",
"title": "3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence",
"url": "https://arxiv.org/abs/2601.06496",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e352048a-670c-4219-af58-0f79cce1e490",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}