dorsal/arxiv
View SchemaFigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
| Authors | Jifeng Song, Arun Das, Pan Wang, Hui Ji, Kun Zhao, Yufei Huang |
|---|---|
| Categories | |
| ArXiv ID | 2601.08026vv1 |
| URL | https://arxiv.org/abs/2601.08026 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Scientific compound figures combine multiple labeled panels into a single image, but captions in real pipelines are often missing or only provide figure-level summaries, making panel-level understanding difficult. In this paper, we propose FigEx2, visual-conditioned framework that localizes panels and generates panel-wise captions directly from the compound figure. To mitigate the impact of diverse phrasing in open-ended captioning, we introduce a noise-aware gated fusion module that adaptively filters token-level features to stabilize the detection query space. Furthermore, we employ a staged optimization strategy combining supervised learning with reinforcement learning (RL), utilizing CLIP-based alignment and BERTScore-based semantic rewards to enforce strict multimodal consistency. To support high-quality supervision, we curate BioSci-Fig-Cap, a refined benchmark for panel-level grounding, alongside cross-disciplinary test suites in physics and chemistry. Experimental results demonstrate that FigEx2 achieves a superior 0.726 mAP@0.5:0.95 for detection and significantly outperforms Qwen3-VL-8B by 0.51 in METEOR and 0.24 in BERTScore. Notably, FigEx2 exhibits remarkable zero-shot transferability to out-of-distribution scientific domains without any fine-tuning.
{
"annotation_id": "e1c2ec07-8153-4f4d-b4de-a14db7a5e841",
"date_created": "2026-02-17T05:53:15.040000Z",
"date_modified": "2026-02-17T05:53:15.040000Z",
"file_hash": "da1dc5f9a9c5758454fd2bde0bdb08ce0ae8353690b1b522a124f8f115097cf7",
"private": false,
"record": {
"abstract": "Scientific compound figures combine multiple labeled panels into a single image, but captions in real pipelines are often missing or only provide figure-level summaries, making panel-level understanding difficult. In this paper, we propose FigEx2, visual-conditioned framework that localizes panels and generates panel-wise captions directly from the compound figure. To mitigate the impact of diverse phrasing in open-ended captioning, we introduce a noise-aware gated fusion module that adaptively filters token-level features to stabilize the detection query space. Furthermore, we employ a staged optimization strategy combining supervised learning with reinforcement learning (RL), utilizing CLIP-based alignment and BERTScore-based semantic rewards to enforce strict multimodal consistency. To support high-quality supervision, we curate BioSci-Fig-Cap, a refined benchmark for panel-level grounding, alongside cross-disciplinary test suites in physics and chemistry. Experimental results demonstrate that FigEx2 achieves a superior 0.726 mAP@0.5:0.95 for detection and significantly outperforms Qwen3-VL-8B by 0.51 in METEOR and 0.24 in BERTScore. Notably, FigEx2 exhibits remarkable zero-shot transferability to out-of-distribution scientific domains without any fine-tuning.",
"arxiv_id": "2601.08026",
"authors": [
"Jifeng Song",
"Arun Das",
"Pan Wang",
"Hui Ji",
"Kun Zhao",
"Yufei Huang"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures",
"url": "https://arxiv.org/abs/2601.08026",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "a8401586-ee96-4758-be44-9dde7eeb1d69",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}