dorsal/arxiv
View SchemaDIFF-MF: A Difference-Driven Channel-Spatial State Space Model for Multi-Modal Image Fusion
| Authors | Yiming Sun, Zifan Ye, Qinghua Hu, Pengfei Zhu |
|---|---|
| Categories | |
| ArXiv ID | 2601.05538vv1 |
| URL | https://arxiv.org/abs/2601.05538 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Multi-modal image fusion aims to integrate complementary information from multiple source images to produce high-quality fused images with enriched content. Although existing approaches based on state space model have achieved satisfied performance with high computational efficiency, they tend to either over-prioritize infrared intensity at the cost of visible details, or conversely, preserve visible structure while diminishing thermal target salience. To overcome these challenges, we propose DIFF-MF, a novel difference-driven channel-spatial state space model for multi-modal image fusion. Our approach leverages feature discrepancy maps between modalities to guide feature extraction, followed by a fusion process across both channel and spatial dimensions. In the channel dimension, a channel-exchange module enhances channel-wise interaction through cross-attention dual state space modeling, enabling adaptive feature reweighting. In the spatial dimension, a spatial-exchange module employs cross-modal state space scanning to achieve comprehensive spatial fusion. By efficiently capturing global dependencies while maintaining linear computational complexity, DIFF-MF effectively integrates complementary multi-modal features. Experimental results on the driving scenarios and low-altitude UAV datasets demonstrate that our method outperforms existing approaches in both visual quality and quantitative evaluation.
{
"annotation_id": "60397bc9-9465-47fc-9785-11ffcebd1553",
"date_created": "2026-02-17T05:53:05.046000Z",
"date_modified": "2026-02-17T05:53:05.046000Z",
"file_hash": "4e0c102108a745185b34394e95d8e07a7bbc55a1a7c5e134c48f805b4e75fb9c",
"private": false,
"record": {
"abstract": "Multi-modal image fusion aims to integrate complementary information from multiple source images to produce high-quality fused images with enriched content. Although existing approaches based on state space model have achieved satisfied performance with high computational efficiency, they tend to either over-prioritize infrared intensity at the cost of visible details, or conversely, preserve visible structure while diminishing thermal target salience. To overcome these challenges, we propose DIFF-MF, a novel difference-driven channel-spatial state space model for multi-modal image fusion. Our approach leverages feature discrepancy maps between modalities to guide feature extraction, followed by a fusion process across both channel and spatial dimensions. In the channel dimension, a channel-exchange module enhances channel-wise interaction through cross-attention dual state space modeling, enabling adaptive feature reweighting. In the spatial dimension, a spatial-exchange module employs cross-modal state space scanning to achieve comprehensive spatial fusion. By efficiently capturing global dependencies while maintaining linear computational complexity, DIFF-MF effectively integrates complementary multi-modal features. Experimental results on the driving scenarios and low-altitude UAV datasets demonstrate that our method outperforms existing approaches in both visual quality and quantitative evaluation.",
"arxiv_id": "2601.05538",
"authors": [
"Yiming Sun",
"Zifan Ye",
"Qinghua Hu",
"Pengfei Zhu"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "DIFF-MF: A Difference-Driven Channel-Spatial State Space Model for Multi-Modal Image Fusion",
"url": "https://arxiv.org/abs/2601.05538",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "dc0489c3-5a3b-4ec7-b537-3b60dfcfe0e9",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}