dorsal/arxiv
View SchemaMMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
| Authors | Aditya Chaudhary, Sneha Barman, Mainak Singha, Ankit Jha, Girish Mishra, Biplab Banerjee |
|---|---|
| Categories | |
| ArXiv ID | 2601.08420vv1 |
| URL | https://arxiv.org/abs/2601.08420 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision-language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic-level understanding. MMLGNet employs modality-specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi-directional contrastive learning. Inspired by CLIP's training paradigm, our approach bridges the gap between high-dimensional remote sensing data and language-guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN-based encoders, outperforming several established multimodal visual-only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI.
{
"annotation_id": "6ba7cc1c-a139-4469-bc10-59c024edf3cc",
"date_created": "2026-02-17T05:53:16.197000Z",
"date_modified": "2026-02-17T05:53:16.197000Z",
"file_hash": "2965574eb2061fa196ff5e84e12f9cd1c25493e83f107b69ce896c637259e741",
"private": false,
"record": {
"abstract": "In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision-language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic-level understanding. MMLGNet employs modality-specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi-directional contrastive learning. Inspired by CLIP\u0027s training paradigm, our approach bridges the gap between high-dimensional remote sensing data and language-guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN-based encoders, outperforming several established multimodal visual-only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI.",
"arxiv_id": "2601.08420",
"authors": [
"Aditya Chaudhary",
"Sneha Barman",
"Mainak Singha",
"Ankit Jha",
"Girish Mishra",
"Biplab Banerjee"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP",
"url": "https://arxiv.org/abs/2601.08420",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "d8f8432c-ccb3-4322-9aa0-a2ce82e95ca1",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}