dorsal/arxiv
View SchemaExploiting DINOv3-Based Self-Supervised Features for Robust Few-Shot Medical Image Segmentation
| Authors | Guoping Xu, Jayaram K. Udupa, Weiguo Lu, You Zhang |
|---|---|
| Categories | |
| ArXiv ID | 2601.08078vv1 |
| URL | https://arxiv.org/abs/2601.08078 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Deep learning-based automatic medical image segmentation plays a critical role in clinical diagnosis and treatment planning but remains challenging in few-shot scenarios due to the scarcity of annotated training data. Recently, self-supervised foundation models such as DINOv3, which were trained on large natural image datasets, have shown strong potential for dense feature extraction that can help with the few-shot learning challenge. Yet, their direct application to medical images is hindered by domain differences. In this work, we propose DINO-AugSeg, a novel framework that leverages DINOv3 features to address the few-shot medical image segmentation challenge. Specifically, we introduce WT-Aug, a wavelet-based feature-level augmentation module that enriches the diversity of DINOv3-extracted features by perturbing frequency components, and CG-Fuse, a contextual information-guided fusion module that exploits cross-attention to integrate semantic-rich low-resolution features with spatially detailed high-resolution features. Extensive experiments on six public benchmarks spanning five imaging modalities, including MRI, CT, ultrasound, endoscopy, and dermoscopy, demonstrate that DINO-AugSeg consistently outperforms existing methods under limited-sample conditions. The results highlight the effectiveness of incorporating wavelet-domain augmentation and contextual fusion for robust feature representation, suggesting DINO-AugSeg as a promising direction for advancing few-shot medical image segmentation. Code and data will be made available on https://github.com/apple1986/DINO-AugSeg.
{
"annotation_id": "82024164-fbea-4e0a-b91f-f88d4865b89a",
"date_created": "2026-02-17T05:53:16.068000Z",
"date_modified": "2026-02-17T05:53:16.068000Z",
"file_hash": "610bca4046a7e48e02cd928ebea4ddd0b7dade74add5cb028d8f8de1dec4695c",
"private": false,
"record": {
"abstract": "Deep learning-based automatic medical image segmentation plays a critical role in clinical diagnosis and treatment planning but remains challenging in few-shot scenarios due to the scarcity of annotated training data. Recently, self-supervised foundation models such as DINOv3, which were trained on large natural image datasets, have shown strong potential for dense feature extraction that can help with the few-shot learning challenge. Yet, their direct application to medical images is hindered by domain differences. In this work, we propose DINO-AugSeg, a novel framework that leverages DINOv3 features to address the few-shot medical image segmentation challenge. Specifically, we introduce WT-Aug, a wavelet-based feature-level augmentation module that enriches the diversity of DINOv3-extracted features by perturbing frequency components, and CG-Fuse, a contextual information-guided fusion module that exploits cross-attention to integrate semantic-rich low-resolution features with spatially detailed high-resolution features. Extensive experiments on six public benchmarks spanning five imaging modalities, including MRI, CT, ultrasound, endoscopy, and dermoscopy, demonstrate that DINO-AugSeg consistently outperforms existing methods under limited-sample conditions. The results highlight the effectiveness of incorporating wavelet-domain augmentation and contextual fusion for robust feature representation, suggesting DINO-AugSeg as a promising direction for advancing few-shot medical image segmentation. Code and data will be made available on https://github.com/apple1986/DINO-AugSeg.",
"arxiv_id": "2601.08078",
"authors": [
"Guoping Xu",
"Jayaram K. Udupa",
"Weiguo Lu",
"You Zhang"
],
"categories": [
"cs.CV",
"cs.CE",
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Exploiting DINOv3-Based Self-Supervised Features for Robust Few-Shot Medical Image Segmentation",
"url": "https://arxiv.org/abs/2601.08078",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e323e125-ef76-4cee-9d9a-53e3c90442b7",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}