dorsal/arxiv
View SchemaEfficientFSL: Enhancing Few-Shot Classification via Query-Only Tuning in Vision Transformers
| Authors | Wenwen Liao, Hang Ruan, Jianbo Yu, Bing Song, YuansongWang, Xiaofeng Yang |
|---|---|
| Categories | |
| ArXiv ID | 2601.08499vv2 |
| URL | https://arxiv.org/abs/2601.08499 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Large models such as Vision Transformers (ViTs) have demonstrated remarkable superiority over smaller architectures like ResNet in few-shot classification, owing to their powerful representational capacity. However, fine-tuning such large models demands extensive GPU memory and prolonged training time, making them impractical for many real-world low-resource scenarios. To bridge this gap, we propose EfficientFSL, a query-only fine-tuning framework tailored specifically for few-shot classification with ViT, which achieves competitive performance while significantly reducing computational overhead. EfficientFSL fully leverages the knowledge embedded in the pre-trained model and its strong comprehension ability, achieving high classification accuracy with an extremely small number of tunable parameters. Specifically, we introduce a lightweight trainable Forward Block to synthesize task-specific queries that extract informative features from the intermediate representations of the pre-trained model in a query-only manner. We further propose a Combine Block to fuse multi-layer outputs, enhancing the depth and robustness of feature representations. Finally, a Support-Query Attention Block mitigates distribution shift by adjusting prototypes to align with the query set distribution. With minimal trainable parameters, EfficientFSL achieves state-of-the-art performance on four in-domain few-shot datasets and six cross-domain datasets, demonstrating its effectiveness in real-world applications.
{
"annotation_id": "23217c6f-4ce8-484b-90fd-9873dd1d9493",
"date_created": "2026-02-17T05:53:15.622000Z",
"date_modified": "2026-02-17T05:53:15.622000Z",
"file_hash": "67a7bc29028b78b92e3116a5fb58adaec8140e5dab09c7baceeb7a26f30c93ac",
"private": false,
"record": {
"abstract": "Large models such as Vision Transformers (ViTs) have demonstrated remarkable superiority over smaller architectures like ResNet in few-shot classification, owing to their powerful representational capacity. However, fine-tuning such large models demands extensive GPU memory and prolonged training time, making them impractical for many real-world low-resource scenarios. To bridge this gap, we propose EfficientFSL, a query-only fine-tuning framework tailored specifically for few-shot classification with ViT, which achieves competitive performance while significantly reducing computational overhead. EfficientFSL fully leverages the knowledge embedded in the pre-trained model and its strong comprehension ability, achieving high classification accuracy with an extremely small number of tunable parameters. Specifically, we introduce a lightweight trainable Forward Block to synthesize task-specific queries that extract informative features from the intermediate representations of the pre-trained model in a query-only manner. We further propose a Combine Block to fuse multi-layer outputs, enhancing the depth and robustness of feature representations. Finally, a Support-Query Attention Block mitigates distribution shift by adjusting prototypes to align with the query set distribution. With minimal trainable parameters, EfficientFSL achieves state-of-the-art performance on four in-domain few-shot datasets and six cross-domain datasets, demonstrating its effectiveness in real-world applications.",
"arxiv_id": "2601.08499",
"authors": [
"Wenwen Liao",
"Hang Ruan",
"Jianbo Yu",
"Bing Song",
"YuansongWang",
"Xiaofeng Yang"
],
"categories": [
"cs.CV",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "EfficientFSL: Enhancing Few-Shot Classification via Query-Only Tuning in Vision Transformers",
"url": "https://arxiv.org/abs/2601.08499",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "20c497b8-e753-4372-93db-ed89c8f36901",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}