dorsal/arxiv
View SchemaFrom Local Windows to Adaptive Candidates via Individualized Exploratory: Rethinking Attention for Image Super-Resolution
| Authors | Chunyu Meng, Wei Long, Shuhang Gu |
|---|---|
| Categories | |
| ArXiv ID | 2601.08341vv1 |
| URL | https://arxiv.org/abs/2601.08341 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Single Image Super-Resolution (SISR) is a fundamental computer vision task that aims to reconstruct a high-resolution (HR) image from a low-resolution (LR) input. Transformer-based methods have achieved remarkable performance by modeling long-range dependencies in degraded images. However, their feature-intensive attention computation incurs high computational cost. To improve efficiency, most existing approaches partition images into fixed groups and restrict attention within each group. Such group-wise attention overlooks the inherent asymmetry in token similarities, thereby failing to enable flexible and token-adaptive attention computation. To address this limitation, we propose the Individualized Exploratory Transformer (IET), which introduces a novel Individualized Exploratory Attention (IEA) mechanism that allows each token to adaptively select its own content-aware and independent attention candidates. This token-adaptive and asymmetric design enables more precise information aggregation while maintaining computational efficiency. Extensive experiments on standard SR benchmarks demonstrate that IET achieves state-of-the-art performance under comparable computational complexity.
{
"annotation_id": "6580e574-7625-449b-b5db-443c05db9b98",
"date_created": "2026-02-17T05:53:15.166000Z",
"date_modified": "2026-02-17T05:53:15.166000Z",
"file_hash": "5ffd6b8f23238b477b40e029bb3e5d507af58ad26584ec72f02c6f82dccd9973",
"private": false,
"record": {
"abstract": "Single Image Super-Resolution (SISR) is a fundamental computer vision task that aims to reconstruct a high-resolution (HR) image from a low-resolution (LR) input. Transformer-based methods have achieved remarkable performance by modeling long-range dependencies in degraded images. However, their feature-intensive attention computation incurs high computational cost. To improve efficiency, most existing approaches partition images into fixed groups and restrict attention within each group. Such group-wise attention overlooks the inherent asymmetry in token similarities, thereby failing to enable flexible and token-adaptive attention computation. To address this limitation, we propose the Individualized Exploratory Transformer (IET), which introduces a novel Individualized Exploratory Attention (IEA) mechanism that allows each token to adaptively select its own content-aware and independent attention candidates. This token-adaptive and asymmetric design enables more precise information aggregation while maintaining computational efficiency. Extensive experiments on standard SR benchmarks demonstrate that IET achieves state-of-the-art performance under comparable computational complexity.",
"arxiv_id": "2601.08341",
"authors": [
"Chunyu Meng",
"Wei Long",
"Shuhang Gu"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "From Local Windows to Adaptive Candidates via Individualized Exploratory: Rethinking Attention for Image Super-Resolution",
"url": "https://arxiv.org/abs/2601.08341",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e3270ce5-51d3-4d1f-9902-4a843f2dd1f3",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}