dorsal/arxiv
View SchemaS3-CLIP: Video Super Resolution for Person-ReID
| Authors | Tamas Endrei, Gyorgy Cserey |
|---|---|
| Categories | |
| ArXiv ID | 2601.08807vv1 |
| URL | https://arxiv.org/abs/2601.08807 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Tracklet quality is often treated as an afterthought in most person re-identification (ReID) methods, with the majority of research presenting architectural modifications to foundational models. Such approaches neglect an important limitation, posing challenges when deploying ReID systems in real-world, difficult scenarios. In this paper, we introduce S3-CLIP, a video super-resolution-based CLIP-ReID framework developed for the VReID-XFD challenge at WACV 2026. The proposed method integrates recent advances in super-resolution networks with task-driven super-resolution pipelines, adapting them to the video-based person re-identification setting. To the best of our knowledge, this work represents the first systematic investigation of video super-resolution as a means of enhancing tracklet quality for person ReID, particularly under challenging cross-view conditions. Experimental results demonstrate performance competitive with the baseline, achieving 37.52% mAP in aerial-to-ground and 29.16% mAP in ground-to-aerial scenarios. In the ground-to-aerial setting, S3-CLIP achieves substantial gains in ranking accuracy, improving Rank-1, Rank-5, and Rank-10 performance by 11.24%, 13.48%, and 17.98%, respectively.
{
"annotation_id": "de87e2e9-6043-4316-897e-ef84019d38cd",
"date_created": "2026-02-17T05:53:16.232000Z",
"date_modified": "2026-02-17T05:53:16.232000Z",
"file_hash": "11479f0c862ea6dcebd3a78831d78c2f03bc839202028fc67bd4fe8dbd226bc6",
"private": false,
"record": {
"abstract": "Tracklet quality is often treated as an afterthought in most person re-identification (ReID) methods, with the majority of research presenting architectural modifications to foundational models. Such approaches neglect an important limitation, posing challenges when deploying ReID systems in real-world, difficult scenarios. In this paper, we introduce S3-CLIP, a video super-resolution-based CLIP-ReID framework developed for the VReID-XFD challenge at WACV 2026. The proposed method integrates recent advances in super-resolution networks with task-driven super-resolution pipelines, adapting them to the video-based person re-identification setting. To the best of our knowledge, this work represents the first systematic investigation of video super-resolution as a means of enhancing tracklet quality for person ReID, particularly under challenging cross-view conditions. Experimental results demonstrate performance competitive with the baseline, achieving 37.52% mAP in aerial-to-ground and 29.16% mAP in ground-to-aerial scenarios. In the ground-to-aerial setting, S3-CLIP achieves substantial gains in ranking accuracy, improving Rank-1, Rank-5, and Rank-10 performance by 11.24%, 13.48%, and 17.98%, respectively.",
"arxiv_id": "2601.08807",
"authors": [
"Tamas Endrei",
"Gyorgy Cserey"
],
"categories": [
"cs.CV",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "S3-CLIP: Video Super Resolution for Person-ReID",
"url": "https://arxiv.org/abs/2601.08807",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "1e0613d6-13e4-4396-92e1-bc26cded6128",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}