dorsal/arxiv
View SchemaBias in the Shadows: Explore Shortcuts in Encrypted Network Traffic Classification
| Authors | Chuyi Wang, Xiaohui Xie, Tongze Wang, Yong Cui |
|---|---|
| Categories | |
| ArXiv ID | 2601.10180vv1 |
| URL | https://arxiv.org/abs/2601.10180 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Pre-trained models operating directly on raw bytes have achieved promising performance in encrypted network traffic classification (NTC), but often suffer from shortcut learning-relying on spurious correlations that fail to generalize to real-world data. Existing solutions heavily rely on model-specific interpretation techniques, which lack adaptability and generality across different model architectures and deployment scenarios. In this paper, we propose BiasSeeker, the first semi-automated framework that is both model-agnostic and data-driven for detecting dataset-specific shortcut features in encrypted traffic. By performing statistical correlation analysis directly on raw binary traffic, BiasSeeker identifies spurious or environment-entangled features that may compromise generalization, independent of any classifier. To address the diverse nature of shortcut features, we introduce a systematic categorization and apply category-specific validation strategies that reduce bias while preserving meaningful information. We evaluate BiasSeeker on 19 public datasets across three NTC tasks. By emphasizing context-aware feature selection and dataset-specific diagnosis, BiasSeeker offers a novel perspective for understanding and addressing shortcut learning in encrypted network traffic classification, raising awareness that feature selection should be an intentional and scenario-sensitive step prior to model training.
{
"annotation_id": "bf271609-56a8-42a4-a2bd-c5f2f713e86f",
"date_created": "2026-02-17T05:53:23.802000Z",
"date_modified": "2026-02-17T05:53:23.802000Z",
"file_hash": "d6126b607d893e0fe6fd5e36c26de7aada4fea303ef17727185a9bca9fc8ea31",
"private": false,
"record": {
"abstract": "Pre-trained models operating directly on raw bytes have achieved promising performance in encrypted network traffic classification (NTC), but often suffer from shortcut learning-relying on spurious correlations that fail to generalize to real-world data. Existing solutions heavily rely on model-specific interpretation techniques, which lack adaptability and generality across different model architectures and deployment scenarios.\n In this paper, we propose BiasSeeker, the first semi-automated framework that is both model-agnostic and data-driven for detecting dataset-specific shortcut features in encrypted traffic. By performing statistical correlation analysis directly on raw binary traffic, BiasSeeker identifies spurious or environment-entangled features that may compromise generalization, independent of any classifier. To address the diverse nature of shortcut features, we introduce a systematic categorization and apply category-specific validation strategies that reduce bias while preserving meaningful information.\n We evaluate BiasSeeker on 19 public datasets across three NTC tasks. By emphasizing context-aware feature selection and dataset-specific diagnosis, BiasSeeker offers a novel perspective for understanding and addressing shortcut learning in encrypted network traffic classification, raising awareness that feature selection should be an intentional and scenario-sensitive step prior to model training.",
"arxiv_id": "2601.10180",
"authors": [
"Chuyi Wang",
"Xiaohui Xie",
"Tongze Wang",
"Yong Cui"
],
"categories": [
"cs.LG",
"cs.NI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Bias in the Shadows: Explore Shortcuts in Encrypted Network Traffic Classification",
"url": "https://arxiv.org/abs/2601.10180",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "96b9c50d-54ac-44a9-b2eb-6f4e9ccefbe1",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}