dorsal/arxiv
View SchemaTwo-step Authentication: Multi-biometric System Using Voice and Facial Recognition
| Authors | Kuan Wei Chen, Ting Yi Lin, Wen Ren Yang, Aryan Kesarwani, Riya Singh |
|---|---|
| Categories | |
| ArXiv ID | 2601.06218vv1 |
| URL | https://arxiv.org/abs/2601.06218 |
| DOI | 10.1049/icp.2024.4141 |
| Journal | IET Conference Proceedings 2024 (22) 11-12 (2025) |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
We present a cost-effective two-step authentication system that integrates face identification and speaker verification using only a camera and microphone available on common devices. The pipeline first performs face recognition to identify a candidate user from a small enrolled group, then performs voice recognition only against the matched identity to reduce computation and improve robustness. For face recognition, a pruned VGG-16 based classifier is trained on an augmented dataset of 924 images from five subjects, with faces localized by MTCNN; it achieves 95.1% accuracy. For voice recognition, a CNN speaker-verification model trained on LibriSpeech (train-other-360) attains 98.9% accuracy and 3.456% EER on test-clean. Source code and trained models are available at https://github.com/NCUE-EE-AIAL/Two-step-Authentication-Multi-biometric-System.
{
"annotation_id": "66c2f856-135a-445e-872f-4c939726c741",
"date_created": "2026-02-17T05:53:07.509000Z",
"date_modified": "2026-02-17T05:53:07.509000Z",
"file_hash": "7a2d7f35546b03def0fe1ccc98e61c3418407f96ead235b34929b8a0909f1805",
"private": false,
"record": {
"abstract": "We present a cost-effective two-step authentication system that integrates face identification and speaker verification using only a camera and microphone available on common devices. The pipeline first performs face recognition to identify a candidate user from a small enrolled group, then performs voice recognition only against the matched identity to reduce computation and improve robustness. For face recognition, a pruned VGG-16 based classifier is trained on an augmented dataset of 924 images from five subjects, with faces localized by MTCNN; it achieves 95.1% accuracy. For voice recognition, a CNN speaker-verification model trained on LibriSpeech (train-other-360) attains 98.9% accuracy and 3.456% EER on test-clean. Source code and trained models are available at https://github.com/NCUE-EE-AIAL/Two-step-Authentication-Multi-biometric-System.",
"arxiv_id": "2601.06218",
"authors": [
"Kuan Wei Chen",
"Ting Yi Lin",
"Wen Ren Yang",
"Aryan Kesarwani",
"Riya Singh"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.MM"
],
"doi": "10.1049/icp.2024.4141",
"journal_ref": "IET Conference Proceedings 2024 (22) 11-12 (2025)",
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Two-step Authentication: Multi-biometric System Using Voice and Facial Recognition",
"url": "https://arxiv.org/abs/2601.06218",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "9fc516d1-690b-46dd-8259-b44840f2dd7d",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}