dorsal/arxiv
View SchemaCredit C-GPT: A Domain-Specialized Large Language Model for Conversational Understanding in Vietnamese Debt Collection
| Authors | Nhung Nguyen Thi Hong, Cuong Nguyen Dang, Tri Le Ngoc |
|---|---|
| Categories | |
| ArXiv ID | 2601.10167vv1 |
| URL | https://arxiv.org/abs/2601.10167 |
| License | http://creativecommons.org/licenses/by-nc-nd/4.0/ |
Abstract
Debt collection is a critical function within the banking, financial services, and insurance (BFSI) sector, relying heavily on large-scale human-to-human conversational interactions conducted primarily in Vietnamese contact centers. These conversations involve informal spoken language, emotional variability, and complex domain-specific reasoning, which pose significant challenges for traditional natural language processing systems. This paper introduces Credit C-GPT, a domain-specialized large language model with seven billion parameters, fine-tuned for conversational understanding in Vietnamese debt collection scenarios. The proposed model integrates multiple conversational intelligence tasks, including dialogue understanding, sentiment recognition, intent detection, call stage classification, and structured slot-value extraction, within a single reasoning-based framework. We describe the data construction process, annotation strategy, and training methodology, and evaluate the model on proprietary human-annotated datasets. Experimental results show consistent improvements over traditional pipeline-based approaches, indicating that domain-specialized conversational language models provide a scalable and privacy-aware solution for real-time assistance and post-call analytics in enterprise contact centers.
{
"annotation_id": "b4e7c69e-2180-4376-a0a0-48df8eec79e1",
"date_created": "2026-02-17T05:53:24.389000Z",
"date_modified": "2026-02-17T05:53:24.389000Z",
"file_hash": "419bb1a6e6a9fb838baa13ad3323a2ac6c252d070ed19477d5f263cc690d33c9",
"private": false,
"record": {
"abstract": "Debt collection is a critical function within the banking, financial services, and insurance (BFSI) sector, relying heavily on large-scale human-to-human conversational interactions conducted primarily in Vietnamese contact centers. These conversations involve informal spoken language, emotional variability, and complex domain-specific reasoning, which pose significant challenges for traditional natural language processing systems. This paper introduces Credit C-GPT, a domain-specialized large language model with seven billion parameters, fine-tuned for conversational understanding in Vietnamese debt collection scenarios. The proposed model integrates multiple conversational intelligence tasks, including dialogue understanding, sentiment recognition, intent detection, call stage classification, and structured slot-value extraction, within a single reasoning-based framework. We describe the data construction process, annotation strategy, and training methodology, and evaluate the model on proprietary human-annotated datasets. Experimental results show consistent improvements over traditional pipeline-based approaches, indicating that domain-specialized conversational language models provide a scalable and privacy-aware solution for real-time assistance and post-call analytics in enterprise contact centers.",
"arxiv_id": "2601.10167",
"authors": [
"Nhung Nguyen Thi Hong",
"Cuong Nguyen Dang",
"Tri Le Ngoc"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by-nc-nd/4.0/",
"title": "Credit C-GPT: A Domain-Specialized Large Language Model for Conversational Understanding in Vietnamese Debt Collection",
"url": "https://arxiv.org/abs/2601.10167",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "eb2321f6-e6a3-4983-b74d-a3cdd3e96144",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}