dorsal/arxiv
View SchemaTowards Comprehensive Semantic Speech Embeddings for Chinese Dialects
| Authors | Kalvin Chang, Yiwen Shao, Jiahong Li, Dong Yu |
|---|---|
| Categories | |
| ArXiv ID | 2601.07274vv1 |
| URL | https://arxiv.org/abs/2601.07274 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Despite having hundreds of millions of speakers, Chinese dialects lag behind Mandarin in speech and language technologies. Most varieties are primarily spoken, making dialect-to-Mandarin speech-LLMs (large language models) more practical than dialect LLMs. Building dialect-to-Mandarin speech-LLMs requires speech representations with cross-dialect semantic alignment between Chinese dialects and Mandarin. In this paper, we achieve such a cross-dialect semantic alignment by training a speech encoder with ASR (automatic speech recognition)-only data, as demonstrated by speech-to-speech retrieval on a new benchmark of spoken Chinese varieties that we contribute. Our speech encoder further demonstrates state-of-the-art ASR performance on Chinese dialects. Together, our Chinese dialect benchmark, semantically aligned speech representations, and speech-to-speech retrieval evaluation lay the groundwork for future Chinese dialect speech-LLMs. We release the benchmark at https://github.com/kalvinchang/yubao.
{
"annotation_id": "57c3d79e-a47b-4d55-a0c3-af0edacdf3a2",
"date_created": "2026-02-17T05:53:11.345000Z",
"date_modified": "2026-02-17T05:53:11.345000Z",
"file_hash": "0218e47115867ac986a491d73f03c1aaf7a3312a95b706eb09fc0d3c341a5a27",
"private": false,
"record": {
"abstract": "Despite having hundreds of millions of speakers, Chinese dialects lag behind Mandarin in speech and language technologies. Most varieties are primarily spoken, making dialect-to-Mandarin speech-LLMs (large language models) more practical than dialect LLMs. Building dialect-to-Mandarin speech-LLMs requires speech representations with cross-dialect semantic alignment between Chinese dialects and Mandarin. In this paper, we achieve such a cross-dialect semantic alignment by training a speech encoder with ASR (automatic speech recognition)-only data, as demonstrated by speech-to-speech retrieval on a new benchmark of spoken Chinese varieties that we contribute. Our speech encoder further demonstrates state-of-the-art ASR performance on Chinese dialects. Together, our Chinese dialect benchmark, semantically aligned speech representations, and speech-to-speech retrieval evaluation lay the groundwork for future Chinese dialect speech-LLMs. We release the benchmark at https://github.com/kalvinchang/yubao.",
"arxiv_id": "2601.07274",
"authors": [
"Kalvin Chang",
"Yiwen Shao",
"Jiahong Li",
"Dong Yu"
],
"categories": [
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects",
"url": "https://arxiv.org/abs/2601.07274",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "c50aef3c-2ba4-44a0-b48b-5f50edd700f5",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}