dorsal/arxiv
View SchemaDiffusion-based Frameworks for Unsupervised Speech Enhancement
| Authors | Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel, Xavier Alameda-Pineda |
|---|---|
| Categories | |
| ArXiv ID | 2601.09931vv1 |
| URL | https://arxiv.org/abs/2601.09931 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
This paper addresses $\textit{unsupervised}$ diffusion-based single-channel speech enhancement (SE). Prior work in this direction combines a score-based diffusion model trained on clean speech with a Gaussian noise model whose covariance is structured by non-negative matrix factorization (NMF). This combination is used within an iterative expectation-maximization (EM) scheme, in which a diffusion-based posterior-sampling E-step estimates the clean speech. We first revisit this framework and propose to explicitly model both speech and acoustic noise as latent variables, jointly sampling them in the E-step instead of sampling speech alone as in previous approaches. We then introduce a new unsupervised SE framework that replaces the NMF noise prior with a diffusion-based noise model, learned jointly with the speech prior in a single conditional score model. Within this framework, we derive two variants: one that implicitly accounts for noise and one that explicitly treats noise as a latent variable. Experiments on WSJ0-QUT and VoiceBank-DEMAND show that explicit noise modeling systematically improves SE performance for both NMF-based and diffusion-based noise priors. Under matched conditions, the diffusion-based noise model attains the best overall quality and intelligibility among unsupervised methods, while under mismatched conditions the proposed NMF-based explicit-noise framework is more robust and suffers less degradation than several supervised baselines. Our code will be publicly available on this $\href{https://github.com/jeaneudesAyilo/enudiffuse}{URL}$.
{
"annotation_id": "94507c3e-2d14-4f8c-9f5c-bf2f0085c619",
"date_created": "2026-02-17T05:53:23.069000Z",
"date_modified": "2026-02-17T05:53:23.069000Z",
"file_hash": "5bfc31ae822f9bdf987a4b04b0a6dac35f395f0535b8bb0c36b7b0ba04428142",
"private": false,
"record": {
"abstract": "This paper addresses $\\textit{unsupervised}$ diffusion-based single-channel speech enhancement (SE). Prior work in this direction combines a score-based diffusion model trained on clean speech with a Gaussian noise model whose covariance is structured by non-negative matrix factorization (NMF). This combination is used within an iterative expectation-maximization (EM) scheme, in which a diffusion-based posterior-sampling E-step estimates the clean speech. We first revisit this framework and propose to explicitly model both speech and acoustic noise as latent variables, jointly sampling them in the E-step instead of sampling speech alone as in previous approaches. We then introduce a new unsupervised SE framework that replaces the NMF noise prior with a diffusion-based noise model, learned jointly with the speech prior in a single conditional score model. Within this framework, we derive two variants: one that implicitly accounts for noise and one that explicitly treats noise as a latent variable. Experiments on WSJ0-QUT and VoiceBank-DEMAND show that explicit noise modeling systematically improves SE performance for both NMF-based and diffusion-based noise priors. Under matched conditions, the diffusion-based noise model attains the best overall quality and intelligibility among unsupervised methods, while under mismatched conditions the proposed NMF-based explicit-noise framework is more robust and suffers less degradation than several supervised baselines. Our code will be publicly available on this $\\href{https://github.com/jeaneudesAyilo/enudiffuse}{URL}$.",
"arxiv_id": "2601.09931",
"authors": [
"Jean-Eudes Ayilo",
"Mostafa Sadeghi",
"Romain Serizel",
"Xavier Alameda-Pineda"
],
"categories": [
"cs.SD"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Diffusion-based Frameworks for Unsupervised Speech Enhancement",
"url": "https://arxiv.org/abs/2601.09931",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "45fdfe5a-62dc-4089-9562-3d32fb3fac3b",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}