dorsal/arxiv
View SchemaAnalysis of the Maximum Prediction Gain of Short-Term Prediction on Sustained Speech
| Authors | Reemt Hinrichs, Muhamad Fadli Damara, Stephan Preihs, Jörn Ostermann |
|---|---|
| Categories | |
| ArXiv ID | 2601.09461vv1 |
| URL | https://arxiv.org/abs/2601.09461 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Signal prediction is widely used in, e.g., economic forecasting, echo cancellation and in data compression, particularly in predictive coding of speech and music. Predictive coding algorithms reduce the bit-rate required for data transmission or storage by signal prediction. The prediction gain is a classic measure in applied signal coding of the quality of a predictor, as it links the mean-squared prediction error to the signal-to-quantization-noise of predictive coders. To evaluate predictor models, knowledge about the maximum achievable prediction gain independent of a predictor model is desirable. In this manuscript, Nadaraya-Watson kernel-regression (NWKR) and an information theoretic upper bound are applied to analyze the upper bound of the prediction gain on a newly recorded dataset of sustained speech/phonemes. It was found that for unvoiced speech a linear predictor always achieves the maximum prediction gain within at most 0.3 dB. On voiced speech, the optimum one-tap predictor was found to be linear but starting with two taps, the maximum achievable prediction gain was found to be about 2 dB to 6 dB above the prediction gain of the linear predictor. Significant differences between speakers/subjects were observed. The created dataset as well as the code can be obtained for research purpose upon request.
{
"annotation_id": "b63dc030-7d9a-4ab4-b075-da454f460988",
"date_created": "2026-02-17T05:53:20.417000Z",
"date_modified": "2026-02-17T05:53:20.417000Z",
"file_hash": "0c7ba5ba07aad1d30f77281fae6d94b3c2a97100b8a37099f5ff45eac3dd56bb",
"private": false,
"record": {
"abstract": "Signal prediction is widely used in, e.g., economic forecasting, echo cancellation and in data compression, particularly in predictive coding of speech and music. Predictive coding algorithms reduce the bit-rate required for data transmission or storage by signal prediction. The prediction gain is a classic measure in applied signal coding of the quality of a predictor, as it links the mean-squared prediction error to the signal-to-quantization-noise of predictive coders. To evaluate predictor models, knowledge about the maximum achievable prediction gain independent of a predictor model is desirable. In this manuscript, Nadaraya-Watson kernel-regression (NWKR) and an information theoretic upper bound are applied to analyze the upper bound of the prediction gain on a newly recorded dataset of sustained speech/phonemes. It was found that for unvoiced speech a linear predictor always achieves the maximum prediction gain within at most 0.3 dB. On voiced speech, the optimum one-tap predictor was found to be linear but starting with two taps, the maximum achievable prediction gain was found to be about 2 dB to 6 dB above the prediction gain of the linear predictor. Significant differences between speakers/subjects were observed.\n The created dataset as well as the code can be obtained for research purpose upon request.",
"arxiv_id": "2601.09461",
"authors": [
"Reemt Hinrichs",
"Muhamad Fadli Damara",
"Stephan Preihs",
"J\u00f6rn Ostermann"
],
"categories": [
"cs.SD"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Analysis of the Maximum Prediction Gain of Short-Term Prediction on Sustained Speech",
"url": "https://arxiv.org/abs/2601.09461",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "fe00a8da-a9bb-4355-8cdc-1df7564d14c2",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}