dorsal/arxiv
View SchemaDistribution Estimation with Side Information
| Authors | Haricharan Balasundaram, Andrew Thangaraj |
|---|---|
| Categories | |
| ArXiv ID | 2601.08535vv1 |
| URL | https://arxiv.org/abs/2601.08535 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
We consider the classical problem of discrete distribution estimation using i.i.d. samples in a novel scenario where additional side information is available on the distribution. In large alphabet datasets such as text corpora, such side information arises naturally through word semantics/similarities that can be inferred by closeness of vector word embeddings, for instance. We consider two specific models for side information--a local model where the unknown distribution is in the neighborhood of a known distribution, and a partial ordering model where the alphabet is partitioned into known higher and lower probability sets. In both models, we theoretically characterize the improvement in a suitable squared-error risk because of the available side information. Simulations over natural language and synthetic data illustrate these gains.
{
"annotation_id": "7ad94f7a-a18e-4a97-8340-ccd1eeab61c6",
"date_created": "2026-02-17T05:53:16.022000Z",
"date_modified": "2026-02-17T05:53:16.022000Z",
"file_hash": "126862444eb68275c01f6874d21dabc41194390640f93210833990f022c4bab2",
"private": false,
"record": {
"abstract": "We consider the classical problem of discrete distribution estimation using i.i.d. samples in a novel scenario where additional side information is available on the distribution. In large alphabet datasets such as text corpora, such side information arises naturally through word semantics/similarities that can be inferred by closeness of vector word embeddings, for instance. We consider two specific models for side information--a local model where the unknown distribution is in the neighborhood of a known distribution, and a partial ordering model where the alphabet is partitioned into known higher and lower probability sets. In both models, we theoretically characterize the improvement in a suitable squared-error risk because of the available side information. Simulations over natural language and synthetic data illustrate these gains.",
"arxiv_id": "2601.08535",
"authors": [
"Haricharan Balasundaram",
"Andrew Thangaraj"
],
"categories": [
"cs.IT",
"math.IT"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Distribution Estimation with Side Information",
"url": "https://arxiv.org/abs/2601.08535",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "613787a0-3e57-46f2-94ab-559021a44dec",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}