dorsal/arxiv
View SchemaTaxoBell: Gaussian Box Embeddings for Self-Supervised Taxonomy Expansion
| Authors | Sahil Mishra, Srinitish Srinivasan, Srikanta Bedathur, Tanmoy Chakraborty |
|---|---|
| Categories | |
| ArXiv ID | 2601.09633vv1 |
| URL | https://arxiv.org/abs/2601.09633 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Taxonomies form the backbone of structured knowledge representation across diverse domains, enabling applications such as e-commerce catalogs, semantic search, and biomedical discovery. Yet, manual taxonomy expansion is labor-intensive and cannot keep pace with the emergence of new concepts. Existing automated methods rely on point-based vector embeddings, which model symmetric similarity and thus struggle with the asymmetric "is-a" relationships that are fundamental to taxonomies. Box embeddings offer a promising alternative by enabling containment and disjointness, but they face key issues: (i) unstable gradients at the intersection boundaries, (ii) no notion of semantic uncertainty, and (iii) limited capacity to represent polysemy or ambiguity. We address these shortcomings with TaxoBell, a Gaussian box embedding framework that translates between box geometries and multivariate Gaussian distributions, where means encode semantic location and covariances encode uncertainty. Energy-based optimization yields stable optimization, robust modeling of ambiguous concepts, and interpretable hierarchical reasoning. Extensive experimentation on five benchmark datasets demonstrates that TaxoBell significantly outperforms eight state-of-the-art taxonomy expansion baselines by 19% in MRR and around 25% in Recall@k. We further demonstrate the advantages and pitfalls of TaxoBell with error analysis and ablation studies.
{
"annotation_id": "327d87d8-7429-4bee-a95f-25d438de365c",
"date_created": "2026-02-17T05:53:20.277000Z",
"date_modified": "2026-02-17T05:53:20.277000Z",
"file_hash": "3258c7ca60d944990f46dc56dc961f1d5d5f7d64b8dbffe4d341812fe35fc215",
"private": false,
"record": {
"abstract": "Taxonomies form the backbone of structured knowledge representation across diverse domains, enabling applications such as e-commerce catalogs, semantic search, and biomedical discovery. Yet, manual taxonomy expansion is labor-intensive and cannot keep pace with the emergence of new concepts. Existing automated methods rely on point-based vector embeddings, which model symmetric similarity and thus struggle with the asymmetric \"is-a\" relationships that are fundamental to taxonomies. Box embeddings offer a promising alternative by enabling containment and disjointness, but they face key issues: (i) unstable gradients at the intersection boundaries, (ii) no notion of semantic uncertainty, and (iii) limited capacity to represent polysemy or ambiguity. We address these shortcomings with TaxoBell, a Gaussian box embedding framework that translates between box geometries and multivariate Gaussian distributions, where means encode semantic location and covariances encode uncertainty. Energy-based optimization yields stable optimization, robust modeling of ambiguous concepts, and interpretable hierarchical reasoning. Extensive experimentation on five benchmark datasets demonstrates that TaxoBell significantly outperforms eight state-of-the-art taxonomy expansion baselines by 19% in MRR and around 25% in Recall@k. We further demonstrate the advantages and pitfalls of TaxoBell with error analysis and ablation studies.",
"arxiv_id": "2601.09633",
"authors": [
"Sahil Mishra",
"Srinitish Srinivasan",
"Srikanta Bedathur",
"Tanmoy Chakraborty"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "TaxoBell: Gaussian Box Embeddings for Self-Supervised Taxonomy Expansion",
"url": "https://arxiv.org/abs/2601.09633",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "6728990c-cfb9-486e-98cd-4dad2b6d4a66",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}