dorsal/arxiv
View SchemaIllusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency
| Authors | Haoming Xu, Ningyuan Zhao, Yunzhi Yao, Weihong Xu, Hongru Wang, Xinle Deng, Shumin Deng, Jeff Z. Pan, Huajun Chen, Ningyu Zhang |
|---|---|
| Categories | |
| ArXiv ID | 2601.05905vv1 |
| URL | https://arxiv.org/abs/2601.05905 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
As Large Language Models (LLMs) are increasingly deployed in real-world settings, correctness alone is insufficient. Reliable deployment requires maintaining truthful beliefs under contextual perturbations. Existing evaluations largely rely on point-wise confidence like Self-Consistency, which can mask brittle belief. We show that even facts answered with perfect self-consistency can rapidly collapse under mild contextual interference. To address this gap, we propose Neighbor-Consistency Belief (NCB), a structural measure of belief robustness that evaluates response coherence across a conceptual neighborhood. To validate the efficiency of NCB, we introduce a new cognitive stress-testing protocol that probes outputs stability under contextual interference. Experiments across multiple LLMs show that the performance of high-NCB data is relatively more resistant to interference. Finally, we present Structure-Aware Training (SAT), which optimizes context-invariant belief structure and reduces long-tail knowledge brittleness by approximately 30%. Code will be available at https://github.com/zjunlp/belief.
{
"annotation_id": "6dcfd290-b30f-4d0c-be36-777cd1d120e0",
"date_created": "2026-02-17T05:53:05.107000Z",
"date_modified": "2026-02-17T05:53:05.107000Z",
"file_hash": "2aea45d227d934aeb1a1eba7140c877aa6488e4b2ed4b4e024ffeb4a8356f735",
"private": false,
"record": {
"abstract": "As Large Language Models (LLMs) are increasingly deployed in real-world settings, correctness alone is insufficient. Reliable deployment requires maintaining truthful beliefs under contextual perturbations. Existing evaluations largely rely on point-wise confidence like Self-Consistency, which can mask brittle belief. We show that even facts answered with perfect self-consistency can rapidly collapse under mild contextual interference. To address this gap, we propose Neighbor-Consistency Belief (NCB), a structural measure of belief robustness that evaluates response coherence across a conceptual neighborhood. To validate the efficiency of NCB, we introduce a new cognitive stress-testing protocol that probes outputs stability under contextual interference. Experiments across multiple LLMs show that the performance of high-NCB data is relatively more resistant to interference. Finally, we present Structure-Aware Training (SAT), which optimizes context-invariant belief structure and reduces long-tail knowledge brittleness by approximately 30%. Code will be available at https://github.com/zjunlp/belief.",
"arxiv_id": "2601.05905",
"authors": [
"Haoming Xu",
"Ningyuan Zhao",
"Yunzhi Yao",
"Weihong Xu",
"Hongru Wang",
"Xinle Deng",
"Shumin Deng",
"Jeff Z. Pan",
"Huajun Chen",
"Ningyu Zhang"
],
"categories": [
"cs.CL",
"cs.AI",
"cs.HC",
"cs.LG",
"cs.MA"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency",
"url": "https://arxiv.org/abs/2601.05905",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "c5cfa715-cfeb-4a74-b573-d0b46cc23091",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}