dorsal/arxiv
View SchemaHallucination Detection and Mitigation in Large Language Models
| Authors | Ahmad Pesaranghader, Erin Li |
|---|---|
| Categories | |
| ArXiv ID | 2601.09929vv1 |
| URL | https://arxiv.org/abs/2601.09929 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Large Language Models (LLMs) and Large Reasoning Models (LRMs) offer transformative potential for high-stakes domains like finance and law, but their tendency to hallucinate, generating factually incorrect or unsupported content, poses a critical reliability risk. This paper introduces a comprehensive operational framework for hallucination management, built on a continuous improvement cycle driven by root cause awareness. We categorize hallucination sources into model, data, and context-related factors, allowing targeted interventions over generic fixes. The framework integrates multi-faceted detection methods (e.g., uncertainty estimation, reasoning consistency) with stratified mitigation strategies (e.g., knowledge grounding, confidence calibration). We demonstrate its application through a tiered architecture and a financial data extraction case study, where model, context, and data tiers form a closed feedback loop for progressive reliability enhancement. This approach provides a systematic, scalable methodology for building trustworthy generative AI systems in regulated environments.
{
"annotation_id": "e9cfc338-778e-4d59-af60-ea2e6215ba51",
"date_created": "2026-02-17T05:53:24.272000Z",
"date_modified": "2026-02-17T05:53:24.272000Z",
"file_hash": "30cc41c392c9da047802121fe4d1497336373c577fb61977773b06d2a7d1372e",
"private": false,
"record": {
"abstract": "Large Language Models (LLMs) and Large Reasoning Models (LRMs) offer transformative potential for high-stakes domains like finance and law, but their tendency to hallucinate, generating factually incorrect or unsupported content, poses a critical reliability risk. This paper introduces a comprehensive operational framework for hallucination management, built on a continuous improvement cycle driven by root cause awareness. We categorize hallucination sources into model, data, and context-related factors, allowing targeted interventions over generic fixes. The framework integrates multi-faceted detection methods (e.g., uncertainty estimation, reasoning consistency) with stratified mitigation strategies (e.g., knowledge grounding, confidence calibration). We demonstrate its application through a tiered architecture and a financial data extraction case study, where model, context, and data tiers form a closed feedback loop for progressive reliability enhancement. This approach provides a systematic, scalable methodology for building trustworthy generative AI systems in regulated environments.",
"arxiv_id": "2601.09929",
"authors": [
"Ahmad Pesaranghader",
"Erin Li"
],
"categories": [
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Hallucination Detection and Mitigation in Large Language Models",
"url": "https://arxiv.org/abs/2601.09929",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "8fe3d3ee-2f54-43cf-9efa-c51db1397c08",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}