dorsal/arxiv
View SchemaOver-Searching in Search-Augmented Large Language Models
| Authors | Roy Xie, Deepak Gopinath, David Qiu, Dong Lin, Haitian Sun, Saloni Potdar, Bhuwan Dhingra |
|---|---|
| Categories | |
| ArXiv ID | 2601.05503vv1 |
| URL | https://arxiv.org/abs/2601.05503 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Search-augmented large language models (LLMs) excel at knowledge-intensive tasks by integrating external retrieval. However, they often over-search -- unnecessarily invoking search tool even when it does not improve response quality, which leads to computational inefficiency and hallucinations by incorporating irrelevant context. In this work, we conduct a systematic evaluation of over-searching across multiple dimensions, including query types, model categories, retrieval conditions, and multi-turn conversations. Our finding shows: (i) search generally improves answer accuracy on answerable queries but harms abstention on unanswerable ones; (ii) over-searching is more pronounced in complex reasoning models and deep research systems, is exacerbated by noisy retrieval, and compounds across turns in multi-turn conversations; and (iii) the composition of retrieved evidence is crucial, as the presence of negative evidence improves abstention. To quantify over-searching, we introduce Tokens Per Correctness (TPC), an evaluation metric that captures the performance-cost trade-off for search-augmented LLMs. Lastly, we investigate mitigation approaches at both the query and retrieval levels and release the OverSearchQA to foster continued research into efficient search-augmented LLMs.
{
"annotation_id": "2b1cf256-59b6-4286-9319-28c07b96682b",
"date_created": "2026-02-17T05:53:04.466000Z",
"date_modified": "2026-02-17T05:53:04.466000Z",
"file_hash": "2d70bca26bbbd9f3f25e84e64d45ee66434aa3a3e5268977fb1e326a4c6c2cc4",
"private": false,
"record": {
"abstract": "Search-augmented large language models (LLMs) excel at knowledge-intensive tasks by integrating external retrieval. However, they often over-search -- unnecessarily invoking search tool even when it does not improve response quality, which leads to computational inefficiency and hallucinations by incorporating irrelevant context. In this work, we conduct a systematic evaluation of over-searching across multiple dimensions, including query types, model categories, retrieval conditions, and multi-turn conversations. Our finding shows: (i) search generally improves answer accuracy on answerable queries but harms abstention on unanswerable ones; (ii) over-searching is more pronounced in complex reasoning models and deep research systems, is exacerbated by noisy retrieval, and compounds across turns in multi-turn conversations; and (iii) the composition of retrieved evidence is crucial, as the presence of negative evidence improves abstention. To quantify over-searching, we introduce Tokens Per Correctness (TPC), an evaluation metric that captures the performance-cost trade-off for search-augmented LLMs. Lastly, we investigate mitigation approaches at both the query and retrieval levels and release the OverSearchQA to foster continued research into efficient search-augmented LLMs.",
"arxiv_id": "2601.05503",
"authors": [
"Roy Xie",
"Deepak Gopinath",
"David Qiu",
"Dong Lin",
"Haitian Sun",
"Saloni Potdar",
"Bhuwan Dhingra"
],
"categories": [
"cs.LG",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Over-Searching in Search-Augmented Large Language Models",
"url": "https://arxiv.org/abs/2601.05503",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "072d11c9-6d11-4c99-81e0-6e7ede829229",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}