dorsal/arxiv
View SchemaBeyond Static Tools: Test-Time Tool Evolution for Scientific Reasoning
| Authors | Jiaxuan Lu, Ziyu Kong, Yemin Wang, Rong Fu, Haiyuan Wan, Cheng Yang, Wenjie Lou, Haoran Sun, Lilong Wang, Yankai Jiang, Xiaosong Wang, Xiao Sun, Dongzhan Zhou |
|---|---|
| Categories | |
| ArXiv ID | 2601.07641vv1 |
| URL | https://arxiv.org/abs/2601.07641 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
The central challenge of AI for Science is not reasoning alone, but the ability to create computational methods in an open-ended scientific world. Existing LLM-based agents rely on static, pre-defined tool libraries, a paradigm that fundamentally fails in scientific domains where tools are sparse, heterogeneous, and intrinsically incomplete. In this paper, we propose Test-Time Tool Evolution (TTE), a new paradigm that enables agents to synthesize, verify, and evolve executable tools during inference. By transforming tools from fixed resources into problem-driven artifacts, TTE overcomes the rigidity and long-tail limitations of static tool libraries. To facilitate rigorous evaluation, we introduce SciEvo, a benchmark comprising 1,590 scientific reasoning tasks supported by 925 automatically evolved tools. Extensive experiments show that TTE achieves state-of-the-art performance in both accuracy and tool efficiency, while enabling effective cross-domain adaptation of computational tools. The code and benchmark have been released at https://github.com/lujiaxuan0520/Test-Time-Tool-Evol.
{
"annotation_id": "ac56c28a-804f-482d-9fcc-7af3baa8a7e9",
"date_created": "2026-02-17T05:53:11.626000Z",
"date_modified": "2026-02-17T05:53:11.626000Z",
"file_hash": "094fee34265cf52cd50934f2d0bab05893af3d21a3c7f011c4cbec74a51ffbf5",
"private": false,
"record": {
"abstract": "The central challenge of AI for Science is not reasoning alone, but the ability to create computational methods in an open-ended scientific world. Existing LLM-based agents rely on static, pre-defined tool libraries, a paradigm that fundamentally fails in scientific domains where tools are sparse, heterogeneous, and intrinsically incomplete. In this paper, we propose Test-Time Tool Evolution (TTE), a new paradigm that enables agents to synthesize, verify, and evolve executable tools during inference. By transforming tools from fixed resources into problem-driven artifacts, TTE overcomes the rigidity and long-tail limitations of static tool libraries. To facilitate rigorous evaluation, we introduce SciEvo, a benchmark comprising 1,590 scientific reasoning tasks supported by 925 automatically evolved tools. Extensive experiments show that TTE achieves state-of-the-art performance in both accuracy and tool efficiency, while enabling effective cross-domain adaptation of computational tools. The code and benchmark have been released at https://github.com/lujiaxuan0520/Test-Time-Tool-Evol.",
"arxiv_id": "2601.07641",
"authors": [
"Jiaxuan Lu",
"Ziyu Kong",
"Yemin Wang",
"Rong Fu",
"Haiyuan Wan",
"Cheng Yang",
"Wenjie Lou",
"Haoran Sun",
"Lilong Wang",
"Yankai Jiang",
"Xiaosong Wang",
"Xiao Sun",
"Dongzhan Zhou"
],
"categories": [
"cs.AI",
"cs.CL",
"cs.MA"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Beyond Static Tools: Test-Time Tool Evolution for Scientific Reasoning",
"url": "https://arxiv.org/abs/2601.07641",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "4008db5a-cf59-4b1e-b471-ac077e4bee99",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}