dorsal/arxiv
View SchemaToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
| Authors | Yutao Mou, Zhangchi Xue, Lijun Li, Peiyang Liu, Shikun Zhang, Wei Ye, Jing Shao |
|---|---|
| Categories | |
| ArXiv ID | 2601.10156vv1 |
| URL | https://arxiv.org/abs/2601.10156 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
While LLM-based agents can interact with environments via invoking external tools, their expanded capabilities also amplify security risks. Monitoring step-level tool invocation behaviors in real time and proactively intervening before unsafe execution is critical for agent deployment, yet remains under-explored. In this work, we first construct TS-Bench, a novel benchmark for step-level tool invocation safety detection in LLM agents. We then develop a guardrail model, TS-Guard, using multi-task reinforcement learning. The model proactively detects unsafe tool invocation actions before execution by reasoning over the interaction history. It assesses request harmfulness and action-attack correlations, producing interpretable and generalizable safety judgments and feedback. Furthermore, we introduce TS-Flow, a guardrail-feedback-driven reasoning framework for LLM agents, which reduces harmful tool invocations of ReAct-style agents by 65 percent on average and improves benign task completion by approximately 10 percent under prompt injection attacks.
{
"annotation_id": "a26a9226-45f0-4e23-98b3-6dbc1fd29e73",
"date_created": "2026-02-17T05:53:23.790000Z",
"date_modified": "2026-02-17T05:53:23.790000Z",
"file_hash": "c00c5c97032514d6177ff5bc471bb3b2a9c1db4ac9af2fb43bbcf424e24ba7d8",
"private": false,
"record": {
"abstract": "While LLM-based agents can interact with environments via invoking external tools, their expanded capabilities also amplify security risks. Monitoring step-level tool invocation behaviors in real time and proactively intervening before unsafe execution is critical for agent deployment, yet remains under-explored. In this work, we first construct TS-Bench, a novel benchmark for step-level tool invocation safety detection in LLM agents. We then develop a guardrail model, TS-Guard, using multi-task reinforcement learning. The model proactively detects unsafe tool invocation actions before execution by reasoning over the interaction history. It assesses request harmfulness and action-attack correlations, producing interpretable and generalizable safety judgments and feedback. Furthermore, we introduce TS-Flow, a guardrail-feedback-driven reasoning framework for LLM agents, which reduces harmful tool invocations of ReAct-style agents by 65 percent on average and improves benign task completion by approximately 10 percent under prompt injection attacks.",
"arxiv_id": "2601.10156",
"authors": [
"Yutao Mou",
"Zhangchi Xue",
"Lijun Li",
"Peiyang Liu",
"Shikun Zhang",
"Wei Ye",
"Jing Shao"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback",
"url": "https://arxiv.org/abs/2601.10156",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "b794329a-4cc6-496a-ac9d-d06188c24535",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}