dorsal/arxiv
View SchemaLost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
| Authors | Zheng Luo, T Pranav Kutralingam, Ogochukwu N Okoani, Wanpeng Xu, Hua Wei, Xiyang Hu |
|---|---|
| Categories | |
| ArXiv ID | 2601.05366vv1 |
| URL | https://arxiv.org/abs/2601.05366 |
| License | http://creativecommons.org/licenses/by-nc-sa/4.0/ |
Abstract
Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls. While recent work reports strong tool-calling performance under standard English-centric evaluations, the robustness of tool calling under multilingual user interactions remains underexplored. In this work, we introduce MLCL, a diagnostic benchmark, and conduct a systematic evaluation of multilingual tool calling across Chinese, Hindi, and the low-resource language Igbo. Through fine-grained error analysis, we show that many failures occur despite correct intent understanding and tool selection. We identify parameter value language mismatch as a dominant failure mode, where models generate semantically appropriate parameter values in the user's language, violating language-invariant execution conventions. We further evaluate several inference-time system strategies and find that while these strategies substantially reduce language-induced execution errors, none of them can fully recover English-level performance.
{
"annotation_id": "64b6afcd-14ac-47ec-a9ea-22a2fbabec7a",
"date_created": "2026-02-17T05:53:04.194000Z",
"date_modified": "2026-02-17T05:53:04.194000Z",
"file_hash": "006985681d8e7254ce27905a9c7ab15e4117d422daef36843cb1493253a5ad8d",
"private": false,
"record": {
"abstract": "Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls. While recent work reports strong tool-calling performance under standard English-centric evaluations, the robustness of tool calling under multilingual user interactions remains underexplored. In this work, we introduce MLCL, a diagnostic benchmark, and conduct a systematic evaluation of multilingual tool calling across Chinese, Hindi, and the low-resource language Igbo. Through fine-grained error analysis, we show that many failures occur despite correct intent understanding and tool selection. We identify parameter value language mismatch as a dominant failure mode, where models generate semantically appropriate parameter values in the user\u0027s language, violating language-invariant execution conventions. We further evaluate several inference-time system strategies and find that while these strategies substantially reduce language-induced execution errors, none of them can fully recover English-level performance.",
"arxiv_id": "2601.05366",
"authors": [
"Zheng Luo",
"T Pranav Kutralingam",
"Ogochukwu N Okoani",
"Wanpeng Xu",
"Hua Wei",
"Xiyang Hu"
],
"categories": [
"cs.CL",
"cs.AI",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by-nc-sa/4.0/",
"title": "Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models",
"url": "https://arxiv.org/abs/2601.05366",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "cf381147-d886-4632-9c00-d13ae68880eb",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}