dorsal/arxiv
View SchemaLLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
| Authors | Hao Li, Yiqun Zhang, Zhaoyan Guo, Chenxu Wang, Shengji Tang, Qiaosheng Zhang, Yang Chen, Biqing Qi, Peng Ye, Lei Bai, Zhen Wang, Shuyue Hu |
|---|---|
| Categories | |
| ArXiv ID | 2601.07206vv1 |
| URL | https://arxiv.org/abs/2601.07206 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Large language model (LLM) routing assigns each query to the most suitable model from an ensemble. We introduce LLMRouterBench, a large-scale benchmark and unified framework for LLM routing. It comprises over 400K instances from 21 datasets and 33 models. Moreover, it provides comprehensive metrics for both performance-oriented routing and performance-cost trade-off routing, and integrates 10 representative routing baselines. Using LLMRouterBench, we systematically re-evaluate the field. While confirming strong model complementarity-the central premise of LLM routing-we find that many routing methods exhibit similar performance under unified evaluation, and several recent approaches, including commercial routers, fail to reliably outperform a simple baseline. Meanwhile, a substantial gap remains to the Oracle, driven primarily by persistent model-recall failures. We further show that backbone embedding models have limited impact, that larger ensembles exhibit diminishing returns compared to careful model curation, and that the benchmark also enables latency-aware analysis. All code and data are available at https://github.com/ynulihao/LLMRouterBench.
{
"annotation_id": "68e9deff-9554-4a69-8ab1-72f5e7a03428",
"date_created": "2026-02-17T05:53:11.757000Z",
"date_modified": "2026-02-17T05:53:11.757000Z",
"file_hash": "c18dd0d6289e9edce33aeafc9f45f9c38c72cc120ed47d4d89054a4e74cfcaff",
"private": false,
"record": {
"abstract": "Large language model (LLM) routing assigns each query to the most suitable model from an ensemble. We introduce LLMRouterBench, a large-scale benchmark and unified framework for LLM routing. It comprises over 400K instances from 21 datasets and 33 models. Moreover, it provides comprehensive metrics for both performance-oriented routing and performance-cost trade-off routing, and integrates 10 representative routing baselines. Using LLMRouterBench, we systematically re-evaluate the field. While confirming strong model complementarity-the central premise of LLM routing-we find that many routing methods exhibit similar performance under unified evaluation, and several recent approaches, including commercial routers, fail to reliably outperform a simple baseline. Meanwhile, a substantial gap remains to the Oracle, driven primarily by persistent model-recall failures. We further show that backbone embedding models have limited impact, that larger ensembles exhibit diminishing returns compared to careful model curation, and that the benchmark also enables latency-aware analysis. All code and data are available at https://github.com/ynulihao/LLMRouterBench.",
"arxiv_id": "2601.07206",
"authors": [
"Hao Li",
"Yiqun Zhang",
"Zhaoyan Guo",
"Chenxu Wang",
"Shengji Tang",
"Qiaosheng Zhang",
"Yang Chen",
"Biqing Qi",
"Peng Ye",
"Lei Bai",
"Zhen Wang",
"Shuyue Hu"
],
"categories": [
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing",
"url": "https://arxiv.org/abs/2601.07206",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "323441f9-ed91-45fa-90a6-c809d667849b",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}