dorsal/arxiv
View SchemaQueueing-Aware Optimization of Reasoning Tokens for Accuracy-Latency Trade-offs in LLM Servers
| Authors | Emre Ozbas, Melih Bastopcu |
|---|---|
| Categories | |
| ArXiv ID | 2601.10274vv1 |
| URL | https://arxiv.org/abs/2601.10274 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
We consider a single large language model (LLM) server that serves a heterogeneous stream of queries belonging to $N$ distinct task types. Queries arrive according to a Poisson process, and each type occurs with a known prior probability. For each task type, the server allocates a fixed number of internal thinking tokens, which determines the computational effort devoted to that query. The token allocation induces an accuracy-latency trade-off: the service time follows an approximately affine function of the allocated tokens, while the probability of a correct response exhibits diminishing returns. Under a first-in, first-out (FIFO) service discipline, the system operates as an $M/G/1$ queue, and the mean system time depends on the first and second moments of the resulting service-time distribution. We formulate a constrained optimization problem that maximizes a weighted average accuracy objective penalized by the mean system time, subject to architectural token-budget constraints and queue-stability conditions. The objective function is shown to be strictly concave over the stability region, which ensures existence and uniqueness of the optimal token allocation. The first-order optimality conditions yield a coupled projected fixed-point characterization of the optimum, together with an iterative solution and an explicit sufficient condition for contraction. Moreover, a projected gradient method with a computable global step-size bound is developed to guarantee convergence beyond the contractive regime. Finally, integer-valued token allocations are attained via rounding of the continuous solution, and the resulting performance loss is evaluated in simulation results.
{
"annotation_id": "33a28302-eb65-4264-abcc-158ca48c95a9",
"date_created": "2026-02-17T05:53:23.338000Z",
"date_modified": "2026-02-17T05:53:23.338000Z",
"file_hash": "a5aeb497ec4abcf97c47990978e4d4fe9e031952759d0d63dba1a7998f77bfc9",
"private": false,
"record": {
"abstract": "We consider a single large language model (LLM) server that serves a heterogeneous stream of queries belonging to $N$ distinct task types. Queries arrive according to a Poisson process, and each type occurs with a known prior probability. For each task type, the server allocates a fixed number of internal thinking tokens, which determines the computational effort devoted to that query. The token allocation induces an accuracy-latency trade-off: the service time follows an approximately affine function of the allocated tokens, while the probability of a correct response exhibits diminishing returns. Under a first-in, first-out (FIFO) service discipline, the system operates as an $M/G/1$ queue, and the mean system time depends on the first and second moments of the resulting service-time distribution. We formulate a constrained optimization problem that maximizes a weighted average accuracy objective penalized by the mean system time, subject to architectural token-budget constraints and queue-stability conditions. The objective function is shown to be strictly concave over the stability region, which ensures existence and uniqueness of the optimal token allocation. The first-order optimality conditions yield a coupled projected fixed-point characterization of the optimum, together with an iterative solution and an explicit sufficient condition for contraction. Moreover, a projected gradient method with a computable global step-size bound is developed to guarantee convergence beyond the contractive regime. Finally, integer-valued token allocations are attained via rounding of the continuous solution, and the resulting performance loss is evaluated in simulation results.",
"arxiv_id": "2601.10274",
"authors": [
"Emre Ozbas",
"Melih Bastopcu"
],
"categories": [
"cs.LG",
"cs.AI",
"cs.IT",
"cs.NI",
"math.IT",
"math.OC"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Queueing-Aware Optimization of Reasoning Tokens for Accuracy-Latency Trade-offs in LLM Servers",
"url": "https://arxiv.org/abs/2601.10274",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "31b55792-7a6f-411c-bcea-0a6e751d8257",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}