dorsal/arxiv
View SchemaMirrorBench: An Extensible Framework to Evaluate User-Proxy Agents for Human-Likeness
| Authors | Ashutosh Hathidara, Julien Yu, Vaishali Senthil, Sebastian Schreiber, Anil Babu Ankisettipalli |
|---|---|
| Categories | |
| ArXiv ID | 2601.08118vv1 |
| URL | https://arxiv.org/abs/2601.08118 |
| License | http://creativecommons.org/licenses/by-sa/4.0/ |
Abstract
Large language models (LLMs) are increasingly used as human simulators, both for evaluating conversational systems and for generating fine-tuning data. However, naive "act-as-a-user" prompting often yields verbose, unrealistic utterances, underscoring the need for principled evaluation of so-called user proxy agents. We present MIRRORBENCH, a reproducible, extensible benchmarking framework that evaluates user proxies solely on their ability to produce human-like user utterances across diverse conversational tasks, explicitly decoupled from downstream task success. MIRRORBENCH features a modular execution engine with typed interfaces, metadata-driven registries, multi-backend support, caching, and robust observability. The system supports pluggable user proxies, datasets, tasks, and metrics, enabling researchers to evaluate arbitrary simulators under a uniform, variance-aware harness. We include three lexical-diversity metrics (MATTR, YULE'S K, and HD-D) and three LLM-judge-based metrics (GTEval, Pairwise Indistinguishability, and Rubric-and-Reason). Across four open datasets, MIRRORBENCH yields variance-aware results and reveals systematic gaps between user proxies and real human users. The framework is open source and includes a simple command-line interface for running experiments, managing configurations and caching, and generating reports. The framework can be accessed at https://github.com/SAP/mirrorbench.
{
"annotation_id": "3550c306-8526-43cf-ae69-fa10e93be905",
"date_created": "2026-02-17T05:53:16.230000Z",
"date_modified": "2026-02-17T05:53:16.230000Z",
"file_hash": "26310c9a124ac9e8398de9ac13121c0fb2d80f9674a051cd1ff965e1dcd0bd4a",
"private": false,
"record": {
"abstract": "Large language models (LLMs) are increasingly used as human simulators, both for evaluating conversational systems and for generating fine-tuning data. However, naive \"act-as-a-user\" prompting often yields verbose, unrealistic utterances, underscoring the need for principled evaluation of so-called user proxy agents. We present MIRRORBENCH, a reproducible, extensible benchmarking framework that evaluates user proxies solely on their ability to produce human-like user utterances across diverse conversational tasks, explicitly decoupled from downstream task success. MIRRORBENCH features a modular execution engine with typed interfaces, metadata-driven registries, multi-backend support, caching, and robust observability. The system supports pluggable user proxies, datasets, tasks, and metrics, enabling researchers to evaluate arbitrary simulators under a uniform, variance-aware harness. We include three lexical-diversity metrics (MATTR, YULE\u0027S K, and HD-D) and three LLM-judge-based metrics (GTEval, Pairwise Indistinguishability, and Rubric-and-Reason). Across four open datasets, MIRRORBENCH yields variance-aware results and reveals systematic gaps between user proxies and real human users. The framework is open source and includes a simple command-line interface for running experiments, managing configurations and caching, and generating reports. The framework can be accessed at https://github.com/SAP/mirrorbench.",
"arxiv_id": "2601.08118",
"authors": [
"Ashutosh Hathidara",
"Julien Yu",
"Vaishali Senthil",
"Sebastian Schreiber",
"Anil Babu Ankisettipalli"
],
"categories": [
"cs.AI",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by-sa/4.0/",
"title": "MirrorBench: An Extensible Framework to Evaluate User-Proxy Agents for Human-Likeness",
"url": "https://arxiv.org/abs/2601.08118",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "c7c970ff-fce4-439a-892c-5e5e5d3cbbe7",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}