dorsal/arxiv
View SchemaRealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction
| Authors | Haonan Bian, Zhiyuan Yao, Sen Hu, Zishan Xu, Shaolei Zhang, Yifu Guo, Ziliang Yang, Xueran Han, Huacan Wang, Ronghao Chen |
|---|---|
| Categories | |
| ArXiv ID | 2601.06966vv1 |
| URL | https://arxiv.org/abs/2601.06966 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
As Large Language Models (LLMs) evolve from static dialogue interfaces to autonomous general agents, effective memory is paramount to ensuring long-term consistency. However, existing benchmarks primarily focus on casual conversation or task-oriented dialogue, failing to capture **"long-term project-oriented"** interactions where agents must track evolving goals. To bridge this gap, we introduce **RealMem**, the first benchmark grounded in realistic project scenarios. RealMem comprises over 2,000 cross-session dialogues across eleven scenarios, utilizing natural user queries for evaluation. We propose a synthesis pipeline that integrates Project Foundation Construction, Multi-Agent Dialogue Generation, and Memory and Schedule Management to simulate the dynamic evolution of memory. Experiments reveal that current memory systems face significant challenges in managing the long-term project states and dynamic context dependencies inherent in real-world projects. Our code and datasets are available at [https://github.com/AvatarMemory/RealMemBench](https://github.com/AvatarMemory/RealMemBench).
{
"annotation_id": "5bcbe515-fe68-47fe-a697-5a9e818874ea",
"date_created": "2026-02-17T05:53:07.904000Z",
"date_modified": "2026-02-17T05:53:07.904000Z",
"file_hash": "dcf1882a666567e76ef391b7e2804c1fa3783ccec68383269fae0b7fc8ab6177",
"private": false,
"record": {
"abstract": "As Large Language Models (LLMs) evolve from static dialogue interfaces to autonomous general agents, effective memory is paramount to ensuring long-term consistency. However, existing benchmarks primarily focus on casual conversation or task-oriented dialogue, failing to capture **\"long-term project-oriented\"** interactions where agents must track evolving goals.\n To bridge this gap, we introduce **RealMem**, the first benchmark grounded in realistic project scenarios. RealMem comprises over 2,000 cross-session dialogues across eleven scenarios, utilizing natural user queries for evaluation.\n We propose a synthesis pipeline that integrates Project Foundation Construction, Multi-Agent Dialogue Generation, and Memory and Schedule Management to simulate the dynamic evolution of memory. Experiments reveal that current memory systems face significant challenges in managing the long-term project states and dynamic context dependencies inherent in real-world projects.\n Our code and datasets are available at [https://github.com/AvatarMemory/RealMemBench](https://github.com/AvatarMemory/RealMemBench).",
"arxiv_id": "2601.06966",
"authors": [
"Haonan Bian",
"Zhiyuan Yao",
"Sen Hu",
"Zishan Xu",
"Shaolei Zhang",
"Yifu Guo",
"Ziliang Yang",
"Xueran Han",
"Huacan Wang",
"Ronghao Chen"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction",
"url": "https://arxiv.org/abs/2601.06966",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "9aaf8bde-7e39-421b-9240-0f21318ae938",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}