dorsal/arxiv
View SchemaTeleMem: Building Long-Term and Multimodal Memory for Agentic AI
| Authors | Chunliang Chen, Ming Guan, Xiao Lin, Jiaxu Li, Luxi Lin, Qiyi Wang, Xiangyu Chen, Jixiang Luo, Changzhi Sun, Dell Zhang, Xuelong Li |
|---|---|
| Categories | |
| ArXiv ID | 2601.06037vv2 |
| URL | https://arxiv.org/abs/2601.06037 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Large language models (LLMs) excel at many NLP tasks but struggle to sustain long-term interactions due to limited attention over extended dialogue histories. Retrieval-augmented generation (RAG) mitigates this issue but lacks reliable mechanisms for updating or refining stored memories, leading to schema-driven hallucinations, inefficient write operations, and minimal support for multimodal reasoning.To address these challenges, we propose TeleMem, a unified long-term and multimodal memory system that maintains coherent user profiles through narrative dynamic extraction, ensuring that only dialogue-grounded information is preserved. TeleMem further introduces a structured writing pipeline that batches, retrieves, clusters, and consolidates memory entries, substantially improving storage efficiency, reducing token usage, and accelerating memory operations. Additionally, a multimodal memory module combined with ReAct-style reasoning equips the system with a closed-loop observe, think, and act process that enables accurate understanding of complex video content in long-term contexts. Experimental results show that TeleMem surpasses the state-of-the-art Mem0 baseline with 19% higher accuracy, 43% fewer tokens, and a 2.1x speedup on the ZH-4O long-term role-play gaming benchmark.
{
"annotation_id": "237959b8-6d5d-4548-ac9d-410cd5a5bb1b",
"date_created": "2026-02-17T05:53:04.378000Z",
"date_modified": "2026-02-17T05:53:04.378000Z",
"file_hash": "2fdd94c2ec67e1ff3e7c06afca03ae6babb6f779dc8bdb58206ff6262a80083e",
"private": false,
"record": {
"abstract": "Large language models (LLMs) excel at many NLP tasks but struggle to sustain long-term interactions due to limited attention over extended dialogue histories. Retrieval-augmented generation (RAG) mitigates this issue but lacks reliable mechanisms for updating or refining stored memories, leading to schema-driven hallucinations, inefficient write operations, and minimal support for multimodal reasoning.To address these challenges, we propose TeleMem, a unified long-term and multimodal memory system that maintains coherent user profiles through narrative dynamic extraction, ensuring that only dialogue-grounded information is preserved. TeleMem further introduces a structured writing pipeline that batches, retrieves, clusters, and consolidates memory entries, substantially improving storage efficiency, reducing token usage, and accelerating memory operations. Additionally, a multimodal memory module combined with ReAct-style reasoning equips the system with a closed-loop observe, think, and act process that enables accurate understanding of complex video content in long-term contexts. Experimental results show that TeleMem surpasses the state-of-the-art Mem0 baseline with 19% higher accuracy, 43% fewer tokens, and a 2.1x speedup on the ZH-4O long-term role-play gaming benchmark.",
"arxiv_id": "2601.06037",
"authors": [
"Chunliang Chen",
"Ming Guan",
"Xiao Lin",
"Jiaxu Li",
"Luxi Lin",
"Qiyi Wang",
"Xiangyu Chen",
"Jixiang Luo",
"Changzhi Sun",
"Dell Zhang",
"Xuelong Li"
],
"categories": [
"cs.CL",
"cs.AI",
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "TeleMem: Building Long-Term and Multimodal Memory for Agentic AI",
"url": "https://arxiv.org/abs/2601.06037",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "68adac26-5e5a-468a-848f-aa882ae13910",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}