dorsal/arxiv
View SchemaLLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
| Authors | Pan Liao, Feng Yang, Di Wu, Jinwen Yu, Yuhua Zhu, Wenhui Zhao |
|---|---|
| Categories | |
| ArXiv ID | 2601.06550vv1 |
| URL | https://arxiv.org/abs/2601.06550 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Traditional Multi-Object Tracking (MOT) systems have achieved remarkable precision in localization and association, effectively answering \textit{where} and \textit{who}. However, they often function as autistic observers, capable of tracing geometric paths but blind to the semantic \textit{what} and \textit{why} behind object behaviors. To bridge the gap between geometric perception and cognitive reasoning, we propose \textbf{LLMTrack}, a novel end-to-end framework for Semantic Multi-Object Tracking (SMOT). We adopt a bionic design philosophy that decouples strong localization from deep understanding, utilizing Grounding DINO as the eyes and the LLaVA-OneVision multimodal large model as the brain. We introduce a Spatio-Temporal Fusion Module that aggregates instance-level interaction features and video-level contexts, enabling the Large Language Model (LLM) to comprehend complex trajectories. Furthermore, we design a progressive three-stage training strategy, Visual Alignment, Temporal Fine-tuning, and Semantic Injection via LoRA to efficiently adapt the massive model to the tracking domain. Extensive experiments on the BenSMOT benchmark demonstrate that LLMTrack achieves state-of-the-art performance, significantly outperforming existing methods in instance description, interaction recognition, and video summarization while maintaining robust tracking stability.
{
"annotation_id": "c9da8a88-205b-4063-b318-b5acb54e0278",
"date_created": "2026-02-17T05:53:07.767000Z",
"date_modified": "2026-02-17T05:53:07.767000Z",
"file_hash": "d06278df5472876ac9b36d3ae080e762f07cff904f057afde1138350b47e65b7",
"private": false,
"record": {
"abstract": "Traditional Multi-Object Tracking (MOT) systems have achieved remarkable precision in localization and association, effectively answering \\textit{where} and \\textit{who}. However, they often function as autistic observers, capable of tracing geometric paths but blind to the semantic \\textit{what} and \\textit{why} behind object behaviors. To bridge the gap between geometric perception and cognitive reasoning, we propose \\textbf{LLMTrack}, a novel end-to-end framework for Semantic Multi-Object Tracking (SMOT). We adopt a bionic design philosophy that decouples strong localization from deep understanding, utilizing Grounding DINO as the eyes and the LLaVA-OneVision multimodal large model as the brain. We introduce a Spatio-Temporal Fusion Module that aggregates instance-level interaction features and video-level contexts, enabling the Large Language Model (LLM) to comprehend complex trajectories. Furthermore, we design a progressive three-stage training strategy, Visual Alignment, Temporal Fine-tuning, and Semantic Injection via LoRA to efficiently adapt the massive model to the tracking domain. Extensive experiments on the BenSMOT benchmark demonstrate that LLMTrack achieves state-of-the-art performance, significantly outperforming existing methods in instance description, interaction recognition, and video summarization while maintaining robust tracking stability.",
"arxiv_id": "2601.06550",
"authors": [
"Pan Liao",
"Feng Yang",
"Di Wu",
"Jinwen Yu",
"Yuhua Zhu",
"Wenhui Zhao"
],
"categories": [
"cs.CV",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models",
"url": "https://arxiv.org/abs/2601.06550",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "c69d126a-e2dc-44fa-8672-028f10fb73a6",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}