dorsal/arxiv
View SchemaLP-LLM: End-to-End Real-World Degraded License Plate Text Recognition via Large Multimodal Models
| Authors | Haoyan Gong, Hongbin Liu |
|---|---|
| Categories | |
| ArXiv ID | 2601.09116vv1 |
| URL | https://arxiv.org/abs/2601.09116 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Real-world License Plate Recognition (LPR) faces significant challenges from severe degradations such as motion blur, low resolution, and complex illumination. The prevailing "restoration-then-recognition" two-stage paradigm suffers from a fundamental flaw: the pixel-level optimization objectives of image restoration models are misaligned with the semantic goals of character recognition, leading to artifact interference and error accumulation. While Vision-Language Models (VLMs) have demonstrated powerful general capabilities, they lack explicit structural modeling for license plate character sequences (e.g., fixed length, specific order). To address this, we propose an end-to-end structure-aware multimodal reasoning framework based on Qwen3-VL. The core innovation lies in the Character-Aware Multimodal Reasoning Module (CMRM), which introduces a set of learnable Character Slot Queries. Through a cross-attention mechanism, these queries actively retrieve fine-grained evidence corresponding to character positions from visual features. Subsequently, we inject these character-aware representations back into the visual tokens via residual modulation, enabling the language model to perform autoregressive generation based on explicit structural priors. Furthermore, combined with the LoRA parameter-efficient fine-tuning strategy, the model achieves domain adaptation while retaining the generalization capabilities of the large model. Extensive experiments on both synthetic and real-world severely degraded datasets demonstrate that our method significantly outperforms existing restoration-recognition combinations and general VLMs, validating the superiority of incorporating structured reasoning into large models for low-quality text recognition tasks.
{
"annotation_id": "14332506-9ba2-47f3-a9f5-8f1bcd041a1e",
"date_created": "2026-02-17T05:53:20.064000Z",
"date_modified": "2026-02-17T05:53:20.064000Z",
"file_hash": "2710e3de7f02ddf94bca088298721ff547bb729c57fd51c3f75fbf12373f1894",
"private": false,
"record": {
"abstract": "Real-world License Plate Recognition (LPR) faces significant challenges from severe degradations such as motion blur, low resolution, and complex illumination. The prevailing \"restoration-then-recognition\" two-stage paradigm suffers from a fundamental flaw: the pixel-level optimization objectives of image restoration models are misaligned with the semantic goals of character recognition, leading to artifact interference and error accumulation. While Vision-Language Models (VLMs) have demonstrated powerful general capabilities, they lack explicit structural modeling for license plate character sequences (e.g., fixed length, specific order). To address this, we propose an end-to-end structure-aware multimodal reasoning framework based on Qwen3-VL. The core innovation lies in the Character-Aware Multimodal Reasoning Module (CMRM), which introduces a set of learnable Character Slot Queries. Through a cross-attention mechanism, these queries actively retrieve fine-grained evidence corresponding to character positions from visual features. Subsequently, we inject these character-aware representations back into the visual tokens via residual modulation, enabling the language model to perform autoregressive generation based on explicit structural priors. Furthermore, combined with the LoRA parameter-efficient fine-tuning strategy, the model achieves domain adaptation while retaining the generalization capabilities of the large model. Extensive experiments on both synthetic and real-world severely degraded datasets demonstrate that our method significantly outperforms existing restoration-recognition combinations and general VLMs, validating the superiority of incorporating structured reasoning into large models for low-quality text recognition tasks.",
"arxiv_id": "2601.09116",
"authors": [
"Haoyan Gong",
"Hongbin Liu"
],
"categories": [
"cs.CV",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "LP-LLM: End-to-End Real-World Degraded License Plate Text Recognition via Large Multimodal Models",
"url": "https://arxiv.org/abs/2601.09116",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "714ce33c-f05b-46ab-913b-48d1f7c328ea",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}