dorsal/arxiv
View SchemaPlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs
| Authors | Zijing Wang, Yongkang Liu, Mingyang Wang, Ercong Nie, Deyuan Chen, Zhengjie Zhao, Shi Feng, Daling Wang, Xiaocui Yang, Yifei Zhang, Hinrich Schütze |
|---|---|
| Categories | |
| ArXiv ID | 2601.07645vv1 |
| URL | https://arxiv.org/abs/2601.07645 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Multimodal Large Language Models (MLLMs) rely on strong linguistic reasoning inherited from their base language models. However, multimodal instruction fine-tuning paradoxically degrades this text's reasoning capability, undermining multimodal performance. To address this issue, we propose a training-free framework to mitigate this degradation. Through layer-wise vision token masking, we reveal a common three-stage pattern in multimodal large language models: early-modal separation, mid-modal alignment, and late-modal degradation. By analyzing the behavior of MLLMs at different stages, we propose a plateau-guided model merging method that selectively injects base language model parameters into MLLMs. Experimental results based on five MLLMs on nine benchmarks demonstrate the effectiveness of our method. Attention-based analysis further reveals that merging shifts attention from diffuse, scattered patterns to focused localization on task-relevant visual regions. Our repository is on https://github.com/wzj1718/PlaM.
{
"annotation_id": "bf0568c3-1031-47d8-8e8b-c5d56796abe9",
"date_created": "2026-02-17T05:53:12.481000Z",
"date_modified": "2026-02-17T05:53:12.481000Z",
"file_hash": "f2fa3b915a04d32c7f9fe936231548d98136d96fcd08b4cb41d70a759a98c853",
"private": false,
"record": {
"abstract": "Multimodal Large Language Models (MLLMs) rely on strong linguistic reasoning inherited from their base language models. However, multimodal instruction fine-tuning paradoxically degrades this text\u0027s reasoning capability, undermining multimodal performance. To address this issue, we propose a training-free framework to mitigate this degradation. Through layer-wise vision token masking, we reveal a common three-stage pattern in multimodal large language models: early-modal separation, mid-modal alignment, and late-modal degradation. By analyzing the behavior of MLLMs at different stages, we propose a plateau-guided model merging method that selectively injects base language model parameters into MLLMs. Experimental results based on five MLLMs on nine benchmarks demonstrate the effectiveness of our method. Attention-based analysis further reveals that merging shifts attention from diffuse, scattered patterns to focused localization on task-relevant visual regions. Our repository is on https://github.com/wzj1718/PlaM.",
"arxiv_id": "2601.07645",
"authors": [
"Zijing Wang",
"Yongkang Liu",
"Mingyang Wang",
"Ercong Nie",
"Deyuan Chen",
"Zhengjie Zhao",
"Shi Feng",
"Daling Wang",
"Xiaocui Yang",
"Yifei Zhang",
"Hinrich Sch\u00fctze"
],
"categories": [
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "PlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs",
"url": "https://arxiv.org/abs/2601.07645",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "cf0d6598-1ffe-423d-96a6-208029481f8e",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}