dorsal/arxiv
View Schema$D^2Prune$: Sparsifying Large Language Models via Dual Taylor Expansion and Attention Distribution Awareness
| Authors | Lang Xiong, Ning Liu, Ao Ren, Yuheng Bai, Haining Fang, BinYan Zhang, Zhe Jiang, Yujuan Tan, Duo Liu |
|---|---|
| Categories | |
| ArXiv ID | 2601.09176vv1 |
| URL | https://arxiv.org/abs/2601.09176 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Large language models (LLMs) face significant deployment challenges due to their massive computational demands. % While pruning offers a promising compression solution, existing methods suffer from two critical limitations: (1) They neglect activation distribution shifts between calibration data and test data, resulting in inaccurate error estimations; (2) They overlook the long-tail distribution characteristics of activations in the attention module. To address these limitations, this paper proposes a novel pruning method, $D^2Prune$. First, we propose a dual Taylor expansion-based method that jointly models weight and activation perturbations for precise error estimation, leading to precise pruning mask selection and weight updating and facilitating error minimization during pruning. % Second, we propose an attention-aware dynamic update strategy that preserves the long-tail attention pattern by jointly minimizing the KL divergence of attention distributions and the reconstruction error. Extensive experiments show that $D^2Prune$ consistently outperforms SOTA methods across various LLMs (e.g., OPT-125M, LLaMA2/3, and Qwen3). Moreover, the dynamic attention update mechanism also generalizes well to ViT-based vision models like DeiT, achieving superior accuracy on ImageNet-1K.
{
"annotation_id": "5b4544ce-3bff-44a3-951e-e9ff2a80fd9d",
"date_created": "2026-02-17T05:53:20.402000Z",
"date_modified": "2026-02-17T05:53:20.402000Z",
"file_hash": "1b333db20afe8908ad51e61d5ff6ede50d0b7b74f109e75433c51121479ce270",
"private": false,
"record": {
"abstract": "Large language models (LLMs) face significant deployment challenges due to their massive computational demands. % While pruning offers a promising compression solution, existing methods suffer from two critical limitations: (1) They neglect activation distribution shifts between calibration data and test data, resulting in inaccurate error estimations; (2) They overlook the long-tail distribution characteristics of activations in the attention module. To address these limitations, this paper proposes a novel pruning method, $D^2Prune$. First, we propose a dual Taylor expansion-based method that jointly models weight and activation perturbations for precise error estimation, leading to precise pruning mask selection and weight updating and facilitating error minimization during pruning. % Second, we propose an attention-aware dynamic update strategy that preserves the long-tail attention pattern by jointly minimizing the KL divergence of attention distributions and the reconstruction error. Extensive experiments show that $D^2Prune$ consistently outperforms SOTA methods across various LLMs (e.g., OPT-125M, LLaMA2/3, and Qwen3). Moreover, the dynamic attention update mechanism also generalizes well to ViT-based vision models like DeiT, achieving superior accuracy on ImageNet-1K.",
"arxiv_id": "2601.09176",
"authors": [
"Lang Xiong",
"Ning Liu",
"Ao Ren",
"Yuheng Bai",
"Haining Fang",
"BinYan Zhang",
"Zhe Jiang",
"Yujuan Tan",
"Duo Liu"
],
"categories": [
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "$D^2Prune$: Sparsifying Large Language Models via Dual Taylor Expansion and Attention Distribution Awareness",
"url": "https://arxiv.org/abs/2601.09176",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "70ae7ee3-94bb-4835-ae67-91bb02638138",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}