dorsal/arxiv
View SchemaDeconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
| Authors | Bo Wang, Junzhuo Li, Hong Chen, Yuanlin Chu, Yuxuan Fan, Xuming Hu |
|---|---|
| Categories | |
| ArXiv ID | 2601.08383vv1 |
| URL | https://arxiv.org/abs/2601.08383 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Mixture-of-Experts (MoE) architectures decouple model capacity from per-token computation, enabling scaling beyond the computational limits imposed by dense scaling laws. Yet how MoE architectures shape knowledge acquisition during pre-training, and how this process differs from dense architectures, remains unknown. To address this issue, we introduce Gated-LPI (Log-Probability Increase), a neuron-level attribution metric that decomposes log-probability increase across neurons. We present a time-resolved comparison of knowledge acquisition dynamics in MoE and dense architectures, tracking checkpoints over 1.2M training steps (~ 5.0T tokens) and 600K training steps (~ 2.5T tokens), respectively. Our experiments uncover three patterns: (1) Low-entropy backbone. The top approximately 1% of MoE neurons capture over 45% of positive updates, forming a high-utility core, which is absent in the dense baseline. (2) Early consolidation. The MoE model locks into a stable importance profile within < 100K steps, whereas the dense model remains volatile throughout training. (3) Functional robustness. Masking the ten most important MoE attention heads reduces relational HIT@10 by < 10%, compared with > 50% for the dense model, showing that sparsity fosters distributed -- rather than brittle -- knowledge storage. These patterns collectively demonstrate that sparsity fosters an intrinsically stable and distributed computational backbone from early in training, helping bridge the gap between sparse architectures and training-time interpretability.
{
"annotation_id": "b4642c88-1a18-478a-9f90-ae0cbd3e3f16",
"date_created": "2026-02-17T05:53:15.967000Z",
"date_modified": "2026-02-17T05:53:15.967000Z",
"file_hash": "5fffc31f95ddbef8a4e2fa41c655b834a0da81b18329521dfcae86c2f93bc05c",
"private": false,
"record": {
"abstract": "Mixture-of-Experts (MoE) architectures decouple model capacity from per-token computation, enabling scaling beyond the computational limits imposed by dense scaling laws. Yet how MoE architectures shape knowledge acquisition during pre-training, and how this process differs from dense architectures, remains unknown. To address this issue, we introduce Gated-LPI (Log-Probability Increase), a neuron-level attribution metric that decomposes log-probability increase across neurons. We present a time-resolved comparison of knowledge acquisition dynamics in MoE and dense architectures, tracking checkpoints over 1.2M training steps (~ 5.0T tokens) and 600K training steps (~ 2.5T tokens), respectively. Our experiments uncover three patterns: (1) Low-entropy backbone. The top approximately 1% of MoE neurons capture over 45% of positive updates, forming a high-utility core, which is absent in the dense baseline. (2) Early consolidation. The MoE model locks into a stable importance profile within \u003c 100K steps, whereas the dense model remains volatile throughout training. (3) Functional robustness. Masking the ten most important MoE attention heads reduces relational HIT@10 by \u003c 10%, compared with \u003e 50% for the dense model, showing that sparsity fosters distributed -- rather than brittle -- knowledge storage. These patterns collectively demonstrate that sparsity fosters an intrinsically stable and distributed computational backbone from early in training, helping bridge the gap between sparse architectures and training-time interpretability.",
"arxiv_id": "2601.08383",
"authors": [
"Bo Wang",
"Junzhuo Li",
"Hong Chen",
"Yuanlin Chu",
"Yuxuan Fan",
"Xuming Hu"
],
"categories": [
"cs.AI",
"cs.CL",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models",
"url": "https://arxiv.org/abs/2601.08383",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "4950ba4c-d19d-4152-95d0-052e0d61a7cf",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}