dorsal/arxiv
View SchemaBenchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
| Authors | Manyi Zhang, Ji-Fu Li, Zhongao Sun, Haoli Bai, Hui-Ling Zhen, Zhenhua Dong, Xianzhi Yu |
|---|---|
| Categories | |
| ArXiv ID | 2601.09555vv1 |
| URL | https://arxiv.org/abs/2601.09555 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Microscaling Floating-Point (MXFP) has emerged as a promising low-precision format for large language models (LLMs). Despite various post-training quantization (PTQ) algorithms being proposed, they mostly focus on integer quantization, while their applicability and behavior under MXFP formats remain largely unexplored. To address this gap, this work conducts a systematic investigation of PTQ under MXFP formats, encompassing over 7 PTQ algorithms, 15 evaluation benchmarks, and 3 LLM families. The key findings include: 1) MXFP8 consistently achieves near-lossless performance, while MXFP4 introduces substantial accuracy degradation and remains challenging; 2) PTQ effectiveness under MXFP depends strongly on format compatibility, with some algorithmic paradigms being consistently more effective than others; 3) PTQ performance exhibits highly consistent trends across model families and modalities, in particular, quantization sensitivity is dominated by the language model rather than the vision encoder in multimodal LLMs; 4) The scaling factor of quantization is a critical error source in MXFP4, and a simple pre-scale optimization strategy can significantly mitigate its impact. Together, these results provide practical guidance on adapting existing PTQ methods to MXFP quantization.
{
"annotation_id": "32981674-58e5-42e8-8891-a1262f0a4abc",
"date_created": "2026-02-17T05:53:19.551000Z",
"date_modified": "2026-02-17T05:53:19.551000Z",
"file_hash": "3b19ddc7424b9e5ac6277164a4d4791e5b856fd6bd770972fd1415cce7a9beb5",
"private": false,
"record": {
"abstract": "Microscaling Floating-Point (MXFP) has emerged as a promising low-precision format for large language models (LLMs). Despite various post-training quantization (PTQ) algorithms being proposed, they mostly focus on integer quantization, while their applicability and behavior under MXFP formats remain largely unexplored. To address this gap, this work conducts a systematic investigation of PTQ under MXFP formats, encompassing over 7 PTQ algorithms, 15 evaluation benchmarks, and 3 LLM families. The key findings include: 1) MXFP8 consistently achieves near-lossless performance, while MXFP4 introduces substantial accuracy degradation and remains challenging; 2) PTQ effectiveness under MXFP depends strongly on format compatibility, with some algorithmic paradigms being consistently more effective than others; 3) PTQ performance exhibits highly consistent trends across model families and modalities, in particular, quantization sensitivity is dominated by the language model rather than the vision encoder in multimodal LLMs; 4) The scaling factor of quantization is a critical error source in MXFP4, and a simple pre-scale optimization strategy can significantly mitigate its impact. Together, these results provide practical guidance on adapting existing PTQ methods to MXFP quantization.",
"arxiv_id": "2601.09555",
"authors": [
"Manyi Zhang",
"Ji-Fu Li",
"Zhongao Sun",
"Haoli Bai",
"Hui-Ling Zhen",
"Zhenhua Dong",
"Xianzhi Yu"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats",
"url": "https://arxiv.org/abs/2601.09555",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "96266a3b-7c28-4fe3-a10c-e1c10f45aa33",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}