dorsal/arxiv
View SchemaMASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
| Authors | Yongtong Gu, Songze Li, Xia Hu |
|---|---|
| Categories | |
| ArXiv ID | 2601.08564vv1 |
| URL | https://arxiv.org/abs/2601.08564 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
The increasing misuse of AI-generated texts (AIGT) has motivated the rapid development of AIGT detection methods. However, the reliability of these detectors remains fragile against adversarial evasions. Existing attack strategies often rely on white-box assumptions or demand prohibitively high computational and interaction costs, rendering them ineffective under practical black-box scenarios. In this paper, we propose Multi-stage Alignment for Style Humanization (MASH), a novel framework that evades black-box detectors based on style transfer. MASH sequentially employs style-injection supervised fine-tuning, direct preference optimization, and inference-time refinement to shape the distributions of AI-generated texts to resemble those of human-written texts. Experiments across 6 datasets and 5 detectors demonstrate the superior performance of MASH over 11 baseline evaders. Specifically, MASH achieves an average Attack Success Rate (ASR) of 92%, surpassing the strongest baselines by an average of 24%, while maintaining superior linguistic quality.
{
"annotation_id": "a152d498-6670-4081-9c86-9b2ea961fa51",
"date_created": "2026-02-17T05:53:15.645000Z",
"date_modified": "2026-02-17T05:53:15.645000Z",
"file_hash": "78e2131741ff7d811b0b66233e3048ba47c7764a5182e0960141df6af042fdde",
"private": false,
"record": {
"abstract": "The increasing misuse of AI-generated texts (AIGT) has motivated the rapid development of AIGT detection methods. However, the reliability of these detectors remains fragile against adversarial evasions. Existing attack strategies often rely on white-box assumptions or demand prohibitively high computational and interaction costs, rendering them ineffective under practical black-box scenarios. In this paper, we propose Multi-stage Alignment for Style Humanization (MASH), a novel framework that evades black-box detectors based on style transfer. MASH sequentially employs style-injection supervised fine-tuning, direct preference optimization, and inference-time refinement to shape the distributions of AI-generated texts to resemble those of human-written texts. Experiments across 6 datasets and 5 detectors demonstrate the superior performance of MASH over 11 baseline evaders. Specifically, MASH achieves an average Attack Success Rate (ASR) of 92%, surpassing the strongest baselines by an average of 24%, while maintaining superior linguistic quality.",
"arxiv_id": "2601.08564",
"authors": [
"Yongtong Gu",
"Songze Li",
"Xia Hu"
],
"categories": [
"cs.CR"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization",
"url": "https://arxiv.org/abs/2601.08564",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "ba996cf0-bb33-4001-891b-bc242d5b38e9",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}