dorsal/arxiv
View SchemaSPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
| Authors | Arion Das, Partha Pratim Saha, Amit Dhanda, Vinija Jain, Aman Chadha, Amitava Das |
|---|---|
| Categories | |
| ArXiv ID | 2601.06238vv1 |
| URL | https://arxiv.org/abs/2601.06238 |
| License | http://creativecommons.org/licenses/by-nc-nd/4.0/ |
Abstract
Direct Preference Optimization (DPO) is a principled, scalable alternative to RLHF for aligning large language models from pairwise preferences, but its internal geometric footprint remains undercharacterized, limiting audits, checkpoint comparisons, and failure prediction. We introduce SPINAL (Scaling-law and Preference Integration in Neural Alignment Layers), a diagnostic that measures how alignment reshapes representations across depth by tracing localized structural change layer by layer. Across model families, DPO produces a layerwise calibration effect concentrated in the final decoder blocks (often layers 21-30), where preference gradients most directly affect the next-token distribution. SPINAL encodes each checkpoint as a depth trace over (layer index, contraction score, transport score). The contraction score summarizes how quickly the tail of a layer's spectrum decays (how fast small modes vanish); higher values indicate stronger contraction into fewer effective directions. The transport score summarizes how much the token distribution shifts between adjacent layers using a bounded overlap measure; lower values indicate shorter, smoother steps through representation space. Aligned checkpoints show a late-layer ramp-up in contraction and a smooth reduction in transport, consistent with tightened and stabilized policy mass, while unaligned models trace higher-curvature, more entropic, and geometrically incoherent depth paths. Overall, alignment is geometrically localized: the final layers encode the dominant preference-induced corrections. SPINAL turns this localization into a practical audit signal, quantifying where alignment concentrates, how strongly it manifests, and when it begins to destabilize during training.
{
"annotation_id": "7f4f3212-9500-429e-8a18-0fe1410ce2d6",
"date_created": "2026-02-17T05:53:07.523000Z",
"date_modified": "2026-02-17T05:53:07.523000Z",
"file_hash": "f9b8c95d0e6e3add69ff5b672d3da97355126576188f8e3a6a81e6ecbc5fe31b",
"private": false,
"record": {
"abstract": "Direct Preference Optimization (DPO) is a principled, scalable alternative to RLHF for aligning large language models from pairwise preferences, but its internal geometric footprint remains undercharacterized, limiting audits, checkpoint comparisons, and failure prediction. We introduce SPINAL (Scaling-law and Preference Integration in Neural Alignment Layers), a diagnostic that measures how alignment reshapes representations across depth by tracing localized structural change layer by layer. Across model families, DPO produces a layerwise calibration effect concentrated in the final decoder blocks (often layers 21-30), where preference gradients most directly affect the next-token distribution. SPINAL encodes each checkpoint as a depth trace over (layer index, contraction score, transport score). The contraction score summarizes how quickly the tail of a layer\u0027s spectrum decays (how fast small modes vanish); higher values indicate stronger contraction into fewer effective directions. The transport score summarizes how much the token distribution shifts between adjacent layers using a bounded overlap measure; lower values indicate shorter, smoother steps through representation space. Aligned checkpoints show a late-layer ramp-up in contraction and a smooth reduction in transport, consistent with tightened and stabilized policy mass, while unaligned models trace higher-curvature, more entropic, and geometrically incoherent depth paths. Overall, alignment is geometrically localized: the final layers encode the dominant preference-induced corrections. SPINAL turns this localization into a practical audit signal, quantifying where alignment concentrates, how strongly it manifests, and when it begins to destabilize during training.",
"arxiv_id": "2601.06238",
"authors": [
"Arion Das",
"Partha Pratim Saha",
"Amit Dhanda",
"Vinija Jain",
"Aman Chadha",
"Amitava Das"
],
"categories": [
"cs.LG",
"cs.AI",
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by-nc-nd/4.0/",
"title": "SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers",
"url": "https://arxiv.org/abs/2601.06238",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "c2a10982-52b7-45e8-b38d-dc217ccd90cb",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}