dorsal/arxiv
View SchemaAntiPaSTO: Self-Supervised Steering of Moral Reasoning
| Authors | Michael J. Clark |
|---|---|
| Categories | |
| ArXiv ID | 2601.07473vv2 |
| URL | https://arxiv.org/abs/2601.07473 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
As models grow more capable, human supervision breaks down: labels don't scale, outputs can be gamed, and training doesn't generalize. Scalable oversight requires steering methods that are internal, self-supervised, and transfer out-of-distribution; existing methods satisfy some but not all three. We introduce AntiPaSTO, which separates representations along an anti-parallel axis ($\alpha=\pm1$ produce opposite shifts), with coherence constraints preventing collapse. Human input is minimal: two contrasting words inserted into template sentences, no preference labels. Using 800 such pairs on Gemma-3-1B, AntiPaSTO beats prompting baselines by $6.9\times$ on DailyDilemmas and maintains bidirectional control where prompting triggers refusal. Code is available at https://github.com/wassname/AntiPaSTO.
{
"annotation_id": "69bc96bb-d299-43a2-b49a-3aa0ad725231",
"date_created": "2026-02-17T05:53:12.365000Z",
"date_modified": "2026-02-17T05:53:12.365000Z",
"file_hash": "cf32c0da93d1be9e5c59aa6a8cf9db180c6e22da447edfd998cd9aabad216ecb",
"private": false,
"record": {
"abstract": "As models grow more capable, human supervision breaks down: labels don\u0027t scale, outputs can be gamed, and training doesn\u0027t generalize. Scalable oversight requires steering methods that are internal, self-supervised, and transfer out-of-distribution; existing methods satisfy some but not all three. We introduce AntiPaSTO, which separates representations along an anti-parallel axis ($\\alpha=\\pm1$ produce opposite shifts), with coherence constraints preventing collapse. Human input is minimal: two contrasting words inserted into template sentences, no preference labels. Using 800 such pairs on Gemma-3-1B, AntiPaSTO beats prompting baselines by $6.9\\times$ on DailyDilemmas and maintains bidirectional control where prompting triggers refusal.\n Code is available at https://github.com/wassname/AntiPaSTO.",
"arxiv_id": "2601.07473",
"authors": [
"Michael J. Clark"
],
"categories": [
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "AntiPaSTO: Self-Supervised Steering of Moral Reasoning",
"url": "https://arxiv.org/abs/2601.07473",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "858b6a30-f714-4831-9293-ef85b2f47fb2",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}