dorsal/arxiv
View SchemaContextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models
| Authors | Abhinaba Basu, Pavan Chakraborty |
|---|---|
| Categories | |
| ArXiv ID | 2601.10460vv1 |
| URL | https://arxiv.org/abs/2601.10460 |
| License | http://creativecommons.org/licenses/by-sa/4.0/ |
Abstract
A model that avoids stereotypes in a lab benchmark may not avoid them in deployment. We show that measured bias shifts dramatically when prompts mention different places, times, or audiences -- no adversarial prompting required. We introduce Contextual StereoSet, a benchmark that holds stereotype content fixed while systematically varying contextual framing. Testing 13 models across two protocols, we find striking patterns: anchoring to 1990 (vs. 2030) raises stereotype selection in all models tested on this contrast (p<0.05); gossip framing raises it in 5 of 6 full-grid models; out-group observer framing shifts it by up to 13 percentage points. These effects replicate in hiring, lending, and help-seeking vignettes. We propose Context Sensitivity Fingerprints (CSF): a compact profile of per-dimension dispersion and paired contrasts with bootstrap CIs and FDR correction. Two evaluation tracks support different use cases -- a 360-context diagnostic grid for deep analysis and a budgeted protocol covering 4,229 items for production screening. The implication is methodological: bias scores from fixed-condition tests may not generalize.This is not a claim about ground-truth bias rates; it is a stress test of evaluation robustness. CSF forces evaluators to ask, "Under what conditions does bias appear?" rather than "Is this model biased?" We release our benchmark, code, and results.
{
"annotation_id": "3670657f-af03-4426-96f8-321fe399f6e4",
"date_created": "2026-02-17T05:53:23.414000Z",
"date_modified": "2026-02-17T05:53:23.414000Z",
"file_hash": "f5040cda6fa35991e5005aeae273a18f49785379edeb950e80b0912e5406e0f8",
"private": false,
"record": {
"abstract": "A model that avoids stereotypes in a lab benchmark may not avoid them in deployment. We show that measured bias shifts dramatically when prompts mention different places, times, or audiences -- no adversarial prompting required.\n We introduce Contextual StereoSet, a benchmark that holds stereotype content fixed while systematically varying contextual framing. Testing 13 models across two protocols, we find striking patterns: anchoring to 1990 (vs. 2030) raises stereotype selection in all models tested on this contrast (p\u003c0.05); gossip framing raises it in 5 of 6 full-grid models; out-group observer framing shifts it by up to 13 percentage points. These effects replicate in hiring, lending, and help-seeking vignettes.\n We propose Context Sensitivity Fingerprints (CSF): a compact profile of per-dimension dispersion and paired contrasts with bootstrap CIs and FDR correction. Two evaluation tracks support different use cases -- a 360-context diagnostic grid for deep analysis and a budgeted protocol covering 4,229 items for production screening.\n The implication is methodological: bias scores from fixed-condition tests may not generalize.This is not a claim about ground-truth bias rates; it is a stress test of evaluation robustness. CSF forces evaluators to ask, \"Under what conditions does bias appear?\" rather than \"Is this model biased?\" We release our benchmark, code, and results.",
"arxiv_id": "2601.10460",
"authors": [
"Abhinaba Basu",
"Pavan Chakraborty"
],
"categories": [
"cs.CL",
"cs.AI",
"cs.CY",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by-sa/4.0/",
"title": "Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models",
"url": "https://arxiv.org/abs/2601.10460",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "d0b944f5-c407-4eaf-800d-f6222cddd950",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}