dorsal/arxiv
View SchemaCLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark
| Authors | Daniil Gurgurov, Yusser Al Ghussin, Tanja Baeumel, Cheng-Ting Chou, Patrick Schramowski, Marius Mosbach, Josef van Genabith, Simon Ostermann |
|---|---|
| Categories | |
| ArXiv ID | 2601.08331vv1 |
| URL | https://arxiv.org/abs/2601.08331 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Understanding and controlling the behavior of large language models (LLMs) is an increasingly important topic in multilingual NLP. Beyond prompting or fine-tuning, , i.e.,~manipulating internal representations during inference, has emerged as a more efficient and interpretable technique for adapting models to a target language. Yet, no dedicated benchmarks or evaluation protocols exist to quantify the effectiveness of steering techniques. We introduce CLaS-Bench, a lightweight parallel-question benchmark for evaluating language-forcing behavior in LLMs across 32 languages, enabling systematic evaluation of multilingual steering methods. We evaluate a broad array of steering techniques, including residual-stream DiffMean interventions, probe-derived directions, language-specific neurons, PCA/LDA vectors, Sparse Autoencoders, and prompting baselines. Steering performance is measured along two axes: language control and semantic relevance, combined into a single harmonic-mean steering score. We find that across languages simple residual-based DiffMean method consistently outperforms all other methods. Moreover, a layer-wise analysis reveals that language-specific structure emerges predominantly in later layers and steering directions cluster based on language family. CLaS-Bench is the first standardized benchmark for multilingual steering, enabling both rigorous scientific analysis of language representations and practical evaluation of steering as a low-cost adaptation alternative.
{
"annotation_id": "7622c122-3427-4ee6-b33c-db5df4fd278d",
"date_created": "2026-02-17T05:53:15.167000Z",
"date_modified": "2026-02-17T05:53:15.167000Z",
"file_hash": "a18ac240f71a3fcafe477f32a0932f2198027b04d4ea51c359f51fefee817b83",
"private": false,
"record": {
"abstract": "Understanding and controlling the behavior of large language models (LLMs) is an increasingly important topic in multilingual NLP. Beyond prompting or fine-tuning, , i.e.,~manipulating internal representations during inference, has emerged as a more efficient and interpretable technique for adapting models to a target language. Yet, no dedicated benchmarks or evaluation protocols exist to quantify the effectiveness of steering techniques. We introduce CLaS-Bench, a lightweight parallel-question benchmark for evaluating language-forcing behavior in LLMs across 32 languages, enabling systematic evaluation of multilingual steering methods. We evaluate a broad array of steering techniques, including residual-stream DiffMean interventions, probe-derived directions, language-specific neurons, PCA/LDA vectors, Sparse Autoencoders, and prompting baselines. Steering performance is measured along two axes: language control and semantic relevance, combined into a single harmonic-mean steering score. We find that across languages simple residual-based DiffMean method consistently outperforms all other methods. Moreover, a layer-wise analysis reveals that language-specific structure emerges predominantly in later layers and steering directions cluster based on language family. CLaS-Bench is the first standardized benchmark for multilingual steering, enabling both rigorous scientific analysis of language representations and practical evaluation of steering as a low-cost adaptation alternative.",
"arxiv_id": "2601.08331",
"authors": [
"Daniil Gurgurov",
"Yusser Al Ghussin",
"Tanja Baeumel",
"Cheng-Ting Chou",
"Patrick Schramowski",
"Marius Mosbach",
"Josef van Genabith",
"Simon Ostermann"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark",
"url": "https://arxiv.org/abs/2601.08331",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "d55866e6-85d0-4cba-b280-14c172638868",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}