dorsal/arxiv
View SchemaPlasticity vs. Rigidity: The Impact of Low-Rank Adapters on Reasoning on a Micro-Budget
| Authors | Zohaib Khan, Omer Tafveez, Zoha Hayat Bhatti |
|---|---|
| Categories | |
| ArXiv ID | 2601.06677vv1 |
| URL | https://arxiv.org/abs/2601.06677 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Recent advances in mathematical reasoning typically rely on massive scale, yet the question remains: can strong reasoning capabilities be induced in small language models ($\leq1.5\text{B}$) under extreme constraints? We investigate this by training models on a single A40 GPU (48GB) for under 24 hours using Reinforcement Learning with Verifiable Rewards (RLVR) and Low-Rank Adaptation (LoRA). We find that the success of this ``micro-budget" regime depends critically on the interplay between adapter capacity and model initialization. While low-rank adapters ($r=8$) consistently fail to capture the complex optimization dynamics of reasoning, high-rank adapters ($r=256$) unlock significant plasticity in standard instruction-tuned models. Our best result achieved an impressive 40.0\% Pass@1 on AIME 24 (an 11.1\% absolute improvement over baseline) and pushed Pass@16 to 70.0\%, demonstrating robust exploration capabilities. However, this plasticity is not universal: while instruction-tuned models utilized the budget to elongate their chain-of-thought and maximize reward, heavily math-aligned models suffered performance collapse, suggesting that noisy, low-budget RL updates can act as destructive interference for models already residing near a task-specific optimum.
{
"annotation_id": "3b5b882b-54f5-47b9-8ae6-64647bd160bd",
"date_created": "2026-02-17T05:53:08.615000Z",
"date_modified": "2026-02-17T05:53:08.615000Z",
"file_hash": "18a3cdbd517d3308424e94202a926c4f807b6ce52e0f0507be9430556bc1359d",
"private": false,
"record": {
"abstract": "Recent advances in mathematical reasoning typically rely on massive scale, yet the question remains: can strong reasoning capabilities be induced in small language models ($\\leq1.5\\text{B}$) under extreme constraints? We investigate this by training models on a single A40 GPU (48GB) for under 24 hours using Reinforcement Learning with Verifiable Rewards (RLVR) and Low-Rank Adaptation (LoRA). We find that the success of this ``micro-budget\" regime depends critically on the interplay between adapter capacity and model initialization. While low-rank adapters ($r=8$) consistently fail to capture the complex optimization dynamics of reasoning, high-rank adapters ($r=256$) unlock significant plasticity in standard instruction-tuned models. Our best result achieved an impressive 40.0\\% Pass@1 on AIME 24 (an 11.1\\% absolute improvement over baseline) and pushed Pass@16 to 70.0\\%, demonstrating robust exploration capabilities. However, this plasticity is not universal: while instruction-tuned models utilized the budget to elongate their chain-of-thought and maximize reward, heavily math-aligned models suffered performance collapse, suggesting that noisy, low-budget RL updates can act as destructive interference for models already residing near a task-specific optimum.",
"arxiv_id": "2601.06677",
"authors": [
"Zohaib Khan",
"Omer Tafveez",
"Zoha Hayat Bhatti"
],
"categories": [
"cs.LG",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Plasticity vs. Rigidity: The Impact of Low-Rank Adapters on Reasoning on a Micro-Budget",
"url": "https://arxiv.org/abs/2601.06677",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "92f50491-ed2f-47e2-98d2-7343f3a55969",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}