dorsal/arxiv
View SchemaTerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback
| Authors | Prithwish Jana, Sam Davidson, Bhavana Bhasker, Andrey Kan, Anoop Deoras, Laurent Callot |
|---|---|
| Categories | |
| ArXiv ID | 2601.08734vv1 |
| URL | https://arxiv.org/abs/2601.08734 |
| DOI | 10.1145/3786583.3786898 |
| License | http://creativecommons.org/licenses/by-nc-nd/4.0/ |
Abstract
Automating Infrastructure-as-Code (IaC) is challenging, and large language models (LLMs) often produce incorrect configurations from natural language (NL). We present TerraFormer, a neuro-symbolic framework for IaC generation and mutation that combines supervised fine-tuning with verifier-guided reinforcement learning, using formal verification tools to provide feedback on syntax, deployability, and policy compliance. We curate two large, high-quality NL-to-IaC datasets, TF-Gen (152k instances) and TF-Mutn (52k instances), via multi-stage verification and iterative LLM self-correction. Evaluations against 17 state-of-the-art LLMs, including ~50x larger models like Sonnet 3.7, DeepSeek-R1, and GPT-4.1, show that TerraFormer improves correctness over its base LLM by 15.94% on IaC-Eval, 11.65% on TF-Gen (Test), and 19.60% on TF-Mutn (Test). It outperforms larger models on both TF-Gen (Test) and TF-Mutn (Test), ranks third on IaC-Eval, and achieves top best-practices and security compliance.
{
"annotation_id": "0fd6395c-ab5b-4e0e-8be8-043987b09202",
"date_created": "2026-02-17T05:53:15.894000Z",
"date_modified": "2026-02-17T05:53:15.894000Z",
"file_hash": "704e94c0df21a4637c15cb2a8a93c57ded703dad32bf163975dd4e98f44ed2b5",
"private": false,
"record": {
"abstract": "Automating Infrastructure-as-Code (IaC) is challenging, and large language models (LLMs) often produce incorrect configurations from natural language (NL). We present TerraFormer, a neuro-symbolic framework for IaC generation and mutation that combines supervised fine-tuning with verifier-guided reinforcement learning, using formal verification tools to provide feedback on syntax, deployability, and policy compliance. We curate two large, high-quality NL-to-IaC datasets, TF-Gen (152k instances) and TF-Mutn (52k instances), via multi-stage verification and iterative LLM self-correction. Evaluations against 17 state-of-the-art LLMs, including ~50x larger models like Sonnet 3.7, DeepSeek-R1, and GPT-4.1, show that TerraFormer improves correctness over its base LLM by 15.94% on IaC-Eval, 11.65% on TF-Gen (Test), and 19.60% on TF-Mutn (Test). It outperforms larger models on both TF-Gen (Test) and TF-Mutn (Test), ranks third on IaC-Eval, and achieves top best-practices and security compliance.",
"arxiv_id": "2601.08734",
"authors": [
"Prithwish Jana",
"Sam Davidson",
"Bhavana Bhasker",
"Andrey Kan",
"Anoop Deoras",
"Laurent Callot"
],
"categories": [
"cs.SE",
"cs.AI"
],
"doi": "10.1145/3786583.3786898",
"license": "http://creativecommons.org/licenses/by-nc-nd/4.0/",
"title": "TerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback",
"url": "https://arxiv.org/abs/2601.08734",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "a46a5832-74b4-4dce-8d6c-07e253094b22",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}