dorsal/arxiv
View SchemaSecureCAI: Injection-Resilient LLM Assistants for Cybersecurity Operations
| Authors | Mohammed Himayath Ali, Mohammed Aqib Abdullah, Mohammed Mudassir Uddin, Shahnawaz Alam |
|---|---|
| Categories | |
| ArXiv ID | 2601.07835vv1 |
| URL | https://arxiv.org/abs/2601.07835 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Large Language Models have emerged as transformative tools for Security Operations Centers, enabling automated log analysis, phishing triage, and malware explanation; however, deployment in adversarial cybersecurity environments exposes critical vulnerabilities to prompt injection attacks where malicious instructions embedded in security artifacts manipulate model behavior. This paper introduces SecureCAI, a novel defense framework extending Constitutional AI principles with security-aware guardrails, adaptive constitution evolution, and Direct Preference Optimization for unlearning unsafe response patterns, addressing the unique challenges of high-stakes security contexts where traditional safety mechanisms prove insufficient against sophisticated adversarial manipulation. Experimental evaluation demonstrates that SecureCAI reduces attack success rates by 94.7% compared to baseline models while maintaining 95.1% accuracy on benign security analysis tasks, with the framework incorporating continuous red-teaming feedback loops enabling dynamic adaptation to emerging attack strategies and achieving constitution adherence scores exceeding 0.92 under sustained adversarial pressure, thereby establishing a foundation for trustworthy integration of language model capabilities into operational cybersecurity workflows and addressing a critical gap in current approaches to AI safety within adversarial domains.
{
"annotation_id": "c7b57e7b-95d0-4240-a12b-f671a2745594",
"date_created": "2026-02-17T05:53:12.488000Z",
"date_modified": "2026-02-17T05:53:12.488000Z",
"file_hash": "eada3aa5f5b39a4d058732c3bf8125d985d694d0d43fa7dfbff6a1c2cd3a052d",
"private": false,
"record": {
"abstract": "Large Language Models have emerged as transformative tools for Security Operations Centers, enabling automated log analysis, phishing triage, and malware explanation; however, deployment in adversarial cybersecurity environments exposes critical vulnerabilities to prompt injection attacks where malicious instructions embedded in security artifacts manipulate model behavior. This paper introduces SecureCAI, a novel defense framework extending Constitutional AI principles with security-aware guardrails, adaptive constitution evolution, and Direct Preference Optimization for unlearning unsafe response patterns, addressing the unique challenges of high-stakes security contexts where traditional safety mechanisms prove insufficient against sophisticated adversarial manipulation. Experimental evaluation demonstrates that SecureCAI reduces attack success rates by 94.7% compared to baseline models while maintaining 95.1% accuracy on benign security analysis tasks, with the framework incorporating continuous red-teaming feedback loops enabling dynamic adaptation to emerging attack strategies and achieving constitution adherence scores exceeding 0.92 under sustained adversarial pressure, thereby establishing a foundation for trustworthy integration of language model capabilities into operational cybersecurity workflows and addressing a critical gap in current approaches to AI safety within adversarial domains.",
"arxiv_id": "2601.07835",
"authors": [
"Mohammed Himayath Ali",
"Mohammed Aqib Abdullah",
"Mohammed Mudassir Uddin",
"Shahnawaz Alam"
],
"categories": [
"cs.CR",
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "SecureCAI: Injection-Resilient LLM Assistants for Cybersecurity Operations",
"url": "https://arxiv.org/abs/2601.07835",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "6b75a2a6-1ea2-43b4-97cd-9e8300b18793",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}