International Journal of Scientific Research and Engineering Development

International Journal of Scientific Research and Engineering Development


( International Peer Reviewed Open Access Journal ) ISSN [ Online ] : 2581 - 7175
Submit Your Manuscript OnlineIJSRED
📑 Paper Information
📑 Paper Title PromptShield AI: A Multi-Agent Architecture for Intelligent Prompt Injection and Jailbreak Attack Detection Using Machine Learning
👤 Authors Jahnavi Somaraju, N. Sree Charan, M. Mythili, T. Reddy Bhargavi, K. Navya Sree
📘 Published Issue Volume 9 Issue 4
📅 Year of Publication 2026
🆔 Unique Identification Number IJSRED-V9I4P53
📝 Abstract
Large language models (LLMs) are increasingly deployed in user-facing applications, which exposes them to prompt injection and jailbreak attacks that override system instructions, exfiltrate data, or elicit disallowed behaviour. Existing defences are largely single-mechanism: a rule filter, a single machine-learning (ML) classifier, or a single LLM-based judge, each of which is comparatively easy to evade once its decision boundary is known. This paper proposes PromptShield AI, a multi-agent architecture that fuses lexical, statistical, semantic, and contextual detection signals through a coordinated set of specialised agents rather than a single monolithic classifier. An Orchestrator Agent decomposes each incoming request and routes it to five parallel detection agents; a Decision Fusion Agent calibrates and combines their outputs; a Response/Mitigation Agent executes the resulting policy (allow, sanitize, or block); and a Logging and Feedback Agent closes the loop for continuous retraining. We describe the architecture, the inter-agent communication protocol, the fusion algorithm, and a reference implementation, and we outline an experimental protocol together with illustrative evaluation results and an ablation study. The results indicate that the coordinated multi-agent design improves detection F1-score and reduces false-positive rate relative to any individual detection mechanism, at an acceptable latency overhead, while remaining more resilient to obfuscation and multi-turn attacks than single-agent baselines.
📝 How to Cite
Jahnavi Somaraju, N. Sree Charan, M. Mythili, T. Reddy Bhargavi, K. Navya Sree, "PromptShield AI: A Multi-Agent Architecture for Intelligent Prompt Injection and Jailbreak Attack Detection Using Machine Learning" International Journal of Scientific Research and Engineering Development, V9(4): Page(471-479) May-June 2026. ISSN: 2581-7175. www.ijsred.com. Published by Scientific and Academic Research Publishing.