Back close

Comprehensive Analysis of Machine Learning and Deep Learning models on Prompt Injection Classification using Natural Language Processing techniques

Publication Type : Journal Article

Publisher : Asian Research Association

Source : International Research Journal of Multidisciplinary Technovation

Url : https://doi.org/10.54392/irjmt2523

Campus : Amritapuri

School : School of Computing

Year : 2025

Abstract : This study addresses the prompt injection attack based vulnerability in large language models, which poses a significant security concern by allowing unauthorized commands by attackers to manipulate the outputs produced by model. Text classification methods used for detecting these malicious prompts are investigated on the prompt injection dataset obtained from Hugging Face datasets, utilizing a combination of natural language processing-based techniques applied on various machine learning and deep learning algorithms. Multiple vectorization approaches, like the Term Frequency-Inverse Document Frequency, Word2Vec, Bag of Words, and embeddings, are implemented to transform textual data into meaningful representations. The performance of several classifiers is assessed, on their ability to identify between malicious and non-malicious prompts. The Recurrent Neural Network model demonstrated high accuracy, achieving a detection rate of 94.74%. Obtained results indicated that deep learning architectures, particularly those that capture sequential dependencies, are highly effective in identifying prompt injection threats. This study contributes to the evolving field of AI security by addressing the issue of defending LLM based systems against adversarial threats in form of prompt injections. The findings highlight the importance of integrating sequential dependencies and contextual understanding in combatting LLM vulnerabilities. By the application of reliable detection mechanisms, this study enhances the security, integrity, and trustworthiness of AI-driven technologies, ensuring their safe use across diverse applications.

Cite this Research Publication : Bhavvya Jain, Pranav Pawar, Dhruv Gada, Tanish Patwa, Pratik Kanani, Deepali Patil, Lakshmi Kurup, Comprehensive Analysis of Machine Learning and Deep Learning models on Prompt Injection Classification using Natural Language Processing techniques, International Research Journal of Multidisciplinary Technovation, Asian Research Association, 2025, https://doi.org/10.54392/irjmt2523

Admissions Apply Now