A Python-based Command Line Interface (CLI) tool designed to scan Large Language Model (LLM) prompts for potential security vulnerabilities, including jailbreaks, prompt injections, and other unsafe behaviors.
- Detect Prompt Injection Attacks: Identifies attempts to manipulate LLM behavior through malicious input.
- Detect Jailbreak Attempts: Recognizes prompts designed to bypass LLM safety guidelines and restrictions.
- Risk Scoring: Assigns a risk level (Low, Medium, High, Critical) to scanned prompts.
- HTML + JSON Report Export: Generates comprehensive scan reports in both HTML and JSON formats.
- Batch Scanning Support: Allows scanning multiple prompts from a file for efficient analysis.
- Color-coded Terminal Output: Provides clear and intuitive feedback directly in the command line.
To get started with the LLM Security Scanner, follow these steps:
-
Clone the repository:
git clone https://github.com/Prem2868/llm-security-scanner.git cd llm-security-scanner -
Create a virtual environment (recommended):
python3 -m venv venv source venv/bin/activate -
Install dependencies:
pip install -r requirements.txt
To scan a single prompt directly from the command line:
python main.py scan "Your LLM prompt here"To scan multiple prompts from a text file (one prompt per line):
python main.py batch-scan prompts.txt(Details on report generation will be added here.)
(Screenshots demonstrating the CLI tool and report output will be added here.)
We welcome contributions! Please see the CONTRIBUTING.md file for guidelines on how to contribute to this project.
This project is licensed under the MIT License - see the LICENSE file for details.
- Pramod Jogdand
- GitHub: Prem2868
- Organization: PremLabs-Security
This project is an open-source research prototype for prompt-risk screening. Its detections are heuristic and should be treated as triage signals, not as a complete security verdict. Results can include false positives and false negatives depending on the prompt language, model behavior, policy, and scanner rules. Do not use a single score as proof that an LLM application is secure.
For meaningful evaluation, maintain a versioned test set containing benign prompts, known injection patterns, jailbreak attempts, multilingual cases, and application-specific policies. Report precision, recall, false-positive rate, and the test-set version alongside any benchmark. Do not include confidential prompts, credentials, customer data, or proprietary system instructions in public examples.
The report-generation section and screenshot examples are intentionally marked as work in progress until reproducible sample outputs and evaluation evidence are added. Contributions that add sanitized fixtures, tests, sample HTML/JSON reports, and documented limitations are especially welcome.
Use this scanner only with prompts and applications that you own or are authorized to assess. It is designed for defensive testing and security review; it does not provide a guarantee of model safety and should be combined with access control, input handling, output validation, logging, and human review.