Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LLM Security Scanner

Python License Stars Issues

A Python-based Command Line Interface (CLI) tool designed to scan Large Language Model (LLM) prompts for potential security vulnerabilities, including jailbreaks, prompt injections, and other unsafe behaviors.

Table of Contents

Features

  • Detect Prompt Injection Attacks: Identifies attempts to manipulate LLM behavior through malicious input.
  • Detect Jailbreak Attempts: Recognizes prompts designed to bypass LLM safety guidelines and restrictions.
  • Risk Scoring: Assigns a risk level (Low, Medium, High, Critical) to scanned prompts.
  • HTML + JSON Report Export: Generates comprehensive scan reports in both HTML and JSON formats.
  • Batch Scanning Support: Allows scanning multiple prompts from a file for efficient analysis.
  • Color-coded Terminal Output: Provides clear and intuitive feedback directly in the command line.

Installation

To get started with the LLM Security Scanner, follow these steps:

  1. Clone the repository:

    git clone https://github.com/Prem2868/llm-security-scanner.git
    cd llm-security-scanner
  2. Create a virtual environment (recommended):

    python3 -m venv venv
    source venv/bin/activate
  3. Install dependencies:

    pip install -r requirements.txt

Usage

Single Prompt Scan

To scan a single prompt directly from the command line:

python main.py scan "Your LLM prompt here"

Batch Scan from File

To scan multiple prompts from a text file (one prompt per line):

python main.py batch-scan prompts.txt

Generating Reports

(Details on report generation will be added here.)

Screenshots

(Screenshots demonstrating the CLI tool and report output will be added here.)

Contributing

We welcome contributions! Please see the CONTRIBUTING.md file for guidelines on how to contribute to this project.

License

This project is licensed under the MIT License - see the LICENSE file for details.

Author

  • Pramod Jogdand
  • GitHub: Prem2868
  • Organization: PremLabs-Security

🔬 Evaluation scope and limitations

This project is an open-source research prototype for prompt-risk screening. Its detections are heuristic and should be treated as triage signals, not as a complete security verdict. Results can include false positives and false negatives depending on the prompt language, model behavior, policy, and scanner rules. Do not use a single score as proof that an LLM application is secure.

For meaningful evaluation, maintain a versioned test set containing benign prompts, known injection patterns, jailbreak attempts, multilingual cases, and application-specific policies. Report precision, recall, false-positive rate, and the test-set version alongside any benchmark. Do not include confidential prompts, credentials, customer data, or proprietary system instructions in public examples.

📄 Report status

The report-generation section and screenshot examples are intentionally marked as work in progress until reproducible sample outputs and evaluation evidence are added. Contributions that add sanitized fixtures, tests, sample HTML/JSON reports, and documented limitations are especially welcome.

⚠️ Safe use

Use this scanner only with prompts and applications that you own or are authorized to assess. It is designed for defensive testing and security review; it does not provide a guarantee of model safety and should be combined with access control, input handling, output validation, logging, and human review.

About

Defensive CLI for triaging prompt-injection and jailbreak risks in LLM prompts with risk-scored reports.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages