Thursday, September 24, 2026

WAInjectBench for measuring prompt injection detection in web agents

WAInjectBench for measuring prompt injection detection in web agents

WAInjectBench is a benchmark suite that evaluates how well text and image prompt-injection detectors protect LLM-driven web agents, serving defenders and researchers in authorized security testing.

ToolNorrrrrrr-lyn/WAInjectBench — a benchmark for prompt injection detection in web agents across text and image modalities
CategoryAI/LLM security benchmarking and evaluation framework (Python)
Primary UseMeasuring and comparing detectors such as PromptGuard, DataSentinel, JailGuard, and custom embedding classifiers against labeled benign/malicious datasets
Safe UseIntended for authorized defensive research, red-team evaluation of web agents you own or are contracted to assess, and academic study of LLM robustness
Telemetry NotePurely an offline evaluation harness: it reads local JSONL datasets and model checkpoints, writes results to a --result_dir, and only contacts external services when an OPENAI_API_KEY-backed detector is selected

As web agents — autonomous systems that browse pages, read their content, and act on instructions — become standard plumbing in enterprise automation, the attack surface shifts from traditional injection against applications to injection against the LLM reasoning loop itself. WAInjectBench is a research-grade answer to the measurement problem that follows: if a malicious webpage contains adversarial text or an adversarial image, does your detector actually catch it? The project, written in Python and hosted at Norrrrrrr-lyn/WAInjectBench, packages a labeled corpus plus reproducible evaluation pipelines for both textual and visual prompt-injection detection, making it a useful fixture for anyone doing authorized robustness assessment of agent stacks.

The scope claimed by the README is notably broader than most injection benchmarks, which tend to stop at text. WAInjectBench covers six families of attacks spread across two modalities: text and image. The dataset layout mirrors that split cleanly under data/, with data/text/benign/ holding four categories of benign content in JSONL files and data/text/malicious/ holding eight attack types; the image side organizes data/image/benign/ into two categories and data/image/malicious/ into seven attack types, stored as subfolders rather than line-delimited records. That asymmetry — more malicious classes than benign ones — signals a corpus designed to stress recall across attack variety rather than to merely balance a binary classification task.

Architecturally the project is a classic harness-plus-adapters design. Two entry points, main_text.py and main_image.py, each accept the same four arguments: --data_dir for the dataset, --detector to select the system under test, --result_dir for outputs, and --gpu for device selection. This uniform interface is the quiet strength of the benchmark — swapping detection backends requires no pipeline surgery, which makes apples-to-apples comparison across heterogeneous detectors feasible. The results directory pattern also means runs are auditable, a property evaluators should insist on when reporting numbers.

The text-side detector roster reads like a who's-who of recent academic and industrial defenses: kad, promptarmor, embedding-t, promptguard, datasentinel, and an ensemble mode. The image side offers gpt-4o-prompt, llava-1.5-7b-prompt, jailguard, embedding-i, a finetuned llava-1.5-7b-ft, and again ensemble. Two design decisions stand out here. First, the benchmark treats prompt-based VLM judges as detectors in their own right, not just as ground-truth oracles, which reflects how many teams deploy detection today. Second, the presence of ensemble in both modalities acknowledges the empirical reality that no single detector dominates across all attack families.

Integration burden varies by detector, and the README is candid about it. PromptArmor and GPT-4o-Prompt simply require an OPENAI_API_KEY environment variable, since they delegate judgment to a hosted model. DataSentinel demands more: you must clone the upstream Open-Prompt-Injection repository from liu00222, download its pretrained weights into WAInjectBench/Open-Prompt-Injection/DataSentinel_Models, and wire the paths in detector_text/datasentinel.py. Similarly, JailGuard requires cloning shiningrain/JailGuard and configuring MiniGPT4 per its own documentation, and the finetuned llava-1.5-7b-ft detector requires fetching the authors' checkpoint and pointing detector_image/llava.py at it. Evaluators should budget time for this dependency wrangling before trusting any cross-detector comparison.

Beyond stock configurations, the project supports in-domain generalization experiments — arguably the most interesting question for practitioners. The README ships in-domain trained variants of both embedding classifiers under model/embedding-t/in-domain and model/embedding-i/in-domain, usable through the same evaluation entry points once you update paths in detector_text/embedding-t.py and detector_image/embedding-i.py. This lets you contrast detectors trained on the benchmark's own distribution against their general-purpose counterparts, which is precisely the overfitting question that plagues security ML benchmarks.

The training code is included, not hidden behind a paper. train/embedding-t.py consumes JSONL records of the form {"text": ..., "label": 1} where 1 marks malicious and 0 benign; train/embedding-i.py takes the same schema with {"path": ..., "label": ...} pointing at images. For teams that want a multimodal detector built on a stronger backbone, the project also includes LLaVA-1.5-7B finetuning via train.py, with flags such as --use_lora, --amp_dtype bf16, --device_mode single, and --gpu_id, using the paper's default hyperparameters. The LoRA option keeps the compute footprint of adapting a 7B VLM within reach of a single-GPU lab rather than a cluster.

Getting started is straightforward in the sanctioned sense: git clone https://github.com/Norrrrrrr-lyn/WAInjectBench.git, then conda env create -f environment.yml and conda activate wainjectbench. Everything runs locally against the bundled datasets; no external target is contacted and no live exploitation is involved. The benchmark is consumable by defenders red-teaming their own agent deployments, by vendors wanting an independent yardstick for their detector, and by researchers reproducing the accompanying paper's numbers.

From a defensive-operations standpoint, the value of WAInjectBench is less in any single score and more in the discipline it imposes. Because the corpus separates modality and attack family, a failed run tells you where a detector breaks — for instance, an embedding classifier that performs well on textual instruction hijacking but collapses on malicious imagery passed through a VLM context window. That granularity converts prompt-injection defense from a vague assurance into a measurable gap analysis, which is exactly what an authorized assessment of a web-agent deployment should deliver.

Caveats worth noting: the repository carries no explicit license at time of writing, so commercial evaluation shops should clarify reuse terms before folding the corpus into proprietary pipelines. The README also omits paper-citation details in the excerpt, though the finetuned-model references and hyperparameter notes make clear the artifacts accompany a published study. With 23 stars and a Python codebase on the main branch, it is an early-stage academic project — expect rough edges in adapter configuration rather than a turnkey product. Even so, as a documented, reproducible yardstick for multimodal prompt-injection detection, it earns a place in the assessment toolkit.

Official project repository for Norrrrrrr-lyn/WAInjectBench.
Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.

Share articleFacebookXLinkedIn

Continue exploring

Browse all articles →

0 comentários:

Post a Comment

Note: Only a member of this blog may post a comment.