
PIForge is an open Python framework that trains reinforcement-learning attacker models against prompt injection defenses and agent benchmarks, intended for authorized LLM security research and red teaming.
| Tool | albert-y1n/PIForge — open framework for RL-based prompt injection red teaming, MIT-licensed, Python |
| Category | LLM adversarial testing / reinforcement learning research framework |
| Primary Use | Training and evaluating RL attacker models against prompt injection defenses on benchmarks like AgentDojo, InjecAgent, and PIArena in authorized research environments |
| Safe Use | For authorized security researchers and red teams testing systems they own or have written permission to assess; benchmark suites are self-contained lab environments, not third-party production targets |
| Telemetry Note | Consumes target OPENAI_API_KEY credentials and GPU resources; trained attackers contact benchmark APIs, so defenders should monitor evaluation traffic in sanctioned lab accounts |
PIForge is the shared codebase behind two research papers, PISmith and Climbing the Hill, both concerned with one of the harder problems in LLM security evaluation: teaching a reinforcement-learning attacker to actually succeed often enough to learn from. Prompt injection attacks against well-defended agent stacks are, by nature, sparsely rewarded — most attempts fail, and an RL policy that almost never sees a success signal cannot improve. PISmith tackles that sparse reward problem directly, while Climbing the Hill extends it with curriculum learning to solve the cold-start problem when the target is a hardened frontier model. The repository, MIT-licensed and written in Python, packages both approaches into a single trainer with a deliberately small interface.
Structurally, the project is refreshingly compact. The benchmarks/ directory carries integration for five evaluation suites — PIArena, InjecAgent, AgentDojo, AgentDyn, and IPI Arena — while the training core lives in train.py, the core/ module tree, and configs/. Entry points are reduced to three shell scripts in scripts/ and eval/: one for single-target training, one for curriculum training, and one for evaluation. This consolidation is the design insight; rather than a new harness per paper, PIForge exposes one shared RL trainer parameterized by benchmark and target, which makes comparative reproduction across papers and defenses practical.
The intended workflow is explicit in the README. You clone the repository, create a conda environment on Python 3.10, and install requirements.txt. Training requires GPUs, which signals the expected audience: this is a research-grade tooling for teams with compute budgets, not a quick script. Configuration is handled almost entirely through environment variables — ATTACKER_MODEL, TRAIN_SUITES, OUTPUT_DIR, LEARNING_RATE, and NUM_TRAIN_EPOCHS are the main knobs. A thoughtful touch for operators is DRY_RUN=1, which prints the commands the wrapper would launch without spinning up any models, useful for verifying pipeline wiring before committing GPU hours.
The curriculum mechanism deserves attention because it is the framework's differentiator. A run like bash scripts/train_curriculum.sh nano-luna starts against a weaker target, then each subsequent stage initializes from the preceding stage's attacker checkpoint as the target hardens. The README reports results from the companion paper: 93.8% and 45.0% attack success rate at ten samples (ASR@10) against two frontier targets where prior RL methods reportedly scored zero. Those numbers are the researchers' own claims from their paper rather than independently verified results, but the released evaluation scripts mean anyone with the compute can attempt reproduction against the same benchmark configurations.
Benchmark coverage is where PIForge earns practical relevance for defensive teams. AgentDojo suites are selectable via TRAIN_SUITES or EVAL_SUITES — workspace, banking, travel, or slack — each a simulated agent environment rather than a live production system. For InjecAgent, the framework integrates with Meta's Meta-SecAlign defended target, which you prepare locally with merge_meta_secalign.py and then serve on a GPU. PIArena supports both the secalign defense and an undefended baseline, with other defenses selectable by configuration name such as promptguard or piguard. The inclusion of named defenses means the tool doubles as a standardized stress test for defense comparisons, not just attack generation.
The target model abstraction is clean. Training scripts take a target identifier (gpt4o-mini, gpt5-nano, secalign) and handle the plumbing, with TARGET_GPU, TRAIN_GPUS, and TARGET_URL available to redirect where the target runs — including pointing at an already-running target server. The default layout starts the local target on GPU 0 and uses GPUs 1–3 for training, a sensible split for a single-box setup. Because targets are pluggable, an authorized team evaluating its own internally deployed agent could, in principle, substitute its own endpoint while keeping the training loop and metrics intact — though the README itself only demonstrates the public benchmark targets.
A significant convenience is the set of released attacker checkpoints in the PIForge Hugging Face collection, built on Qwen3-4B backbones. The eval.sh script accepts a Hugging Face model ID directly, so an evaluator can measure a pre-trained attacker against a chosen target without running the expensive RL training at all. From a defensive research standpoint this is arguably the most useful artifact in the project: fixed, versioned adversarial policies provide a repeatable regression baseline, letting defenders check whether a prompt hardening change or a new filter actually moves ASR on a consistent adversary rather than ad hoc prompts.
For defenders, the value of PIForge is inverted from its attacker framing. An attack success rate measured by a persistent RL policy is a far stronger signal than manual prompt fiddling, because the policy has been optimized to find the weak spots a human would miss. Teams building agent platforms with tool access — the exact scenario AgentDojo simulates — can use the curriculum trainer to measure whether their defenses degrade under adaptive pressure, which is precisely the failure mode that static benchmarks miss. The sparse-reward framing of PISmith is itself a defensive insight: if your defense is so robust that even an RL-trained attacker rarely succeeds in a lab, you have quantitative evidence rather than vibes.
Operational cautions are worth stating plainly. The tool consumes an OPENAI_API_KEY for hosted targets, so all evaluation traffic is attributable to that account and should be run under a sanctioned research key, never against endpoints you do not own or have authorization to test. The trained attacker models, while small at 4B parameters, are purpose-built adversarial artifacts; treat released checkpoints with the same handling discipline as any offensive research output. There is no telemetry or covert channel here — everything is explicit API traffic and logged training runs — but the artifacts themselves deserve care.
Maturity-wise, the project is young but anchored by peer review: PISmith was accepted to COLM 2026, and the curriculum follow-up was released in September 2026, indicating active development aligned with the publication cycle. At 32 stars it is early in adoption, which cuts both ways — less community hardening, but the codebase is small enough for a competent reviewer to audit the training loop and benchmark integrations directly before relying on its numbers. For an authorized LLM security team that wants a rigorous, reproducible answer to how well its prompt injection defenses hold under adaptive attack, PIForge is one of the more complete open frameworks currently available.
albert-y1n/PIForge.Educational analysis for authorized security professionals. Use only in controlled, authorized environments.
Related coverage
0 comentários:
Post a Comment
Note: Only a member of this blog may post a comment.