Tuesday, September 22, 2026

Learning prompt injection against six AI personas with ai-ctf

Learning prompt injection against six AI personas with ai-ctf

ai-ctf is a self-contained local Capture-the-Flag platform from mubix that teaches prompt injection, tool-call abuse, and OSINT against simulated chatbot personas for authorized security training.

Toolmubix/ai-ctf — local AI Capture-the-Flag platform with guided prompt-injection lessons and six chatbot personas
CategoryAI security training / CTF platform (Python, FastAPI, Docker)
Primary UseRunning hands-on workshops and labs that teach prompt injection, tool authorization abuse, business-logic manipulation, and OSINT against fictional bots
Safe UseAll targets are simulated in-memory fixtures inside Docker on the operator's own hardware — no real employees, mail servers, or external services are involved; suited for authorized corporate training and team exercises
Telemetry NoteFully local: Ollama inference runs in its own container, no host ports published except the web UI on 18080; the organizer can export a complete chat-log HTML report via dump_chats.py for event forensics

ai-ctf is Rob Fuller's (mubix) answer to a recurring training problem: everyone is shipping LLM-powered assistants, but few practitioners have a safe place to actually practice attacking one. The repo, MIT-licensed and largely AI-generated code with human-owned design, packages a complete local Capture-the-Flag environment where players learn prompt injection by exploiting six simulated AI personas — one of them hidden — guarding 20 flags. What distinguishes it from most CTF bundles is the pedagogical layer: beyond the self-directed practice labs, there is a guided learning path with hints, saved attempts, tool-call evidence, completion feedback, and protected tool comparisons. This is a teaching instrument first and a game second.

Architecturally the platform is a three-service Docker Compose stack: a web service built from a FastAPI application, an ollama container doing local model inference, and a tiny decoy nginx serving two flag pages. The default model is qwen2.5:7b-instruct-q4_K_M, a roughly 4.7 GB download that runs quantized on CPU — no GPU passthrough is configured. After the build and model pull, the core stack works fully offline, which matters for air-gapped training rooms and conference networks. Only the web UI publishes a host port (18080 by default, overridable via CTF_WEB_PORT); Ollama and the decoy stay on the internal Compose network.

The persona roster covers a genuinely broad attack surface. The Customer Service bot is the prompt-extraction and hidden-coupon warmup; the HR bot guards a seeded SQLite employee database seeded via data/init.sql, and players learn to coax SQL errors that leak schema information; the Code Review bot invites injection embedded in code comments and docstrings — the most realistic scenario in the set, since it mirrors indirect prompt injection through untrusted content. The Checkout bot is a business-logic playground where the goal is a $0 order and then a negative-total order, teaching that LLM-mediated authorization failures are often economic rather than linguistic. The Web Retrieval bot adds allowlist escapes, internal LAN reachability, and file:// LFI triggered through the LLM itself.

The guided path is where the design shows the most care. Rather than free-form prompting against static bots, lessons use isolated fictional fixtures — an HR tool-access scenario and an editable knowledge article that can redirect a simulated reply. The platform executes schema-validated model action requests against fixture tools, records the tool calls, and validates their results. Crucially, players can replay identical arguments through a permission check and inspect a legitimate-use control, so the lesson isn't just "I made the bot misbehave" but "here is exactly where the tool boundary broke." Nothing is emailed and no real HR service is ever connected; the beginner profile deliberately trusts a claimed support-operator role so first attempts succeed, and the platform distinguishes solving a lesson from merely viewing the worked example.

The standout scenario is **Email Joe in Product Sales**, a simulated inbox and Windows-style desktop. Players write an email that Joe's AI assistant reads, then observe the actual model summary alongside the recorded actions the assistant took. Three objectives — a misleading sales brief, disclosure of a fictional internal file, and an unauthorized discount — force the player to think like an attacker targeting agentic workflows: indirect injection arriving as untrusted email content, acted upon by a tool-wielding assistant. The File Explorer uses familiar paths like C:\Users\Joe\Documents\Sales, but these identify in-memory fixtures that never touch the host filesystem. An "Enforce Joe's tool permissions" toggle lets the same objectives be rerun with application-level checks, then verified against a clean legitimate example — a compact demonstration of why input filtering and output authorization must be layered.

Four of the twenty flags (#14–17) require external infrastructure the organizer pushes themselves, documented in external_artifacts/README.md: a fake supply-chain package (anvil_chatkit), a public Gist, and a DNS TXT record on a domain you control. This OSINT chain — fingerprinting the platform from HTTP headers, finding the GitHub repo, digging through commit history, following links to the Gist, querying DNS — is the most operationally demanding part to run, but also the most instructive, because it teaches that LLM applications leak identity through their entire surrounding footprint, not just the chat interface. The remaining 16 flags are fully self-contained in the platform, so the external chain is skippable for a quick lab.

Deployment is a standard clone-and-compose flow: git clone https://github.com/mubix/ai-ctf.git, generate a SESSION_SECRET into platform/.env, then docker compose build, pull the model, and bring the stack up. Registration generates a 12-character one-time password, and progress persists across sessions. One operational wrinkle worth noting: application and template changes require rebuilding the web image (docker compose up -d --build web) — a restart alone won't pick up code changes — and lesson tables are created on startup without deleting existing accounts or chat history, so mid-event updates are survivable.

Customization is refreshingly direct. All 20 flag values live in a single platform/flags.toml, and a helper script apply_flags.py stamps those values into the non-runtime files: the SQL seed, the decoy HTML pages, EXIF metadata in app/static/logo.jpg (requiring exiftool), and the external artifacts. The README is honest about the sharp edge here: the seed uses INSERT OR IGNORE, so changing the CEO salary flag in an existing installation needs a targeted migration, not a database wipe. For deeper changes there is no DSL — the README states plainly that the personas in platform/app/personas/__init__.py and platform/app/tools.py ARE the game, with difficulty tuned by loosening guardrail clauses in persona prompts or swapping OLLAMA_MODEL entirely.

For blue-team readers, the platform doubles as a detection laboratory. The dump_chats.py script exports the bind-mounted SQLite chat log as a self-contained HTML report with per-message flag-detection chips — and the README correctly cautions that these are substring matches, not proof an assistant disclosed a secret or executed a tool. That distinction between "the string appeared" and "the action executed" is exactly the forensic discipline needed when auditing real LLM deployments, and the guided lessons keep validated completion records that demonstrate what evidence-based scoring looks like. The Web recon flags (HTTP headers, robots.txt, EXIF metadata) are equally useful as a checklist of what your own AI-facing services expose.

The docs/ directory is what elevates this from a code dump to a runnable event: SETUP_RUNBOOK.md for the week-of checklist, EVENT_DAY_NOTES.md with a player-briefing script and tiered hint catalog, CHEAT_SHEET.md with working solutions restricted to the game master, ANSWER_KEY.md for verification, and OPERATIONS.md covering backups and pre-event checks that explicitly re-test normal tasks, worked examples, and native tool calls after any model change. The README's note that model size alone does not establish difficulty is a mature observation — a bigger model can be either more resistant to injection or more capable of following attacker instructions.

With 57 stars and a single-maintainer footprint, ai-ctf is a boutique project rather than an ecosystem, but that is appropriate: it is event-grade training material from a well-known operator, designed to be cloned, customized, and run by an instructor who reads the runbook. For security teams onboarding engineers onto LLM threat modeling, for CTF organizers wanting an AI track that doesn't depend on a third-party API, and for red teamers who need a rehearsal environment before testing a client's chatbot deployment, ai-ctf compresses a lot of hard-won instructional design into a stack you can carry on a laptop and run without internet.

Official project repository for mubix/ai-ctf.
Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.

Share articleFacebookXLinkedIn

Continue exploring

Browse all articles →

0 comentários:

Post a Comment

Note: Only a member of this blog may post a comment.