
garak is NVIDIA's open-source vulnerability scanner for large language models, running probes for hallucination, prompt injection, data leakage, and jailbreaks against models you are authorized to assess.
| Tool | NVIDIA/garak — Generative AI Red-teaming & Assessment Kit, an LLM vulnerability scanner (≈9.2k stars, Apache-2.0, Python) |
| Category | LLM security scanner / automated red-teaming framework |
| Primary Use | Running systematic probes (--spec probes.dan, probes.encoding, probes.promptinject) against huggingface, openai, replicate, bedrock, and REST-accessible models to measure failure rates before release |
| Safe Use | Intended for security teams and developers evaluating models they own or are contracted to test, in labs, staging environments, and authorized AI red-team engagements |
| Telemetry Note | Writes garak.log and a detailed .jsonl run log locally; remote-facing footprint is API calls to the target model provider, visible in provider-side request logs and token billing |
garak — the Generative AI Red-teaming & Assessment Kit from NVIDIA — occupies a niche that classical tooling simply cannot: it applies nmap-style systematic scanning to large language models, checking whether an LLM can be made to fail in ways its operators don't want. The README is explicit about the analogy to nmap and msf/Metasploit Framework, but the targets here are dialog systems and generative models rather than network services. The probe surface covers hallucination, data leakage, prompt injection, misinformation, toxicity generation, and jailbreaks, and the project couples static, dynamic, and adaptive probing techniques to explore those weaknesses. At roughly 9,265 stars, an Apache-2.0 license, and backing from a major vendor, it has become the reference open-source scanner in the LLM evaluation space.
Architecturally, garak is a plugin-driven system with two main axes: generators and probes. Generators abstract the model under test — the README lists Hugging Face hub models (local via the Pipeline API, or remote via huggingface.InferenceAPI and huggingface.InferenceEndpoint), replicate, the openai API, AWS bedrock, litellm, NIM endpoints from build.nvidia.com, groq, cohere, ggml/llama.cpp models, and, critically, a flexible rest.RestGenerator that can point at essentially any REST endpoint returning plaintext or JSON with a short YAML descriptor. Probes generate adversarial prompts; detectors evaluate the responses. Each probe ships with recommended detectors, so you rarely need to wire evaluation logic yourself.
Installation is deliberately boring: python -m pip install -U garak from PyPI, or python -m pip install -U git+https://github.com/NVIDIA/garak.git@main for the fresher development tip. The README also documents a Conda-based source workflow (conda create --name garak "python>=3.11,<=3.13"), which hints at the dependency footprint — pulling in the transformers/Hugging Face ecosystem locally will bring weight with it. The tool is developed on Linux and OSX, runs on Windows per the CI badges, and has an active test matrix across all three platforms, which is a good hygiene signal for something you'd want to integrate into a release pipeline.
The invocation model is straightforward: garak <options> with --target_type selecting the generator family and --target_name the specific model. Run bare, garak attempts every probe it knows against the target — which is a real cost consideration against metered APIs. Scoping is done via --spec: --spec probes.promptinject runs only the PromptInject framework's methods, while --spec probes.lmrc.SlurUsage narrows to a single plugin implementing the Language Model Risk Cards check for slur generation. garak --list_probes enumerates the arsenal. This granularity matters for CI, where you want fast, deterministic subsets rather than a full sweep on every commit.
Credentials are handled through conventional environment variables per provider: OPENAI_API_KEY, REPLICATE_API_TOKEN, COHERE_API_KEY, GROQ_API_KEY, NIM_API_KEY, HF_INFERENCE_TOKEN, and BEDROCK_API_KEY with optional BEDROCK_REGION (defaulting to us-east-1). The Bedrock generator uses the Converse API for unified access across Anthropic Claude, Meta Llama, Amazon Titan, AI21, Cohere, and Mistral families. Nothing exotic here — the design keeps secrets out of command lines and configs, which is the correct default for something that may run inside automation.
The README's worked examples are instructive because they show the comparative-analysis workflow the tool enables. One demonstrated run of the encoding probe family found a GPT-3-era model vulnerable only to quoted-printable and MIME encoding injections, while a more recent model was substantially more susceptible to encoding-based injection overall — a counterintuitive result that illustrates why empirical probing beats assumptions about model robustness. This is the core value proposition: garak turns vague safety concerns into comparable, per-probe failure rates you can track across model versions.
Output handling is well thought out for assessment documentation. Each probe prints a progress bar during generation, then a per-detector result row; any prompt attempt that produced undesirable behavior is marked FAIL with a failure rate. The notation like 840/840 reflects total generations versus acceptable ones, with 10 generations per prompt by default — that multiplicity exists to catch probabilistic failures that a single sample would miss. Errors land in garak.log, the full run is serialized to a .jsonl file named at run start and end, and analyse/analyse_log.py ranks the probes and prompts that produced the most hits. That JSONL artifact is what you'd attach to a report or diff between regression runs.
The generator zoo includes two deliberately synthetic targets worth knowing about for calibration: test.Blank, which always emits the empty string, and test.Repeat, which echoes the prompt back. These let you validate probe wiring and understand which detectors demand refutation-style output — test.Blank fails any probe requiring the model to actively rebut a contentious claim. In an authorized workflow, these make good smoke tests before spending budget on a real model, and they demonstrate that the framework treats "the model must do something" as distinct from "the model must not do something bad."
For defenders and model owners, the telemetry profile is clean: everything stays local except API calls to whichever provider hosts the target, which will show up in request logs and token billing on the provider side. The .jsonl run logs and garak.log are written to your own filesystem, making garak suitable for environments where exfiltrating transcripts to a third party would be unacceptable — you'd want to check the license and logging posture of alternative hosted scanners against this. The project also publishes documentation at docs.garak.ai, an arXiv paper (arXiv:2406.11036), DEF CON AI Village slides, and an active Discord, which speaks to both academic grounding and operational community.
Where garak fits in a professional workflow is squarely pre-deployment: gate a model or a guardrail change with a --spec-scoped probe run in CI, treat rising failure rates on probes.dan or probes.encoding as regressions, and attach the analyse_log.py output to your model card or assessment record. Its breadth — hallucination and misinformation probes sit alongside injection and jailbreak probes — also makes it useful for the softer trust-and-safety dimensions that pure security scanners skip. Two caveats from the README deserve attention: the full-probe default can be expensive against commercial APIs, and whitelist-based model support for openai means unsupported models fail with an informative error rather than silently misbehaving — open an issue or send a PR, as the maintainers explicitly invite.
As an educational artifact, garak is also one of the best-documented bodies of LLM attack taxonomy in open source: browsing --list_probes and the probe plugin tree is effectively a guided tour of how LLMs fail. For red teamers moving into AI assessments, model developers building release gates, and defenders trying to understand what an LLM compromise even looks like, garak provides a shared, reproducible vocabulary — and a scanner that turns that vocabulary into numbers.
NVIDIA/garak.Educational analysis for authorized security professionals. Use only in controlled, authorized environments.
0 comentários:
Post a Comment
Note: Only a member of this blog may post a comment.