Skip to content
AI Defense Lab

AI Defense Lab

Your AI Security Learning Platform

Start with one experiment

Same input. Different policy. What changes?

Open the Threat Lab, run its ready-made example, and inspect what the detectors noticed. Then try Monitor Only to see why detecting an attack and blocking it are different decisions.

Try your first experiment

Use the supplied examples: detection and policy evaluation run for real, while dangerous tool actions are simulated. Do not paste passwords, personal data, or confidential documents. An allowed result does not prove an input is safe.

What You'll Learn

Explore how attacks influence AI systems, what pattern-based detectors notice, and how policies respond. Follow the evidence, including what these defenses miss.

Learning Modules

Lab Activity

These totals are shared across visitors to this lab, not your personal learning progress. Block rate is the share of policy-evaluated runs that were blocked; RAG retrieval queries are excluded because no policy judges them. It is not detector accuracy: the submitted inputs have no labelled ground truth against which to measure misses.

—

Total Runs

—

Runs Blocked

—

Active Policies

—

Block Rate

What these detectors actually catch

Measured against a labelled corpus of 34 attacks and 30 benign inputs, written from the threat model rather than from the detectors’ own pattern lists — so the test and the code cannot agree by construction. Regenerate with pytest tests/test_detection_efficacy.py -s -k report. Measured on 2026-10-09 on main commit f3f0a12 plus the change that regenerated this file.

82%
recall · 28 of 34 attacks
95% interval 66.5–91.7%
0%
false positives · 0 of 30 benign inputs
95% interval 0–11.4%

Intervals are 95% Wilson score intervals on the counts above: how much a count this small can say, not whether the cases were well chosen. 0 cases errored (a detector raised); errors are counted as neither caught nor blocked and stay in the denominator.

Read this as a regression baseline, not a benchmark. The corpus was assembled by the same people who wrote the detectors, so 82% here is not comparable to a published evaluation on an independent adversarial dataset. It answers “has detection got worse?”, not “how would this fare against a real attacker?” For scale: At a 1% false-positive rate, five previously published detectors caught between 1.97% and 20.37% of injections on the PromptShield evaluation split, and between 0% and 9.39% at 0.1%. The ledger entry states the conditions behind that range, and why it is not comparable to this one (ACM CODASPY 2025 (extended technical report)).

By attack class

Data exfiltration
6/6
Prompt injection
13/19
RAG poisoning
3/3
Tool hijacking
6/6

By whether patterns can reach it

Normalisable
7/7

Character-level obfuscation — homoglyphs, zero-width characters, spacing, leetspeak. Folding the input to a canonical form reaches these.

Literal
20/20

The hostile wording is plain English. A pattern can match it directly, so a miss here is missing coverage.

Semantic only
1/7

Paraphrase, other languages, or intent split across turns. No pattern list closes these; they need a model that understands meaning.

OWASP Top 10 for LLM Applications

2026 revision — verified 2026-10-07 against genai.owasp.org. Every identifier below links to its own source page.

“Detectors” below means one or more regex detectors in this lab emit that identifier. It is not a claim that the risk is handled. At a 1% false-positive rate, five previously published detectors caught between 1.97% and 20.37% of injections on the PromptShield evaluation split, and between 0% and 9.39% at 0.1%. The authors built that split and calibrated the thresholds on it, their training data are English only, and multi-turn and function-calling attacks were out of scope. One of the same detectors scores 71% on its vendor's own data where this study measured 12.8%: a detection rate depends on the data and the operating point. At least two of the five are learned classifiers, not pattern matchers. This lab's 82% on its own 34-case corpus is a third, different measurement and is not comparable. ACM CODASPY 2025 (extended technical report) (independent study, not a vendor claim).

IDNameDescriptionDetectors
LLM01Prompt Injection

User input alters the model's behaviour or output in ways the developer did not intend.

override_detector, encoding_detector, poisoning_detector
LLM02Sensitive Information Disclosure

The model or its surrounding application exposes PII, credentials, business data, or its own training material.

exfil_detector
LLM03Excessive Agency

The model is granted more functionality, permission, or autonomy than the task requires.

tool_detector
LLM04Supply Chain

Compromised models, datasets, adapters, or packages enter the application through its dependencies.

No detector
LLM05Data and Model Poisoning

Training or fine-tuning data is manipulated to implant bias, backdoors, or degraded behaviour in the model itself.

No detector
LLM06Unbounded Consumption

Uncontrolled inference volume or size drives denial of service, cost exhaustion, or model extraction.

No detector
LLM07Misinformation

The model produces confident, plausible, and wrong output that downstream users or systems act on.

No detector
LLM08Hidden Context Exposure

Hidden context — the system prompt and the rules, tools and data placed beside it — is disclosed, revealing instructions, logic, or secrets.

override_detector
LLM09Vector and Embedding Weaknesses

The retrieval layer itself is the attack surface — poisoned documents, cross-tenant leakage, or inversion of stored embeddings.

poisoning_detector
LLM10Improper Output Handling

Model output is passed to a downstream system without validation, turning the LLM into an injection relay.

No detector