Skip to content
AI Defense Lab

Threat Lab

Learn how AI threat detection works, run attack scenarios, and explore the results step by step.

LEARN— What is Threat Detection?

AI systems can be tricked by specially crafted inputs called “prompt injections.” Detectors scan for known attack patterns in real time.

PRACTICE— Try an Attack

Choose an Attack Scenario

What to Expect

Technique: Direct instruction override via 'ignore previous' and role reassignment phrasing

Look for: Look for instruction-override patterns in the evidence. Compare the detector findings with the policy's recorded reason for blocking or allowing this input.

The selected policy determines the decision. These learning prompts describe the supplied example, not a guaranteed outcome for edited input.

Attack Input

Edit this text — try modifying the attack!

Ready to Detect

Choose an attack scenario above and press Run Detection. Your input goes to 5 regex detectors, each scoring it independently and returning a confidence, a severity, and the exact spans it matched. The policy engine then turns those findings into one decision. Everything you see below is the real pipeline output, recorded as a replayable trace.