Skip to content
AI Defense Lab

Threat Lab

Learn how AI threat detection works, run attack scenarios, and explore the results step by step.

LEARN— What is Threat Detection?

AI systems can be tricked by specially crafted inputs called “prompt injections.” Detectors scan for known attack patterns in real time.

PRACTICE— Try an Attack

Choose an Attack Scenario

What to Expect

Technique: Direct instruction override via 'ignore previous' and role reassignment phrasing

Expected: BLOCK — Override Detector matches several patterns; the highest-weighted ('ignore all previous instructions') carries it to 95% confidence, severity high

Attack Input

Edit this text — try modifying the attack!

Ready to Detect

Choose an attack scenario above and press Run Detection. Your input goes to 5 regex detectors, each scoring it independently and returning a confidence, a severity, and the exact spans it matched. The policy engine then turns those findings into one decision. Everything you see below is the real pipeline output, recorded as a replayable trace.