Safety & Security · Software component
Jailbreak Detector
Software componentSafety & SecuritySafety, Security & GovernanceVariation point (abstract)arc:JailbreakDetector
A detection component that identifies jailbreak and prompt-injection attempts in user input, such as role-play manipulations or known attack preambles.
Responsibility. Detects attempts to subvert the model's safety behaviour through crafted input.
Also known as: Jailbreak detection, Prompt injection detector
Variants
| Variant | When to choose |
|---|---|
| Classifier Jailbreak Detector | Choose for deeper analysis in high-stakes applications (financial services, healthcare, government). |
| Heuristic Jailbreak Detector | Choose for fast initial screening of all traffic; effective against documented attacks but not novel ones. |
Relationships
is invoked by dependency
escalates to dynamic
is evaluated by assurance
Design guidance
- MUST NOT be treated as reliable on its own; novel semantic variations, encodings and multi-turn erosion evade heuristics.
- SHOULD combine heuristics, fine-tuned classifiers and human review for high-stakes domains, and update detection rules continuously.
- SHOULD set detection thresholds by explicitly trading false negatives against false positives.
Quantitative guidance
As stated by the sources; verify before use.
- Layered heuristics + classifiers + human review achieve 95%+ detection with false positives below 1% (Ch7.1B).
Classification
- Patterns
- Perplexity analysisSignature matchingLayered detection with human review of ambiguous cases
- Risks mitigated
- JailbreakPrompt injectionSQL injection via promptsInvisible-character injectionUnicode homoglyph substitutionMulti-turn split prohibited terms
Sources
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
- Ch8.2B: T. Nguyen, "NeMo Guardrails Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.2B. ISBN: 9798244538229.
- Ch10.5: T. Nguyen, "Human-over-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.5. ISBN: 9798244538229.
- Ref7.03: NVIDIA, "Overview," NVIDIA NeMo Guardrails Library Developer Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/guardrails/about-nemo-guardrails-library/overview
- Ref9.01: "AI Safety Frameworks for Agent Systems," unpublished reference note (01-AI-Safety-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note