Safety & Security · Software component

Jailbreak Detector

Software componentSafety & SecuritySafety, Security & GovernanceVariation point (abstract)arc:JailbreakDetector

A detection component that identifies jailbreak and prompt-injection attempts in user input, such as role-play manipulations or known attack preambles.

Responsibility. Detects attempts to subvert the model's safety behaviour through crafted input.

Also known as: Jailbreak detection, Prompt injection detector

is invoked byescalates tois invoked byis evaluated byis specialized byis specialized byGuardrail Orchestrator: is invoked byGuardrail OrchestratorHuman Specialist: escalates toHuman SpecialistInput Rail: is invoked byInput RailRed Team Tester: is evaluated byRed Team TesterHeuristic Jailbreak Detector: is specialized byHeuristic Jailbreak Dete…Classifier Jailbreak Detector: is specialized byClassifier Jailbreak Det…
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Classifier Jailbreak DetectorChoose for deeper analysis in high-stakes applications (financial services, healthcare, government).
Heuristic Jailbreak DetectorChoose for fast initial screening of all traffic; effective against documented attacks but not novel ones.

Relationships

is invoked by dependency

escalates to dynamic

is evaluated by assurance

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Perplexity analysisSignature matchingLayered detection with human review of ambiguous cases
Risks mitigated
JailbreakPrompt injectionSQL injection via promptsInvisible-character injectionUnicode homoglyph substitutionMulti-turn split prohibited terms

Sources

  1. Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
  2. Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
  3. Ch8.2B: T. Nguyen, "NeMo Guardrails Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.2B. ISBN: 9798244538229.
  4. Ch10.5: T. Nguyen, "Human-over-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.5. ISBN: 9798244538229.
  5. Ref7.03: NVIDIA, "Overview," NVIDIA NeMo Guardrails Library Developer Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/guardrails/about-nemo-guardrails-library/overview
  6. Ref9.01: "AI Safety Frameworks for Agent Systems," unpublished reference note (01-AI-Safety-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note