Safety & Security · Human role

Red Team Tester

Human roleSafety & SecuritySafety, Security & Governancearc:RedTeamTester

A security researcher who attempts to bypass guardrails with novel jailbreaks, subtle prompt injections and social engineering so that discovered bypasses can be fixed.

Responsibility. Finds guardrail bypasses through adversarial testing.

Also known as: Security researcher, Red team, Container security firm, Breach simulation tester, Value red teamer

evaluatesevaluatesevaluatesevaluatessends data toevaluatesevaluatesevaluatesinvokessends data toAgent Controller: evaluatesAgent ControllerGuardrail Orchestrator: evaluatesGuardrail OrchestratorExecution Sandbox: evaluatesExecution SandboxReward Model: evaluatesReward ModelPreference Dataset: sends data toPreference DatasetContent Safety Filter: evaluatesContent Safety FilterConstitutionally Aligned Model: evaluatesConstitutionally Aligned…Jailbreak Detector: evaluatesJailbreak DetectorSandbox Validation Runner: invokesSandbox Validation RunnerPrinciple Adherence Test Suite: sends data toPrinciple Adherence Test…
Direct neighbourhood (hover for relationship types)

Relationships

invokes dependency

sends data to dynamic

evaluates assurance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Red-team exercisesBreach simulation

Sources

  1. Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.
  2. Ch9.1: T. Nguyen, "Output Filtering and Content Moderation," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.1. ISBN: 9798244538229.
  3. Ch9.3: T. Nguyen, "Sandboxing and Transparency Foundations," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.3. ISBN: 9798244538229.
  4. Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
  5. Ch9.6: T. Nguyen, "Value Alignment Frameworks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.6. ISBN: 9798244538229.
  6. Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.
  7. Ch10.5: T. Nguyen, "Human-over-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.5. ISBN: 9798244538229.
  8. Ref9.01: "AI Safety Frameworks for Agent Systems," unpublished reference note (01-AI-Safety-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  9. Ref9.08: "Safety Incident Response for AI Systems," unpublished reference note (08-Safety-Incident-Response.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  10. Ref9.10: "Chapter 9 Summary: Safety, Ethics, and Compliance," unpublished reference note (10-Chapter-9-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note