Part 9 — Safety & Governance

10 chapters · 35.4 study hours allocated in the Study Plan · 0 slide decks · 5 videos · 27 code example files

On this page
  1. Chapters
  2. Chapter summaries
    1. 9.1. Output Filtering
    2. 9.2. Action Constraints
    3. 9.3. Sandboxing and Isolation
    4. 9.4. Fairness Foundations
    5. 9.5. Constitutional AI Principles
    6. 9.6. Standards, Certifications, and Frameworks
    7. 9.7. GDPR Foundations
    8. 9.8. Standards and Frameworks for AI Governance
    9. 9.16. Chapter Summary and Integration
    10. 9.17. Chapter Summary and Integration
    11. Labs

Chapters

Rating tags show which certification knowledge maps rate the chapter H (highly relevant) in at least one item: NV NCP-AAI · AWS AIP-C01 · DBX Databricks GenAI Engineer · GCP Professional ML Engineer · MS AI-102. See Certifications.

Ch. Title Hours Slides Quiz Videos Figures Code H-rated for
9.1 Output Filtering 2.5 — Quiz 2 7 6 NV AWS GCP MS
9.2 Action Constraints 2.2 — Quiz 0 6 21 NV AWS MS
9.3 Sandboxing and Isolation 4.4 — Quiz 3 6 — NV AWS GCP MS
9.4 Fairness Foundations 6.5 — Quiz 0 9 — NV AWS DBX GCP MS
9.5 Constitutional AI Principles 7.6 — — 0 12 — NV AWS GCP MS
9.6 Standards, Certifications, and Frameworks 3.5 — — 0 9 — NV AWS DBX MS
9.7 GDPR Foundations 3.7 — Quiz 0 12 — NV AWS DBX GCP MS
9.8 Standards and Frameworks for AI Governance 5.0 — — 0 12 — NV AWS DBX GCP MS
9.16 Chapter Summary and Integration § — — Quiz — — — NV
9.17 Chapter Summary and Integration § — — Quiz — — — NV

Notes. The Videos column counts the videos shown under each chapter summary below, out of the unique direct links in Part_09_YoutubeVideos.md (“3 of 5”). A video is left out when its link is dead, embedding is disabled, or YouTube’s title does not match the entry; see the link check. Chapters can also list search suggestions instead of links.

† Linked by chapter-family number, not an exact ID match: the deck, quiz, or figure set is numbered differently from this chapter in the source files (for example a quiz or deck numbered 6.2 for chapters 6.2A and 6.2B).

‡ A combined deck that covers more than one chapter.

A chapter that is missing from a certification’s mapping file shows no tag for that certification: the NVIDIA file omits 4.1 and 10.6, and the other four omit 1.8, 9.16, and 9.17.

§ Has no section of its own in Study_Plan.md; the title comes from the Study Plan’s table of contents or a cross-reference there, or (9.16, 9.17) from the quiz list.

Chapter summaries

Summaries are excerpted from Study_Plan.md, which also lists each chapter’s key concepts and self-check questions.

9.1. Output Filtering

Output filtering serves as the critical last line of defense in AI safety, intercepting LLM outputs before delivery to users to prevent harmful content including hate speech, harassment, misinformation, and regulatory violations. This chapter explores multi-layered defense architectures, implementation techniques from keyword matching to ML classifiers, human-in-the-loop moderation workflows, NeMo Guardrails integration, and domain-specific compliance requirements for regulated domains like healthcare and financial services.

Videos (2)
Code examples (6 files)

9.2. Action Constraints

This chapter addresses the fundamental vulnerability of autonomous agents operating with excessive permissions, where the machine-paced execution of 1,000-10,000 operations per minute combined with dynamic behavior synthesis creates risks that traditional human-centric permission models cannot address. It provides comprehensive frameworks for implementing least-privilege permissions, multi-layered defense architectures, and human oversight mechanisms to contain the blast radius of agent misbehavior or compromise.

No videos are shown for this chapter: its list has only search suggestions, or its links failed the link check.

Code examples (21 files)

9.3. Sandboxing and Isolation

Sandboxing represents a fundamental paradigm shift from detection-based safety approaches to containment-based approaches, providing structural guarantees that even perfectly compromised agents cannot escape designated boundaries or propagate damage beyond defined limits. Through layered defense combining process isolation, resource restrictions, filesystem virtualization, and network isolation, sandboxing implements the principle that perfect detection is impossible and systems must design for failure.

Videos (3)

9.4. Fairness Foundations

Fairness in AI extends beyond non-discrimination to address emergent biases from multi-agent interactions and systems-level patterns, requiring continuous demographic auditing, fairness-aware data preparation, constraint-based training, and runtime monitoring rather than one-time testing during development. The challenge involves navigating fundamental mathematical trade-offs between incompatible fairness metrics while implementing comprehensive detection and correction across the entire system lifecycle.

9.5. Constitutional AI Principles

Constitutional AI addresses RLHF’s critical limitations (implicit values, annotation bottleneck, psychological costs) by replacing preference-based learning with explicit, inspectable ethical principles that guide behavior throughout training. The two-phase approach—Phase 1 supervised self-critique with constitutional principles and Phase 2 reinforcement learning from AI feedback—makes values transparent while enabling scalability without human annotation, though implementation remains constrained by principle ambiguity, incomplete coverage, and fundamental value conflicts.

No videos are shown for this chapter: its list has only search suggestions, or its links failed the link check.

9.6. Standards, Certifications, and Frameworks

Value alignment frameworks operationalize abstract ethical principles into systematic technical and governance approaches that ensure AI systems maintain long-term consistency with human values across deployment contexts. The World Economic Forum framework emphasizes that effective alignment requires transparency at every development stage, continuous stakeholder participation beyond initial design, ongoing monitoring to detect drift, and explicit documentation of value conflicts rather than pretending technical methods eliminate inherent tensions.

No videos are shown for this chapter: its list has only search suggestions, or its links failed the link check.

9.7. GDPR Foundations

The General Data Protection Regulation represents a paradigm shift in global data protection, establishing principles-based requirements applicable worldwide to any organization processing EU resident data since May 25, 2018. GDPR compliance requires embedding data protection into operational culture as ongoing governance commitment rather than temporary initiative, with organizations adapting implementations to context while maintaining consistent data protection standards.

No videos are shown for this chapter: its list has only search suggestions, or its links failed the link check.

9.8. Standards and Frameworks for AI Governance

NIST AI Risk Management Framework and ISO/IEC 42001 provide complementary governance approaches where NIST delivers flexible operational risk management while ISO 42001 establishes formal management system structure, with implementation of both frameworks creating more robust governance than either alone. These standards integrate with sector-specific regulations and other management systems into unified governance addressing complete AI system lifecycles.

No videos are shown for this chapter: its list has only search suggestions, or its links failed the link check.

9.16. Chapter Summary and Integration

No summary in the Study Plan for this chapter.

9.17. Chapter Summary and Integration

No summary in the Study Plan for this chapter.

Labs

No lab or legacy example exists for this Part yet. See Labs and Contributing.