Part 6 — Retrieval-Augmented Generation (RAG)

14 chapters · 26.2 study hours allocated in the Study Plan · 8 slide decks · 21 videos · 59 code example files

On this page
  1. Chapters
  2. Chapter summaries
    1. 6.1A. RAG Chunking and Embeddings
    2. 6.1C. RAG Implementation
    3. 6.2A. Vector Database Selection
    4. 6.2B. Production Vector Database Deployment
    5. 6.3A. ETL Fundamentals
    6. 6.3B. ETL Load and Integration
    7. 6.3C. ETL Practice
    8. 6.4. Data Quality Fundamentals
    9. 6.4B. Data Quality Practice
    10. 6.5. Production RAG Architecture
    11. 6.5B. Production RAG Practice
    12. 6.6. Query Decomposition and Adaptive Retrieval
    13. 6.6A. Reranking Implementation
    14. 6.6C. Advanced Retrieval Practice
    15. Additional worked examples
    16. Labs

Chapters

Rating tags show which certification knowledge maps rate the chapter H (highly relevant) in at least one item: NV NCP-AAI · AWS AIP-C01 · DBX Databricks GenAI Engineer · GCP Professional ML Engineer · MS AI-102. See Certifications.

Ch. Title Hours Slides Quiz Videos Figures Code H-rated for
6.1A RAG Chunking and Embeddings 1.0 PDF† Quiz 6.1† 5† 4† 3 NV AWS DBX MS
6.1C RAG Implementation § — PDF† Quiz 6.1† see 6.1A† 4† 9 NV AWS DBX GCP MS
6.2A Vector Database Selection 2.6 PDF Quiz 6.2† 5 12 3 NV AWS DBX GCP MS
6.2B Production Vector Database Deployment 1.7 PDF Quiz 6.2† 0 12 12 NV AWS DBX GCP MS
6.3A ETL Fundamentals 2.9 PDF Quiz 6.3† 1 12 8 NV AWS DBX GCP MS
6.3B ETL Load and Integration 1.3 PDF Quiz 6.3† 0 12 7 NV AWS DBX MS
6.3C ETL Practice § — — Quiz 3 — — NV AWS DBX GCP MS
6.4 Data Quality Fundamentals 4.4 PDF Quiz 0 12 9 NV AWS DBX GCP MS
6.4B Data Quality Practice 2.2 — — 2 — — NV AWS DBX GCP MS
6.5 Production RAG Architecture 3.8 PDF Quiz 1 6 — NV AWS DBX GCP MS
6.5B Production RAG Practice 4.0 — — 0 — — NV AWS DBX GCP MS
6.6 Query Decomposition and Adaptive Retrieval 1.5 PDF Quiz 3 7 8 NV AWS DBX GCP MS
6.6A Reranking Implementation 0.8 — — 1 — — NV AWS DBX GCP MS
6.6C Advanced Retrieval Practice § — — — — — — NV AWS GCP MS

Notes. The Videos column counts the videos shown under each chapter summary below, out of the unique direct links in Part_06_YoutubeVideos.md (“3 of 5”). A video is left out when its link is dead, embedding is disabled, or YouTube’s title does not match the entry; see the link check. Chapters can also list search suggestions instead of links.

† Linked by chapter-family number, not an exact ID match: the deck, quiz, or figure set is numbered differently from this chapter in the source files (for example a quiz or deck numbered 6.2 for chapters 6.2A and 6.2B).

‡ A combined deck that covers more than one chapter.

A chapter that is missing from a certification’s mapping file shows no tag for that certification: the NVIDIA file omits 4.1 and 10.6, and the other four omit 1.8, 9.16, and 9.17.

§ Has no section of its own in Study_Plan.md; the title comes from the Study Plan’s table of contents or a cross-reference there, or (9.16, 9.17) from the quiz list.

Chapter summaries

Summaries are excerpted from Study_Plan.md, which also lists each chapter’s key concepts and self-check questions.

6.1A. RAG Chunking and Embeddings

This chapter establishes the technical foundation for semantic search in RAG systems through embeddings that convert text into high-dimensional vectors where semantic meaning is preserved through spatial relationships. It covers embedding fundamentals, comparing leading embedding models, building production pipelines, implementing hybrid search combining dense and sparse methods, and understanding performance optimization through GPU acceleration.

Videos (5)

Covers chapters 6.1A, 6.1C.

Code examples (3 files)

6.1C. RAG Implementation

No summary in the Study Plan for this chapter.

The videos for this chapter family are shown under chapter 6.1A.

Code examples (9 files)

6.2A. Vector Database Selection

Vector databases represent a fundamental architectural paradigm shift enabling semantic search through specialized indexing optimized for high-dimensional vectors. This chapter covers the vector database landscape comparing six dominant platforms, systematic database selection frameworks, HNSW algorithm parameters, distance metrics, and practical decision frameworks for choosing appropriate infrastructure as systems scale from prototypes to enterprise deployments.

Videos (5)
Code examples (3 files)

6.2B. Production Vector Database Deployment

Production vector database deployments require careful orchestration across connectivity, authentication, persistence, performance, and observability dimensions. This chapter covers Docker Compose configuration for both REST and gRPC endpoints, implements secure Python clients with proper authentication and timeout handling, and demonstrates batch ingestion achieving 10-20x throughput improvements. Comprehensive monitoring and high-availability clustering patterns enable reliable production operation.

Code examples (12 files)

6.3A. ETL Fundamentals

ETL pipelines solve the critical enterprise data integration problem, bridging 70+ fragmented organizational data sources into AI-ready vector databases for RAG systems. This chapter establishes the three-stage systematic approach (Extract, Transform, Load) with specialized connectors for different source types, chunking strategies balancing context and specificity, and quality validation ensuring only clean data enters production knowledge bases. Real-world examples demonstrate measurable business impact: accuracy improvement from 67% to 92%, response time reduction to 3.2 seconds, and $2.1M annual savings through reduced escalations.

Videos (1)
Code examples (8 files)

6.3B. ETL Load and Integration

The load phase completes ETL pipelines by translating transformed data into vector database operations. This chapter explains vector database architecture fundamentals, collection schema design optimizing for RAG retrieval patterns, batch insertion strategies achieving 10-38x throughput improvements, and indexing decisions determining 20-100x performance variance. Production patterns address incremental updates through state management, graceful error handling enabling partial success, comprehensive monitoring for operational visibility, and GPU acceleration for billion-document scale systems.

Code examples (7 files)

6.3C. ETL Practice

No summary in the Study Plan for this chapter.

Videos (3)

6.4. Data Quality Fundamentals

Data quality represents the hidden variable determining whether production RAG systems deliver reliable value or generate catastrophic failures. This chapter establishes the five-dimensional framework (completeness, accuracy, consistency, timeliness, validity) and demonstrates how quality failures amplify through RAG systems. A $4.2 million financial services deployment failed within 72 hours due to 12% duplicate articles with conflicting information, 8% corrupted formatting, and 5% outdated regulatory guidance. The chapter translates abstract quality goals into concrete SLAs: 98% completeness minimum, 99% accuracy, 0.5% duplicate threshold, 95% content reflecting 24-hour changes, 99.9% schema conformance. Comprehensive validation across these dimensions at ingestion, transformation, post-loading, and monitoring stages ensures production-grade reliability.

Code examples (9 files)

6.4B. Data Quality Practice

This chapter provides comprehensive practical implementation of quality validation frameworks for production RAG systems. Through guided and independent practice, learners implement multi-dimensional quality checking, automated remediation, quality monitoring dashboards, and deploy these patterns in real-world scenarios involving diverse data sources and domain-specific requirements. The chapter addresses realistic production failures and establishes patterns to recognize and avoid.

Videos (2)

6.5. Production RAG Architecture

This chapter covers the complete architectural design of production RAG systems operating at enterprise scale with sub-second latency, 90%+ accuracy, strict cost controls, and guaranteed availability. The chapter establishes critical patterns for layered system design (ingestion, storage, retrieval, generation, API, observability), addresses fundamental production challenges absent in prototypes, and implements fault tolerance and deployment strategies ensuring reliable operations.

Videos (1)

6.5B. Production RAG Practice

This chapter moves from production RAG theory to practical implementation through guided and independent challenges. Learners implement health checks validating every critical dependency, conduct load testing to identify bottlenecks, optimize latency and costs, design A/B testing infrastructure for data-driven decisions, and build comprehensive monitoring dashboards. The chapter establishes advanced retrieval techniques including reranking and their cost-benefit analysis, with emphasis on measuring effectiveness before production deployment.

6.6. Query Decomposition and Adaptive Retrieval

This chapter addresses advanced retrieval challenges in production RAG systems through two complementary techniques. Query decomposition breaks complex multi-part questions into focused sub-queries, enabling targeted retrieval of comprehensive context for each component. Adaptive retrieval recognizes when external knowledge is genuinely needed, reducing unnecessary API calls and improving latency. Together, these techniques significantly improve answer quality for complex queries while optimizing efficiency for queries where parametric memory suffices.

Videos (3)
Code examples (8 files)

6.6A. Reranking Implementation

This chapter provides comprehensive understanding and practical implementation of cross-encoder based reranking as a production-grade advanced retrieval technique. The chapter explains the architectural differences between bi-encoders and cross-encoders, guides implementation of two-stage retrieval systems, compares commercial APIs with self-hosted approaches, and establishes production error handling patterns for resilient systems.

Videos (1)

6.6C. Advanced Retrieval Practice

No summary in the Study Plan for this chapter.

Additional worked examples

From more_examples/part_06/:

Labs

No finished lab exists for this Part yet. These legacy example files are prose excerpts with embedded code, kept as source material; they do not count as lab coverage. See Labs.