Part 4 — Production Deployment & Scaling
7 chapters · 24.4 study hours allocated in the Study Plan · 7 slide decks · 39 videos · 45 code example files
On this page
Chapters
Rating tags show which certification knowledge maps rate the chapter H (highly relevant) in at least one item: NV NCP-AAI · AWS AIP-C01 · DBX Databricks GenAI Engineer · GCP Professional ML Engineer · MS AI-102. See Certifications.
| Ch. | Title | Hours | Slides | Quiz | Videos | Figures | Code | H-rated for |
|---|---|---|---|---|---|---|---|---|
| 4.1 | AI Agent Deployment and Scaling | 3.2 | Quiz 4.1A†, Quiz 4.1B† | 9 | 6 | — | AWS GCP MS | |
| 4.2 | Deployment & Scaling | 5.8 | — | 1 | 12 | 4 | NV AWS GCP | |
| 4.3 | Container Orchestration and Edge Deployment | 2.1 | Quiz | 7 | 12 | 7 | NV AWS GCP MS | |
| 4.4 | Performance Profiling and Optimization | 6.0 | — | 7 | 12 | 12 | NV AWS GCP MS | |
| 4.5 | NVIDIA NIM and Triton Inference Server | 2.1 | Quiz | 5 | 6 | — | NV AWS GCP MS | |
| 4.6 | TensorRT-LLM and NVIDIA Fleet Command | 1.5 | Quiz | 7 | 9 | — | NV AWS GCP MS | |
| 4.7 | Scaling Strategies | 3.7 | — | 3 | 6 | — | NV AWS GCP MS |
Notes. The Videos column counts the videos shown under each chapter summary below, out of the unique direct links in Part_04_YoutubeVideos.md (“3 of 5”). A video is left out when its link is dead, embedding is disabled, or YouTube’s title does not match the entry; see the link check. Chapters can also list search suggestions instead of links.
† Linked by chapter-family number, not an exact ID match: the deck, quiz, or figure set is numbered differently from this chapter in the source files (for example a quiz or deck numbered 6.2 for chapters 6.2A and 6.2B).
‡ A combined deck that covers more than one chapter.
A chapter that is missing from a certification’s mapping file shows no tag for that certification: the NVIDIA file omits 4.1 and 10.6, and the other four omit 1.8, 9.16, and 9.17.
Chapter summaries
Summaries are excerpted from Study_Plan.md, which also lists each chapter’s key concepts and self-check questions.
4.1. AI Agent Deployment and Scaling
This chapter introduces the essential infrastructure and operational practices for deploying and scaling multi-agent systems in production, covering message queue architectures, vector database selection, observability patterns, API gateway implementations, MLOps for agentic systems, and CI/CD pipeline automation.
Videos (9)
4.2. Deployment & Scaling
Chapter 4.2 details deployment patterns for agentic systems, examining microservices and serverless approaches, message queue architecture selection, vector database deployment options, observability implementation, and CI/CD pipeline construction. The chapter provides production-ready guidance for scaling systems while maintaining reliability through progressive deployment and comprehensive monitoring.
Videos (1)
Code examples (4 files)
4.3. Container Orchestration and Edge Deployment
Chapter 4.3 covers Kubernetes orchestration for production multi-agent deployments and edge model optimization strategies. The chapter explains how Kubernetes automates deployment, scaling, and healing of containerized agents while covering model optimization techniques (quantization, pruning, distillation) that enable efficient edge deployment on resource-constrained devices.
Videos (7)
Code examples (7 files)
01_autoscaling_metrics_monitoring.py02_weighted_load_balancer.py03_rag_caching_configuration.pyjetson_deployment_runtime_code_04_jetson_deployment_runtime.pymodel_pruning_code_02_model_pruning.pyquantization_optimization_code_01_quantization_optimization.pytensorrt_object_detector_code_03_tensorrt_object_detector.py
4.4. Performance Profiling and Optimization
Performance profiling represents a critical but often-neglected step between deployment and production stability. AI agent systems introduce unique challenges compared to traditional inference workloads because their multi-stage execution pattern creates bottlenecks distributed across components that simple metrics cannot reveal. Measurement-driven optimization transforms deployment from one-time event into continuous cycle of improvement, replacing assumptions with data to guide effort toward high-impact optimizations.
Videos (7)
Code examples (12 files)
08_react_agent_inference.py11_quantization_baseline_fp16.py12_int8_quantization_evaluation.py13_accuracy_validation_quantization.py14_attention_profiling_baseline.py15_flash_attention_optimization.py16_paged_attention_batch_scaling.py17_optimized_throughput_measurement.py18_predictability_analysis_speculative_decoding.py19_speculative_decoding_deployment.py20_speculative_output_quality_validation.py21_mlflow_registration_artifacts.py
4.5. NVIDIA NIM and Triton Inference Server
NVIDIA NIM represents a paradigm shift in LLM deployment by collapsing the months-long gap between “agent works locally” and “agent serves production traffic” through pre-optimized containerized microservices. NIM bundles a complete, enterprise-grade inference stack while Triton serves as a unified multi-framework serving platform, enabling production-quality deployments without extensive optimization expertise.
Videos (5)
4.6. TensorRT-LLM and NVIDIA Fleet Command
TensorRT-LLM addresses fundamental inference challenges through optimization pipeline orchestrating multiple complementary optimizations achieving 3-8x speedup while reducing memory 50-75%. Fleet Command enables orchestration of edge AI deployments at scale through hybrid-cloud architecture, one-touch provisioning, and zero-trust security, transforming edge deployment from operational burden to managed platform.
Videos (7)
4.7. Scaling Strategies
Horizontal scaling addresses capacity expansion through creating multiple agent instances operating in parallel, enabling nearly linear capacity improvements. Strategic scaling requires effective load balancing, sophisticated batching decisions, multi-tier caching architectures, and cost optimization while maintaining high availability across distributed infrastructure.
Videos (3)
Other code examples
These files are named for a chapter number that is not in the current chapter list (an older numbering), so they are not attached to a chapter above.
22 files
Part_04_Chapter_4.1C_01_cuda_apt_installation.shPart_04_Chapter_4.1C_02_nsys_version_verification.shPart_04_Chapter_4.1C_03_nsight_docker_pull.shPart_04_Chapter_4.1C_04_nsight_workspace_setup.shPart_04_Chapter_4.1C_05_docker_run_nsight_container.shPart_04_Chapter_4.1C_06_nvidia_smi_gpu_verification.shPart_04_Chapter_4.1C_07_nsys_profile_command.shPart_04_Chapter_4.1C_09_react_profiling_execution.shPart_04_Chapter_4.1C_10_nsys_gui_launch.shPart_04_Chapter_4.1C_22_mlflow_production_transition.shPart_04_Chapter_4.1C_23_kustomization_staging_overlay.yamlPart_04_Chapter_4.1C_24_kustomization_production_overlay.yamlPart_04_Chapter_4.1C_25_argo_rollout_canary_deployment.yamlPart_04_Chapter_4.1C_26_analysis_template_success_rate_check.yamlPart_04_Chapter_4.2A_01_langchain_nim_configuration.pyPart_04_Chapter_4.2A_02_bert_torchscript_conversion.pyPart_04_Chapter_4.2A_03_triton_sentiment_realtime_client.pyPart_04_Chapter_4.2A_04_triton_sentiment_batch_processing.pyPart_04_Chapter_4.2B_01_gpt2_baseline_performance_measurement.pyPart_04_Chapter_4.2B_02_tensorrt_llm_fp16_optimization.pyPart_04_Chapter_4.2B_03_int8_quantization_calibration.pyPart_04_Chapter_4.2B_04_production_kv_cache_configuration.py
Labs
No finished lab exists for this Part yet. These legacy example files are prose excerpts with embedded code, kept as source material; they do not count as lab coverage. See Labs.
