Architecture profiles
Named, coherent configurations (ISO/IEC/IEEE 42010-style views) that select components for a class of use cases.
Accuracy-critical FP16-baseline serving
When to choose. Choose for medical, legal or financial applications; start at FP16 and evaluate FP8 only on Hopper with acceptable benchmark accuracy.
Engine Builder Engine Build Configuration FP16 Inference Engine Evaluation Harness Evaluation Baseline Benchmark Suite Manifest Paged KV Cache Allocator Inference Server
Alternative to: Throughput-critical quantized LLM serving
Adaptive RASC with confidence-based escalation
When to choose. Choose for cost-constrained, high-stakes production (medical coding, compliance review, customer support) where sample budgets must adapt to difficulty, reasoning quality should weight votes, and low-agreement cases must go to humans with auditable reasoning.
Reasoning Strategy Router Query Complexity Classifier Adaptive Sample Allocator Self-Consistency Sampling Policy Reasoning Path Sampler LLM Inference Service Final Answer Extractor Reasoning Path Quality Classifier Quality-Weighted Vote Aggregator Rationale Selector Explanation Presenter Confidence Estimator Confidence Gate Escalation Threshold Policy Human Specialist Audit Log Store
Alternative to: CoT + Self-Consistency (majority vote)
Sources: Ch5.3
Adaptive retrieval RAG (retrieve-vs-generate routing)
When to choose. Choose when a substantial share of queries (e.g., 20-40%) are answerable from parametric knowledge and retrieval cost or latency matters.
RAG Query Orchestrator Query-Type Retrieval Router Retrieval Routing Rule Set Parametric Answer Generator Retriever abstract Question Decomposer Parallel Sub-Query Retrieval Controller Multi-Hop Answer Synthesizer Answer Synthesizer abstract LLM Inference Service
Sources: Ch6.6
Adaptive rule learning (explainable adaptation)
When to choose. Choose when decisions must remain explainable but patterns evolve (e.g., fraud detection, clinical guideline drift, pricing), so learned rules are proposed, validated and expert-approved before entering the rule base.
Rule-Based Decision Engine Symbolic Logic Engine Production Rule Base Rule Firing Trace Audit Log Store Rule Outcome Monitor Misclassified Case Store Labeled Decision Case Dataset Rule Learner abstract Inductive Rule Learner Case-Based Rule Refiner Candidate Rule Queue Rule Validation Orchestrator Rule Consistency Checker Rule Quality Scorer Rule Acceptance Threshold Policy Holdout Evaluation Set A/B Test Traffic Splitter Domain Rule Expert Rule Base Updater Rejected Rule Log
Sources: Ch5.11
Adaptive utility-based recommender
When to choose. Choose when a personalization agent must balance relevance against diversity/novelty and adapt trade-off weights per user and context from observed behaviour.
Utility-Based Decision Maker Utility Function Specification Outcome Probability Estimator Contextual Weight Adapter Revealed Preference Learner User Feedback Store User Activity History Store Pareto Frontier Optimizer Explanation Presenter
Sources: Ch5.10
Agent CI/CD with progressive quality gates and canary release
When to choose. Choose for agent systems committed to multiple times daily that need fast feedback yet must block behavioural regressions, vulnerable images and unstable builds before full production exposure.
Continuous Integration Runner Evaluation Workflow Definition Static Code Analyzer Static Security Scanner Agent Test Runner Unit Test Suite Integration Test Suite Evaluation Harness Evaluation Dataset Chain-of-Thought Judge Regression Gate Regression Threshold Policy Container Image Builder Container and Model Artifact Registry Container Image Scanner Vulnerability Gate Policy Model and Agent Release Registry Staging Environment Smoke Tester SLO Monitor Release Approver Rollout Manager abstract Experiment Guardrail Monitor Online Evaluator Statistical Comparator Deployment Notifier
Sources: Ch4.2
Agent-driven tool chaining
When to choose. Choose for open-ended problems where the tool sequence depends on intermediate results and cannot be predicted in advance.
ReAct Agent Controller Reasoning Engine Native Function-Calling API abstract Tool Call Dispatcher abstract Sequential Tool Dispatcher Parallel Tool Dispatcher Tool Executor Tool Result Transformer Working Memory Buffer abstract Tool Schema Tool Registry
Sources: Ch2.6
Agent-local state
When to choose. Choose when individual workflows complete in seconds to minutes and need not be distributed across instances.
State-Graph Orchestrator Agent State Schema Working Memory Buffer abstract In-Memory State Store
Alternative to: Distributed state for horizontally scaled agents
Sources: Ch1.6
Balanced-performance agent configuration
When to choose. Choose as general-purpose production default: large model, temperature 0.3, compressed 5K context, 8 iterations (~91%, 3.5s, $0.06).
Large Language Model Tier Inference Serving Configuration abstract Summarizing History Compressor Context Compressor Iteration Limit Policy
Alternative to: Latency-optimized agent configuration, Cost-optimized agent configuration
Sources: Ch3.4
Batch incremental multi-source knowledge ETL
When to choose. Choose for enterprise knowledge bases needing multi-source integration, quality validation and deduplication where latency tolerance exceeds minutes; avoid for sub-second real-time freshness, small single-source applications, on-demand embedding of dynamic queries, or vector stores without batch operations.
Pipeline Scheduler Batch Extraction Policy Incremental Refresh Policy Ingestion Pipeline Orchestrator ETL Pipeline Configuration Incremental Watermark Store Source Change Detector Source Extractor Interface Source Record Extractor Paginated API Extractor File Store Extractor Document Quality Filter Text Normalizer Exact Hash Deduplicator Overlapping Window Chunker Semantic Boundary Chunker Chunk Metadata Extractor Text Embedding Service Chunk Schema Validator Vector Batch Ingestor Vector Index Builder IVF Index Configuration Self-Managed Vector Index Store CPU Dataframe Engine Dead Letter Queue Metrics Collector Alert Manager
Alternative to: GPU-accelerated web-scale data curation
Bidirectional interactive streaming agent
When to choose. Choose for agents expecting mid-stream user feedback or interruption, collaborative multi-user sessions, or multi-agent outputs streamed over one connection.
Conversational (Chat) Interface WebSocket Channel Response Streamer Stream Connection Manager Stream Interrupt Handler Conversation State Store abstract Agent Controller abstract Worker Agent abstract Explanation Presenter Native Function-Calling API abstract LLM Inference Service Metrics Collector Trace Collector
Sources: Ch2.9
Bottom-up (learned) value alignment
When to choose. Choose when contextual nuance and adaptation to unanticipated scenarios matter more than verifiable guarantees and large volumes of high-quality feedback are available; values remain implicit and opaque.
Preference Annotator abstract Preference Dataset Reward Model Trainer Reward Model RLHF Policy Optimizer Inverse Reward Learner Revealed Preference Learner Fine-Tuned Agent Model
Alternative to: Top-down (rule/principle-based) value alignment, Hybrid value alignment (hard constraints + learned nuance + continuous monitoring)
Sources: Ch9.6
Centralized multi-agent orchestration
When to choose. Choose for well-defined sequential workflows with stable task structures where predictability and centralised state outweigh scalability (up to about a dozen workers).
Supervisor Agent Worker Agent abstract Static Delegation Interface Static Rule Task Router Task State Ledger Trace Collector
Alternative to: Decentralized (peer-to-peer) multi-agent orchestration, Hierarchical multi-agent orchestration, Federated multi-agent orchestration
Sources: Ch1.3
Centrally managed edge AI fleet
When to choose. Choose for large-scale (50+ sites), geographically distributed NVIDIA edge deployments without local technical staff, with frequent model updates, high-security or mission-critical availability needs; avoid for <10 locations, cloud-only workloads or rarely updated static models.
Edge Fleet Manager Edge Provisioning Service Edge Device Agent Edge Application Definition Edge Update Orchestrator Staged Rollout Policy Edge GPU Device Cluster Consensus Coordinator GPU Partition GPU Partition Layout abstract Container and Model Artifact Registry Certificate Authority Key Management Service Encrypted Model Volume Secure Boot Verifier Inference Server Optimized Inference Engine abstract Container Health Prober Alert Manager Alert Rule Set Metrics Collector
Alternative to: Pre-optimized LLM microservice on Kubernetes
Sources: Ch4.6
Clinical decision support output filtering
When to choose. Choose for healthcare or medical-device agents where outputs affect patient safety and HIPAA/FDA obligations apply: prioritise recall, fact-check clinical claims and require human approval for high-risk actions.
Output Risk Stratifier Output Rail PII Redactor Pattern PII Detector PII Pattern Library Fact Checking Rail abstract Bias Evaluator Toxicity Classifier Moderation Triage Router Content Moderator Action Policy Engine Action Risk Tier Policy Approval Gateway Human Approver Audit Log Store
Sources: Ch9.1
Competitive (game-theoretic) multi-agent system
When to choose. Choose for strategic scenarios with conflicting interests such as resource scheduling, marketplaces, adversarial security testing, or economic simulation.
Strategic Agent Auction Task Allocator Collusion Monitor
Alternative to: Swarm intelligence multi-agent system
Sources: Ch1.3
Compiled, quantized LLM engine behind a multi-framework server on Kubernetes
When to choose. Choose for long-running, high-throughput or latency-sensitive production serving of >7B-parameter models where inference cost dominates; avoid for rapid prototyping, research iteration, <1B models or low-traffic applications.
Engine Builder Model Quantizer Quantization Calibrator Quantization Calibration Dataset Optimized Inference Engine abstract KV Cache Manager Speculative Decoder In-Flight Batch Scheduler LLM Generation Backend Inference Server Model Repository Continuous Integration Runner Evaluation Harness Inference Performance Analyzer Inference Metrics Endpoint Metrics Collector Custom Metrics Adapter Metric-Driven Autoscaler GPU Node abstract
Constitutional AI alignment (SL-CAI + RLAIF)
When to choose. Choose when alignment must scale without extensive human labeling of harmful content and values must be explicit and inspectable by regulators and stakeholders.
Constitution abstract Constitution Authoring Board Red-Team Prompt Dataset Critique-Revision Generator Self-Reflection Critic Critique Rubric Critique-Revision Dataset Fine-Tuning Pipeline abstract Candidate Response Sampler AI-Feedback Preference Labeler Preference Dataset Reward Model Trainer Reward Model Constitutional Reward Scorer RLHF Policy Optimizer Constitutionally Aligned Model Foundation LLM abstract LLM Inference Service Training Pipeline Orchestrator Principle Adherence Evaluator
Alternative to: Human-feedback RLHF alignment, Hybrid Constitutional AI then RLHF refinement
Sources: Ch9.5
Constitutional defense-in-depth runtime
When to choose. Choose for production deployments of aligned models, since no single alignment method eliminates harmful outputs; combine training-time alignment with runtime rails, moderation, human escalation and monitoring.
Constitutionally Aligned Model LLM Inference Service Agent Controller abstract Guardrail Orchestrator Guardrail Policy Canonical Form Matcher Input Rail Dialog Rail Retrieval Rail Execution Rail Output Rail Third-Party Moderation Service Violation Response Policy abstract Escalation Handler Human Specialist Guardrail Violation Monitor Adversarial Robustness Evaluator Alignment Drift Monitor Audit Log Store
Containerised microservices agent deployment
When to choose. Choose when components have fundamentally different scaling needs (e.g., CPU retrieval vs GPU generation), multiple teams deploy independently, latency must be consistent with warm instances, and the organisation already operates Kubernetes.
Agent API Gateway Layer-7 Load Balancer Identity Provider Rate Limiter Container Orchestrator Metric-Driven Autoscaler Autoscaling Policy Container Health Prober Liveness Endpoint Intent Router Knowledge Retrieval Agent Answer Synthesizer abstract Escalation Handler Model Router Model Routing Policy Message Queue Circuit Breaker Retry Handler Trace Collector Log Aggregator Metrics Collector Alert Manager GPU Node abstract
Alternative to: Serverless event-driven agents
Sources: Ch4.2
Continual learning with experience replay
When to choose. Choose when a learned policy must acquire new skills in production without catastrophic forgetting of earlier ones.
Experience Replay Buffer Continual Learning Trainer Policy Network Event-Based Episode Encoder
Sources: Ch5.7
Continuous compliance automation
When to choose. Choose when compliance must be a continuous, embedded process with automated monitoring, enforcement, alerting and reporting, keeping humans for judgment.
Metrics Collector Log Aggregator Continuous Compliance Monitor Action Policy Engine Bias Evaluator Alert Manager Incident Manager Compliance Dashboard Compliance Report Generator Compliance Evidence Repository Compliance Gate Continuous Integration Runner
Continuous delivery with human release gate
When to choose. Choose when a human must approve promotion from staging to production after automated validation.
Continuous Integration Runner Static Code Analyzer Evaluation Harness Load Test Runner Regression Gate Container Image Builder Container Image Scanner Container and Model Artifact Registry Configuration Repository Staging Environment Release Approver Rollout Manager abstract
Alternative to: Fully automated continuous deployment
Sources: Ch4.1
Conversation-driven multi-agent collaboration
When to choose. Choose for multi-agent workflows where flexible dialogue-driven collaboration and readable transcripts aid development, exploration and human oversight; not for deterministic, auditable production decisions.
Conversational Agent Coordinator Worker Agent abstract Human Proxy Agent Code Execution Runner Conversation State Store abstract Dual-Agent Critic Shared Message History Termination Checker Iteration Limit Policy Execution Sandbox abstract LLM Inference Service Human Approver
Alternative to: Role-based hierarchical crew, Plugin-routed enterprise agent platform, Role-based sequential team, Role-based hierarchical team with manager quality gates
Cooperative iterative neural-symbolic refinement
When to choose. Choose when symbolic knowledge must direct further neural analysis to resolve ambiguity and iterative refinement justifies extra cost and latency (e.g., circuit-board defect diagnosis).
Iterative Refinement Coordinator Perception Interpreter Neural Perception Model Symbolic Logic Engine Iteration Limit Policy Termination Checker
Alternative to: Sequential neural-to-symbolic pipeline, Parallel paradigms with result fusion, Embedded neural modules within a symbolic program
Sources: Ch5.13
Coordinator-managed working-memory budgets
When to choose. Choose when multiple worker agents compete for limited context capacity and global allocation should override individual over-consumption, accepting coordinator overhead and a potential single point of failure.
Agent Context Budget Coordinator Worker Agent abstract
Sources: Ch5.9
Cost-optimised routed agent (router-first, slim context)
When to choose. Choose for high-volume workloads with power-law query complexity (most queries simple) where token and API costs threaten viability, e.g., loan pre-screening or omnichannel support.
Model Router Query Complexity Classifier Small Language Model Tier Large Language Model Tier LLM Inference Service KV Cache Manager Dynamic Tool Loader Tool Schema Trajectory Pruner Summarizing History Compressor Tool Result Cache Request Batcher Parallel Agent Coordinator Token Budget Enforcer Token Budget Policy Token Cost Meter Execution Profiler Evaluation Baseline Quality Drift Detector Alert Manager
Sources: Ch3.10
Cost-optimized agent configuration
When to choose. Choose for high-volume simple queries where ~84% accuracy suffices and per-query cost must be minimal (~$0.01).
Small Language Model Tier Inference Serving Configuration abstract Context Compressor Iteration Limit Policy
Sources: Ch3.4
Cost-optimized conversational agent (token economics)
When to choose. Choose for high-volume, multi-turn conversational agents with large static context (system prompt, tool definitions, user profile) whose token cost must fall without degrading quality; applied progressively (caching, retrieval, output limits, routing) with quality validation after each step.
Token Cost Meter Cost Rate Card Cost Attribution Aggregator Cost Reporting Dashboard Cacheable Prefix Prompt Layout KV Cache Manager Prompt Context Builder Vector Retriever abstract Output Token Limit Policy Output Format Specification Model Router Rule-Based Complexity Classifier Small Language Model Tier Large Language Model Tier Regression Gate
Sources: Ch8.3
Cost-optimized inference serving
When to choose. Choose under budget constraints or at large concurrent scale, accepting ~5% quality loss for 40-60% infrastructure savings.
Evaluation Harness Evaluation Dataset Small Language Model Tier Model Quantizer INT8 Quantized Engine Inference Server Metric-Driven Autoscaler Autoscaling Policy
Alternative to: Throughput-optimized batch inference serving, Latency-optimized interactive inference serving
Sources: Ch7.2
Cost-sensitive LLM serving
When to choose. Choose when cost per inference is the priority and some quality and latency can be traded away.
Small Language Model Tier Model Quantizer INT4 Quantized Engine Request Batcher Response Cache abstract Spot GPU Node Metric-Driven Autoscaler Scheduled Scaler Token Cost Meter
Alternative to: Quality-critical LLM serving
CoT + Self-Consistency (majority vote)
When to choose. Choose for complex multi-step reasoning with definable correct answers (math, legal analysis, medical diagnosis) where base accuracy is insufficient and errors are path-specific; offers the best cost-accuracy trade-off without tree search or graph refinement.
Chain-of-Thought Prompt abstract Reasoning Path Sampler Self-Consistency Sampling Policy LLM Inference Service Final Answer Extractor Majority Vote Aggregator Confidence Estimator
Alternative to: Adaptive RASC with confidence-based escalation
Sources: Ch5.3
Curated multi-pass long-source analysis
When to choose. Choose when source material (large codebases, long contracts, multi-paper corpora) approaches or exceeds the context window or suffers lost-in-the-middle effects, and sequential-pass latency is acceptable.
Prompt Context Builder Working Memory Buffer abstract Vector Retriever abstract Reranker abstract Context Compressor Context Assembler abstract Reasoning Consistency Checker Reasoning Engine Episodic Memory Store Memory Retriever Context Budget Allocator Code Dependency Analyzer LLM Inference Service
Sources: Ch5.9
Customer-service transcript curation for support-agent training
When to choose. Choose when curating ASR-produced call transcripts to train a support agent: tolerate higher perplexity, use higher deduplication thresholds, and classify out sales/survey calls.
Speech Transcriber Partitioned Dataset Reader Document Quality Filter Perplexity Filter Exact Hash Deduplicator Near-Duplicate Detector Domain Relevance Classifier Document PII Redactor Data Curator Curated Training Corpus Fine-Tuning Pipeline abstract
Alternative to: GPU-accelerated web-corpus curation for a domain agent
Sources: Ch7.5
Decentralized (peer-to-peer) multi-agent orchestration
When to choose. Choose for large-scale systems where central coordination is infeasible, agent populations change frequently, and resilience to individual agent failure is critical.
Worker Agent abstract Agent Capability Registry Capability-Matching Allocator Discoverable Delegation Interface Agent Message Bus abstract Trace Collector
Alternative to: Centralized multi-agent orchestration, Hierarchical multi-agent orchestration, Federated multi-agent orchestration
Sources: Ch1.3
Decomposed RAG (query decomposition + parallel retrieval + synthesis)
When to choose. Choose for multi-part or comparison questions with several distinct information needs, when measured accuracy gains justify two extra LLM calls.
RAG Query Orchestrator Query Complexity Classifier Question Decomposer Decomposition Prompt Template Parallel Sub-Query Retrieval Controller Retriever abstract Multi-Hop Answer Synthesizer Synthesis Prompt Template LLM Inference Service
Alternative to: Standard single-query RAG
Sources: Ch6.6
Dedicated full-GPU single-tenant serving
When to choose. Choose for 70B+ models or one high-concurrency, latency-critical agent needing full memory bandwidth and large continuous batches.
Unpartitioned GPU Layout Dedicated GPU Device Inference Server In-Flight Batch Scheduler Paged KV Cache Allocator
Alternative to: Multi-tenant SaaS on uniform GPU partitions, Multi-tier platform on mixed GPU partitions
Sources: Ch7.6
Defense-in-depth reasoning verification
When to choose. Choose for high-stakes domains (finance, healthcare, legal) where plausible but wrong reasoning is unacceptable and automated verification must be layered with human review.
Reasoning Verifier abstract Fine-Tuned Step Verifier Business Rule Logic Verifier Citation Verifier Reasoning Consistency Checker Verifier Ensemble Aggregator Self-Reflection Critic Confidence Estimator Confidence Gate Human Specialist Trace Annotation Console
Sources: Ch3.6
Dependency-governed concurrent multi-agent workflow
When to choose. Choose when several specialist agents share workflow state and run many concurrent sessions in production (e.g., 50+), where deadlock, out-of-order execution and lost updates emerge that low-concurrency testing does not reveal.
Agent Dependency Graph Dependency Cycle Validator Phase Barrier Executor Worker Agent abstract Shared Blackboard Store Optimistic Lock Controller State Merge Policy Trace Context Propagator Trace Exporter Telemetry Gateway Trace Store Evaluation Failure Analyzer Bottleneck Analyzer
Sources: Ch8.2A
Design-time Pareto analysis with runtime scalarized utility
When to choose. Choose when objectives are incommensurable or stakeholders dispute weights: compute Pareto frontiers at design time, let stakeholders choose trade-offs, then deploy a scalarized utility for real-time operation.
Pareto Frontier Optimizer Pareto Frontier Set Explanation Presenter Decision Stakeholder Decision Sensitivity Analyzer Utility Function Specification Utility-Based Decision Maker Evaluation Harness Decision Scenario Simulator
Sources: Ch5.10
Deterministic auditable function-calling agent
When to choose. Choose when regulation requires identical decisions for identical inputs and clear decision trails, e.g., loan application processing calling credit, income and fraud services.
Direct Tool-Calling Controller Tool Schema Tool Executor External Service API Audit Log Store Native Function-Calling API abstract
Alternative to: Conversation-driven multi-agent collaboration
Development direct-export tracing
When to choose. Choose for development or small teams: agents export traces directly to a lightweight backend without a collector, inspected ad hoc.
Trace Collector Trace Schema Trace Exporter Trace Store Trace Visualizer Workflow Execution Debugger Agent Developer
Alternative to: Production centralized tracing
Sources: Ch3.6
Distributed billion-scale vector search
When to choose. Choose for enterprise deployments with billions of vectors needing horizontal scaling, high availability and GPU acceleration, operated by teams experienced with distributed systems and container orchestration.
Distributed Vector Index Store Cluster Consensus Coordinator Object Store Event Stream Log Container Orchestrator GPU Node abstract Vector Index Build Configuration abstract Approximate Vector Search Retriever
Alternative to: Highly available replicated vector store cluster
Sources: Ch6.2A
Distributed state for horizontally scaled agents
When to choose. Choose for long-running workflows (hours or days) or high-throughput systems with hundreds of concurrent requests that need multiple agent instances.
State-Graph Orchestrator Agent State Schema Working Memory Buffer abstract Distributed Cache State Store Database State Store State Concurrency Controller abstract State Retention Policy
Alternative to: Agent-local state
Sources: Ch1.6
Document + chart analysis agent
When to choose. Choose when critical quantitative data resides in charts within uploaded or indexed documents (e.g., earnings reports).
Document Ingestor Chart Data Extractor Vision-Language Model abstract Inference Server Agent Controller abstract LLM Inference Service
Sources: Ch7.5
Documented content moderation at scale
When to choose. Choose when an AI system makes millions of daily moderation decisions across languages and jurisdictions and must remain transparent, contestable and auditable.
Toxicity Classifier Regional Moderation Policy Confidence Gate Human Specialist Audit Log Store Moderation Appeal Tracker Bias Evaluator Model Card System Card Transparency Report Incident Manager
Sources: Ch9.8
DPO offline preference alignment
When to choose. Choose for pure preference learning from static datasets, particularly for organizations with limited RL expertise or computational resources.
Foundation LLM abstract Instruction Demonstration Dataset Fine-Tuning Pipeline abstract Reference Policy Model Alignment Prompt Dataset Candidate Response Sampler Preference Annotation Console Pairwise Comparison Format Preference Annotator abstract Annotation Quality Monitor Preference Dataset Direct Preference Optimizer Preference Optimization Config Fine-Tuned Agent Model
Alternative to: PPO-based RLHF alignment pipeline
Sources: Ch10.3
Embedded neural modules within a symbolic program
When to choose. Choose when deep integration (e.g., knowledge-grounded question answering) provides capabilities that justify reduced modularity and dual-paradigm expertise.
Symbolic Program Executor Natural Language to Logic Translator Vector Retriever abstract Symbolic Logic Engine Logic Conclusion Verbalizer Domain Ontology Knowledge Graph Store abstract
Alternative to: Sequential neural-to-symbolic pipeline, Parallel paradigms with result fusion, Cooperative iterative neural-symbolic refinement
Sources: Ch5.13
Ensemble (complexity-routed) scaling
When to choose. Choose when query complexity varies and most queries can be answered by smaller models, to cut cost 40-60% while escalating complex queries.
Model Router Query Complexity Classifier Embedding Service abstract Small Language Model Tier Standard Language Model Tier Large Language Model Tier LLM Inference Service
Alternative to: Model sharding across GPUs
Sources: Ch1.8
Explainable decision support (agent as assistant)
When to choose. Choose when an agent pre-screens or recommends high-stakes decisions (diagnostic imaging, loan screening, legal case prioritization, borderline moderation) while a human expert makes the final call; low-risk integration that is easy to turn off (Ref10.06 assistant pattern).
Decision Engine abstract Confidence Estimator Confidence Calibrator Trace Collector Trace Store Explanation Presenter Layered Explanation View abstract Proactive Explanation Expander Explanation Method Selector Attribution Analyzer Counterfactual Explainer Precedent Case Retriever Confidence Indicator Decision Override Control Acknowledgment Friction Gate Feedback Collector abstract Audit Log Store Override Pattern Analyzer Human Approver
Fact-checking-emphasis guardrails
When to choose. Choose for educational apps that emphasize factual accuracy while relaxing dialog restrictions.
Guardrail Policy Input Rail Output Rail Fact Checking Rail abstract LLM Inference Service
Alternative to: Full six-rail defense-in-depth
Sources: Ch7.1A
Fairness-assured high-stakes decision agent
When to choose. Choose for agents making consequential decisions about individuals (healthcare access, lending, hiring, criminal justice, education) where disparate impact creates legal and ethical exposure; combines layered fairness rails, group-wise auditing, continuous fairness monitoring with SLOs, active debiasing, human override and consent-governed demographic data.
Agent Controller abstract Predictive Decision Model Inference Server Guardrail Orchestrator Input Rail Retrieval Rail Output Rail Execution Rail Output Bias Detector abstract Bias Evaluator Demographic Audit Dataset Fairness Threshold Policy Fairness Monitor Alert Manager Bias Mitigator abstract Risk Gate Human Approver Decision Appeal Service Audit Log Store Fairness Auditor Demographic Data Store Purpose-Based Access Controller Consent Registry Model Card Oversight Governance Committee
Federated multi-agent orchestration
When to choose. Choose for regulated or multi-stakeholder settings where autonomous domains (organisations, partners) must keep local governance and coordinate only through contracts.
Supervisor Agent Worker Agent abstract Agent Service API abstract Agent Message Contract Schema Registry
Alternative to: Centralized multi-agent orchestration, Decentralized (peer-to-peer) multi-agent orchestration, Hierarchical multi-agent orchestration
Sources: Ch1.3
Feedback-driven continuous improvement flywheel
When to choose. Choose for production agents with real users where automated metrics miss satisfaction, comprehension and efficiency: collect explicit and implicit feedback, analyse themes and root causes, augment tests, refine, validate offline and via A/B, redeploy.
Feedback Collector abstract Behavioral Signal Tracker User Feedback Store Sentiment Classifier Feedback Theme Clusterer Feedback Prioritizer Evaluation Failure Analyzer Evaluation Dataset Evaluation Harness Regression Gate A/B Test Traffic Splitter Experiment Guardrail Monitor Rollout Manager abstract
Sources: Ch3.2
Few-shot in-context adaptation
When to choose. Choose for specialized domains or novel formats where zero-shot fails, high-quality demonstrations are available, and tasks are pattern-recognizable (classification, structured extraction, format transformation); not when knowledge is missing or deep behavioural change is needed.
System Prompt Template Prompt Exemplar Set Similarity Exemplar Selector Prompt Context Builder Trajectory Harvester Data Curator LLM Judge abstract LLM Inference Service Foundation LLM abstract Evaluation Harness
Alternative to: Retrieval-grounded adaptation, Fine-tuned specialist agent
Sources: Ch3.5
Financial services compliance output filtering
When to choose. Choose for banking/fintech agents subject to FINRA, FCRA and fair lending rules: layer deny-list, semantic compliance, PII filtering, disclaimers, bias detection and full audit trails.
Guardrail Orchestrator Guardrail Policy LLM Self-Check Output Rail Domain Compliance Rail abstract Pattern Compliance Checker Semantic Compliance Classifier Disclaimer Injector PII Redactor Deny-List Content Filter Output Bias Detector abstract Template Response Generator Audit Log Store Trace Collector Compliance Officer
Sources: Ch9.1
Fine-tuned specialist agent
When to choose. Choose only when behavioural-consistency problems resist prompt engineering and RAG, the domain is well structured, and roughly 1,000+ (ideally 5,000-10,000) quality trajectories are available.
Agent Trajectory Dataset Synthetic Data Generator Domain Expert Annotator Data Curator Continued Pretrainer Domain Text Corpus LoRA Fine-Tuner abstract LoRA Adapter Fine-Tuned Agent Model Training Pipeline Orchestrator Evaluation Harness Bias Evaluator Online Evaluator Rollout Manager abstract Inference Server System Prompt Template
Sources: Ch3.5
Full agentic system sandboxing
When to choose. Choose for high-risk deployments needing maximum containment even if the agent is compromised via prompt injection; requires secret injection and mediated production APIs.
Agent Controller abstract Tool Executor Code Execution Runner Execution Sandbox abstract Just-in-Time Credential Broker Secrets Vault Pod Network Policy Runtime Security Policy Enforcer Sandbox Anomaly Detector
Alternative to: Individual code snippet sandboxing
Sources: Ch9.3
Full six-rail defense-in-depth
When to choose. Choose for high-security environments (healthcare, finance) that accept 50-150ms latency as compliance cost.
Guardrail Policy Input Rail Dialog Rail Retrieval Rail Execution Rail Output Rail Fact Checking Rail abstract LLM Inference Service
Alternative to: Input/output rails only, Fact-checking-emphasis guardrails
Sources: Ch7.1A
Fully automated continuous deployment
When to choose. Choose when automated validation is trusted to promote passing builds from staging to production without a human gate.
Continuous Integration Runner Static Code Analyzer Evaluation Harness Load Test Runner Regression Gate Container Image Builder Container Image Scanner Container and Model Artifact Registry Configuration Repository Staging Environment Canary Rollout Controller
Alternative to: Continuous delivery with human release gate
Sources: Ch4.1
GDPR Art. 32 layered security for personal data and AI
When to choose. Choose when handling sensitive personal data (financial, health); protection scales with data sensitivity and processing risk.
Data Classification Policy Data Classifier Field-Level Encryptor Key Management Service Authorization Policy Decision Point Identity Provider Approval Gateway Audit Log Store Access Anomaly Detector Security Analyst Penetration Tester Incident Manager Incident Escalation Policy Personal Data Breach Notifier Personal Data Breach Register Input Rail PII Redactor Training Data Lineage Store Model and Agent Release Registry Agent API Gateway
GDPR data-subject rights and consent management
When to choose. Choose for any system processing personal data of EU residents, regardless of organization size or location: granular consent, rights-request fulfilment and retention enforcement.
Data Subject Privacy Portal Privacy Notice Consent Manager Consent Registry Consent Enforcement Gate Lawful Basis Record abstract Records of Processing Register Data Subject Request Handler Data Subject Identity Verifier Data Subject Request Queue Erasure Orchestrator Personal Data Exporter Personal Data Store Backup Archive Store Personal Data Retention Policy Data Retention Enforcer Data Processing Agreement External Service API Audit Log Store Data Protection Officer
GDPR-compliant automated decision-making
When to choose. Choose when an AI model makes decisions with legal or similarly significant effects on individuals (credit, insurance claims, hiring): DPIA, minimized inputs, fairness monitoring, borderline human review, explanations and appeals.
DPIA Record Purpose-Bound Data Schema Predictive Decision Model Decision Engine abstract Confidence Gate Human Approver Bias Evaluator Fairness Threshold Policy Proxy Feature Detector Decision Factor Explainer Decision Appeal Service Human Specialist Authorization Policy Decision Point Audit Log Store Data Protection Officer
Governed clinical decision-support AI (e.g., patient deterioration prediction)
When to choose. Choose when an AI system augments high-stakes clinical judgment under healthcare regulation: clinicians retain final authority, predictions are explained, performance and fairness are monitored continuously, and the system is withdrawn on threshold breach.
Oversight Governance Committee Accountable AI Executive AI Governance Policy AI Impact Assessment Record Harm Risk Register Compliance Requirement Crosswalk Bias Evaluator Fairness Threshold Policy Data Drift Detector Continual Learning Trainer Regression Gate Holdout Evaluation Set Approval Gateway Human Approver Feedback Collector abstract Audit Log Store Incident Manager AI Incident Classification Policy Feature Flag Service
Sources: Ch9.8
GPU-accelerated multimodal production stack
When to choose. Choose for production multimodal agents needing high throughput and low latency on self-hosted GPUs, with model-agnostic APIs, GPU vector search, health probing, metrics and zero-downtime model updates.
Ingestion Pipeline Orchestrator Inference Server OpenAI-Compatible Inference API Engine Builder Model Quantizer Optimized Inference Engine abstract Foundation LLM abstract Vision-Language Model abstract Chart-to-Table Model Joint Multimodal Embedding Service GPU-Accelerated Vector Index Store Metadata-Filtered Retriever GPU Node abstract Container Orchestrator Container Health Prober Load Balancer abstract Metrics Collector Rollout Manager abstract
GPU-accelerated web-corpus curation for a domain agent
When to choose. Choose when building a pretraining or domain corpus from large noisy web scrapes (billions of documents) for a single-language domain agent.
Partitioned Dataset Reader Language Identification Filter Document Quality Filter Perplexity Filter Exact Hash Deduplicator Near-Duplicate Detector Domain Relevance Classifier Document PII Redactor Synthetic Data Generator Data Curator GPU-Accelerated Dataframe Engine Curated Training Corpus
Alternative to: Customer-service transcript curation for support-agent training
Sources: Ch7.5
GPU-accelerated web-scale data curation
When to choose. Choose when processing billions of documents or terabytes where CPU pipelines take days, for regular full reprocessing and fuzzy deduplication at scale, especially with existing GPU clusters; stay on CPU below ~10 million documents.
Ingestion Pipeline Orchestrator Full Refresh Policy GPU-Accelerated Dataframe Engine GPU Node abstract Document Quality Filter Text Normalizer Exact Hash Deduplicator Near-Duplicate Detector Text Embedding Service Vector Batch Ingestor
Alternative to: Batch incremental multi-source knowledge ETL
Sources: Ch6.3B
GPU-enabled container cluster with telemetry
When to choose. Choose for production container clusters running GPU-accelerated workloads that need automated GPU enablement and GPU observability.
Container Orchestrator Accelerator Operator GPU Device Plugin Node Capability Labeler Accelerator Container Runtime GPU Telemetry Exporter Metrics Collector Metrics Dashboard abstract Alert Manager GPU Node abstract
Graph-based iterative workflow agent
When to choose. Choose when the workflow combines sequential stages with iterative refinement (generate-test-regenerate, feedback loops), conditional branching, and complex structured state requiring custom merge logic.
State-Graph Orchestrator Workflow State Graph Agent State Schema Rule-Based Transition Router State Checkpoint Store abstract Reasoning Engine Output Verifier Code Execution Runner Feedback Collector abstract Failure Analyzer Iteration Limit Policy Workflow Execution Debugger Execution Sandbox abstract Working Memory Buffer abstract
Alternative to: Conversation-driven multi-agent collaboration, Role-based hierarchical crew, Plugin-routed enterprise agent platform
Graph-of-Thought synthesis reasoning module
When to choose. Choose when problems decompose into independent subproblems whose solutions must be merged and synthesis adds value beyond the best individual exploration.
Graph-of-Thought Controller Graph of Operations Thought Generator abstract Thought Aggregator Thought Refiner Thought State Evaluator abstract Graph Reasoning State Store Reasoning Graph Analyzer Procedural Memory Store Plan Template Extractor Reasoning Strategy Refiner State Checkpoint Store abstract LLM Inference Service
Alternative to: Tree-of-Thought deliberate reasoning module
Sources: Ch5.2
Graph-routed customer support agent
When to choose. Choose when inquiries follow a decision tree: classification determines which specialised branch handles the query, with multi-turn state and escalation to humans under sub-second latency.
State-Graph Orchestrator Workflow State Graph Agent State Schema Working Memory Buffer abstract Intent Router Rule-Based Transition Router Worker Agent abstract Escalation Handler Human Specialist System Prompt Template LLM Inference Service Foundation LLM abstract Optimized Inference Engine abstract Inference Serving Configuration abstract Trace Collector
Sources: Ch1.5B
Ground-to-text multimodal RAG with metadata
When to choose. Choose for information-dense visuals (financial reports, scientific charts, technical diagrams) requiring precise retrieval, accepting one-time preprocessing cost while keeping text retrieval infrastructure unchanged.
Ingestion Pipeline Orchestrator Document Ingestor Multimodal Content Router Heuristic Image Type Classifier VLM-based Image Type Classifier Image Captioner Chart Data Extractor Vision-Language Model abstract Chart-to-Table Model Semantic Boundary Chunker Text Embedding Service Vector Index Store abstract Source Media Store Multimodal Chunk Metadata Schema Grounded Text Retriever Multimodal Context Assembler Multimodal Answer Synthesizer Inference Server OpenAI-Compatible Inference API Throughput-Oriented Batching Config Latency-Oriented Batching Config
Alternative to: Unified embedding space multimodal RAG, Separate modality stores with cross-modal reranking
Guardrail-sandwiched ReAct RAG agent
When to choose. Choose when a retrieval-augmented agent must enforce safety and compliance on both user input and agent output.
Input Rail ReAct Agent Controller Retriever abstract LLM Inference Service OpenAI-Compatible Inference API Output Rail Tool Protocol Server
Sources: Ref7.14
Guardrails-wrapped self-hosted inference microservice
When to choose. Choose for production LLM applications needing optimized self-hosted inference plus layered runtime safety, with safety policies and inference engines updated independently.
Guardrail Orchestrator Guardrail Policy Input Rail Dialog Rail Retrieval Rail Output Rail Fact Checking Rail abstract OpenAI-Compatible Inference API Inference Server Inference Engine Selector Pre-compiled Engine Backend Portable LLM Runtime Backend Model-Specific Inference Image Metric-Driven Autoscaler Container Orchestrator
Sources: Ch7.1B
Heuristic fast path with rule-based escalation
When to choose. Choose when most cases are routine and must be decided in milliseconds while borderline, high-stakes cases still need careful multi-rule evaluation (e.g., loan approval).
Heuristic Decision Engine Heuristic Rule Set Rule-Based Decision Engine Symbolic Logic Engine Production Rule Base
Sources: Ch5.11
Hierarchical aggregation multi-agent
When to choose. Choose for very large decomposable tasks (e.g., multi-document synthesis) where no single agent should hold all raw inputs, accepting coordination overhead that grows with hierarchy depth.
Supervisor Agent Worker Agent abstract
Alternative to: Multi-agent shared context pool, Message passing with selective context sharing
Sources: Ch5.9
Hierarchical backbone with flat leaves
When to choose. Choose when high-level workflow structure benefits from HTN scaling but primitive-level operations need sophisticated local optimization (e.g., manufacturing batch sequencing plus machine scheduling).
HTN Planner Decomposition Method Library Task Network Local Constraint Scheduler Flat Planner Plan Executor
Alternative to: Progressive deepening planning
Sources: Ch5.4
Hierarchical multi-agent orchestration
When to choose. Choose for large systems that partition naturally into subsystems or mirror team structures, balancing central oversight with distributed execution.
Supervisor Agent Worker Agent abstract Agent Delegation Interface abstract Task State Ledger Trace Collector
Alternative to: Centralized multi-agent orchestration, Decentralized (peer-to-peer) multi-agent orchestration, Federated multi-agent orchestration
Sources: Ch1.3
Hierarchical multi-agent pool scaling
When to choose. Choose for complex workflows where specialised sub-agents with different load patterns justify independent pool scaling despite orchestration and monitoring overhead.
Supervisor Agent Worker Agent abstract Task Queue State Checkpoint Store abstract Metric-Driven Autoscaler Layer-7 Load Balancer Metrics Collector
Sources: Ch1.8
Hierarchical replanning
When to choose. Choose when execution outcomes can deviate from predictions (failures, unexpected effects, resource conflicts) and valid higher-level decisions should be preserved during recovery.
HTN Planner Task Network Plan Executor Plan Deviation Monitor Tool Error Classifier Retry Handler Replanner abstract
Sources: Ch5.4
Highly available replicated vector store cluster
When to choose. Choose when 99.9%+ uptime is required so single-node failures must not interrupt service; add shards and nodes when capacity, not only availability, must scale.
Self-Managed Vector Index Store Persistent Volume API Key Authenticator Vector Store gRPC API Gossip Membership Service Shard Replication Manager Shard Query Router Replication and Sharding Configuration Load Balancer abstract Readiness Endpoint Pod Network Policy Container Orchestrator Metrics Collector Alert Manager Vector Store Capacity Monitor
Sources: Ch6.2B
Human-feedback RLHF alignment
When to choose. Choose when capturing contextual nuance and diverse, empirical human preferences outweighs annotation cost, annotator exposure to harmful content and value opacity.
Preference Annotator abstract Preference Annotation Console Candidate Response Sampler Preference Dataset Reward Model Trainer Reward Model Preference Reward Scorer RLHF Policy Optimizer Foundation LLM abstract LLM Inference Service
Alternative to: Constitutional AI alignment (SL-CAI + RLAIF), Hybrid Constitutional AI then RLHF refinement
Sources: Ch9.5
Human-in-the-loop pre-execution approval
When to choose. Choose when consequential, often regulated decisions (coverage, large payments, production changes) must not execute without explicit human authorization, and request volumes fit planned reviewer capacity.
Agent Controller abstract Approval Proposal Builder Confidence Estimator Oversight Gate abstract Confidence Gate Risk Gate Exception Gate Escalation Threshold Policy Approval Gateway State-Graph Orchestrator State Checkpoint Store abstract Approval Request Store Approval Router Approval Authority Matrix Approver Notifier Approval Review Console Explanation Presenter Approval Outcome Router Approval Escalation Scheduler Escalation Chain Policy Approval SLA Policy Approval Timeout Fallback Policy abstract Action Policy Engine Human Approver Human Specialist Tool Executor Audit Log Store Oversight Performance Monitor Approval Pattern Auditor Adaptive Threshold Tuner Override Pattern Analyzer
Alternative to: Human-on-the-loop monitoring with reactive intervention
Human-in-the-loop synchronous approval
When to choose. Choose for high-consequence, irreversible, binary decisions (production deletions, large financial transfers, treatment changes, legal notices) where decision latency of hours is acceptable and volume is low.
Risk Gate Approval Gateway Approval Request Store Approval Router Approval Authority Matrix Approver Notifier Approval Queue API Approval Review Console Approval Escalation Scheduler Escalation Chain Policy Human Approver Audit Log Store
Alternative to: Human-on-the-loop asynchronous monitoring, Human-over-the-loop strategic governance
Sources: Ch9.2
Human-on-the-loop asynchronous monitoring
When to choose. Choose for medium-consequence reversible actions (content moderation, access changes, automated emails) at volumes where per-decision approval is impractical; unsuitable for irreversible actions.
Agent Controller abstract Audit Log Store Execution Monitor Console Human Supervisor Action Rollback Service
Alternative to: Human-in-the-loop synchronous approval, Human-over-the-loop strategic governance
Sources: Ch9.2
Human-on-the-loop monitoring with reactive intervention
When to choose. Choose for high-volume operations (thousands of transactions per hour) where per-action pre-approval is impractical: actions execute immediately under continuous observation and humans intervene after the fact when anomalies emerge.
Agent Controller abstract Decision Telemetry Event Schema Trace Collector Metrics Collector Log Aggregator Trace Store Time-Series Metrics Store Centralized Log Store Behavioral Baseline Builder Behavioral Baseline Agent Behavior Anomaly Detector Telemetry Change-Point Detector Agent Interaction Graph Analyzer Alert Manager Alert Rule Set Alert Consolidator Alert Suppression Rule Set Incident Manager Execution Monitor Console Trace Visualizer Change Impact Attributor Agent Intervention Controller Intervention Authority Policy Agent Circuit Breaker Guardrail Orchestrator Guardrail Violation Monitor Oversight Intensity Policy Human Supervisor Override Pattern Analyzer Audit Log Store
Alternative to: Human-in-the-loop pre-execution approval
Sources: Ch10.4
Human-over-the-loop strategic governance
When to choose. Choose for low-consequence routine actions at very high scale (report generation, retrieval, log analysis) where only aggregate patterns matter and detection delay until periodic review is acceptable.
Agent Controller abstract Audit Log Store Metrics Dashboard abstract Policy Adherence Evaluator Oversight Governance Committee
Alternative to: Human-in-the-loop synchronous approval, Human-on-the-loop asynchronous monitoring
Sources: Ch9.2
Hybrid 3D parallelism
When to choose. Choose for scaling very large models to thousands of GPUs: data parallel across nodes, tensor parallel within NVLink nodes, pipeline parallel for massive models.
Data Parallel Executor Tensor Parallel Executor Pipeline Parallel Executor Collective Communication Library GPU Switch Fabric Inter-Node Network Fabric GPU Node abstract
Sources: Ch7.1A
Hybrid Constitutional AI then RLHF refinement
When to choose. Choose (the text's recommended practice) to establish principled, scalable foundations via critique-revision and RLAIF, then refine contextual nuance with human-preference RLHF, alternating iteratively.
Constitution abstract Critique-Revision Generator Critique-Revision Dataset Fine-Tuning Pipeline abstract Candidate Response Sampler AI-Feedback Preference Labeler Preference Annotator abstract Preference Annotation Console Preference Dataset Reward Model Trainer Reward Model Constitutional Reward Scorer Preference Reward Scorer RLHF Policy Optimizer Constitutionally Aligned Model Training Pipeline Orchestrator Principle Adherence Evaluator Alignment Drift Monitor Direct Preference Optimizer Fine-Tuned Agent Model Content Safety Filter abstract
Alternative to: Constitutional AI alignment (SL-CAI + RLAIF), Human-feedback RLHF alignment
Hybrid cost-optimized customer service stack
When to choose. Choose for high-volume customer service with repetitive queries and stable prompts: combine prompt caching, complexity-based model routing, FAQ response caching and TTL tool-result caching.
KV Cache Manager Model Router Query Complexity Classifier Small Language Model Tier Large Language Model Tier Response Cache abstract Tool Result Cache Cache Policy
Sources: Ch3.4
Hybrid edge inference with cloud retraining
When to choose. Choose when time-critical decisions need edge inference (sub-50 ms, offline, privacy, bandwidth) while periodic cloud synchronization supports retraining, analytics and version control.
Edge Device abstract Edge Inference Runtime Edge-Optimized Model Edge Device Agent Edge Update Orchestrator Device Inventory Registry Model Quantizer Model Pruner Engine Builder Data Curator Training Pipeline Orchestrator Quality Drift Detector
Sources: Ch4.3
Hybrid RAG + fine-tuned model
When to choose. Choose when both current and domain knowledge are needed, highest accuracy is the priority and budget allows both fine-tuning and retrieval infrastructure (Ref6.01).
Domain Text Corpus Data Curator Fine-Tuning Pipeline abstract Fine-Tuned Agent Model Inference Server Retriever abstract Vector Index Store abstract Context Assembler abstract Answer Synthesizer abstract Knowledge Base Refresher Evaluation Harness Holdout Evaluation Set RAG Evaluator
Alternative to: Retrieval-grounded adaptation, Fine-tuned specialist agent
Sources: Ref6.01
Hybrid RAG + knowledge graph
When to choose. Choose when queries consistently need both semantic understanding and precise relationship tracking (financial compliance, fraud detection, research analysis); highest complexity (p95 ~180ms).
Document Chunker abstract Embedding Service abstract Vector Index Store abstract Vector Retriever abstract Entity Recognizer Entity Linker abstract Relation Extractor abstract Knowledge Graph Loader Property Graph Store Graph Retriever Hybrid Retriever abstract Context Assembler abstract Synthesis Prompt Template Answer Synthesizer abstract LLM Inference Service Subgraph Cache Circuit Breaker
Alternative to: Vector RAG only, Knowledge graph only (text-to-graph-query QA)
Sources: Ch1.7B
Hybrid task queue + event stream messaging
When to choose. Choose when a system needs both reliable task orchestration/request-response (broker queue) and high-throughput replayable event streaming for data pipelines.
Message Queue Event Stream Log Dead Letter Queue Event Retention Policy Worker Agent abstract
Sources: Ch4.1
Hybrid value alignment (hard constraints + learned nuance + continuous monitoring)
When to choose. Choose for consequential domains (healthcare, lending, autonomous vehicles): explicit rules for non-negotiable constraints combined with RLHF-learned contextual judgment, continuous drift monitoring and stakeholder feedback; the chapter's recommended approach.
Constitution abstract Principle Priority Policy Principle Conflict Register Operational Norm Set Formal Rule Specification Rule Constraint Filter Guardrail Policy Preference Annotator abstract Preference Dataset Reward Model Trainer Reward Model RLHF Policy Optimizer Fine-Tuned Agent Model Alignment Drift Monitor Principle Adherence Evaluator Red Team Tester User Complaint Intake Oversight Governance Committee Decision Stakeholder Bias Evaluator External Auditor
Alternative to: Top-down (rule/principle-based) value alignment, Bottom-up (learned) value alignment
Sources: Ch9.6
Hybrid vector + graph memory
When to choose. Choose when both broad semantic retrieval and precise structured reasoning are needed and justify maintaining two stores (e.g., medical diagnosis, legal research, literature synthesis, enterprise entity unification, episode timelines).
Semantic Memory Store Episodic Memory Store Vector Index Store abstract Knowledge Graph Store abstract Retrieval-Augmented Graph Retriever Graph-Constrained Vector Retriever Entity Recognizer Relation Extractor abstract Multi-Signal Entity Resolver Text Embedding Service
Alternative to: Vector-only memory, Knowledge-graph memory
Hybrid vector-plus-graph RAG
When to choose. Choose when answers need both rich semantic context from documents and precise quantitative or relational facts from a knowledge graph, grounding generation to reduce hallucination.
Text Embedding Service Parallel Fusion Retriever Vector Retriever abstract Vector Index Store abstract Text-to-Graph-Query Translator Graph Retriever Knowledge Graph Store abstract Context Assembler abstract Answer Synthesizer abstract LLM Inference Service
Sources: Ch5.13
Individual code snippet sandboxing
When to choose. Choose when generated code is the primary attack surface and the agent system itself is protected by input filtering and prompt isolation; simplest setup since only the sandbox service is configured.
Agent Controller abstract Sandbox Execution API Code Execution Runner Execution Sandbox abstract Static Security Scanner
Alternative to: Full agentic system sandboxing
Sources: Ch9.3
Input/output rails only
When to choose. Choose for consumer apps with non-state-modifying agents, where execution rails can be disabled.
Guardrail Policy Input Rail Output Rail LLM Inference Service
Alternative to: Full six-rail defense-in-depth
Sources: Ch7.1A
Integrated memory-perception agent
When to choose. Choose for stateful, personalised interactions where the agent must resolve implicit references to past interactions, learn from successful resolutions, and combine current perception with separate episodic, semantic and procedural memory.
Agent Controller abstract Perception Interpreter Memory Retriever Memory Consolidator Memory Write Validator Memory Lifecycle Manager Working Memory Buffer abstract State Checkpoint Store abstract Context Window Manager abstract Semantic Memory Store Episodic Memory Store Procedural Memory Store LLM Inference Service Embedding Service abstract
Sources: Ch1.4
Integrated NIST AI RMF + ISO/IEC 42001 governance program
When to choose. Choose when an organization needs both operationally agile AI risk management and a formal, certifiable management system with accountability and external validation, integrating sector regulations into one program.
Oversight Governance Committee Accountable AI Executive AI Governance Policy Risk Tolerance Policy AI Use Inventory AI Risk Tier Classifier AI System Context Record AI Impact Assessment Record Harm Risk Register Risk Scorer abstract Risk Monitor Risk Treatment Plan AI Control Catalog Compliance Requirement Crosswalk Control Effectiveness Evaluator Continuous Compliance Monitor Compliance Evidence Collector Compliance Evidence Repository Internal Auditor External Auditor Remediation Tracker Model Card Dataset Datasheet System Card Transparency Report Audit Log Store Incident Manager AI Incident Classification Policy
Alternative to: Proportionate small/medium-organization AI governance
Knowledge graph only (text-to-graph-query QA)
When to choose. Choose for relationship queries and multi-hop traversal over well-modeled entities (p95 ~80ms, medium complexity).
Document Chunker abstract Entity Recognizer Entity Linker abstract Relation Extractor abstract Knowledge Graph Loader Property Graph Store Graph Schema Introspector Graph Schema Text-to-Graph-Query Translator Graph Retriever Answer Synthesizer abstract LLM Inference Service
Alternative to: Vector RAG only, Hybrid RAG + knowledge graph
Knowledge-graph memory
When to choose. Choose when reasoning requires explicit relationships, multi-hop inference or guaranteed consistency and knowledge is already structured (e.g., supply chain planning, clinical decision support).
Semantic Memory Store Knowledge Graph Store abstract Knowledge Graph Rule Set Graph Rule Inferencer Graph Consistency Validator Graph Retriever Knowledge Graph Loader Subgraph Cache
Alternative to: Vector-only memory, Hybrid vector + graph memory
Latency-critical HITL approval (e.g., real-time fraud review)
When to choose. Choose when approvals must complete within seconds (real-time trading, emergency response, fraud prevention): dedicated reviewer pools, simplified criteria, fast-track escalation, pre-approved categories and one-click interfaces.
Confidence Gate Approval Gateway Pre-Approved Category Policy Approval Router Approver Notifier Approval Review Console Approval SLA Policy Oversight Performance Monitor Approval Escalation Scheduler Human Approver
Alternative to: Throughput-critical HITL approval (e.g., batch claims, bulk onboarding)
Sources: Ch10.4
Latency-critical LLM serving
When to choose. Choose when time to first token is the priority (interactive chat, code completion).
LLM Inference Service Speculative Decoder Speculative Draft Model abstract Latency-Oriented Batching Config Inference Queue Policy Latency-Optimized Deployment Profile
Alternative to: Throughput-critical LLM serving, Cost-sensitive LLM serving, Quality-critical LLM serving
Latency-critical low-batch serving
When to choose. Choose for premium single users with batch size 1-4, predictable sequence lengths (e.g., <512-token translation) and strict per-token latency SLAs, when GPU memory is abundant.
Engine Builder Engine Build Configuration Contiguous KV Cache Allocator Fused Block-wise Attention Kernel Inference Server
Alternative to: Throughput-critical quantized LLM serving
Sources: Ch7.4
Latency-critical templated agent
When to choose. Choose when latency is a hard workflow constraint (e.g., sub-2-second clinical documentation) and task types are recurring enough for pre-compiled context.
Model Router Small Language Model Tier Large Language Model Tier LLM Inference Service Precompiled Context Template Prompt Context Builder Conversation State Store abstract Response Streamer Token Cost Meter Metrics Collector
Sources: Ch3.10
Latency-optimized agent configuration
When to choose. Choose for interactive, real-time applications where sub-2-second response is mandatory (small model, compressed context, 5 iterations).
Small Language Model Tier Inference Serving Configuration abstract Context Compressor Iteration Limit Policy
Alternative to: Cost-optimized agent configuration
Sources: Ch3.4
Latency-optimized interactive inference serving
When to choose. Choose for interactive applications (chat, code completion, real-time moderation) needing sub-200ms to sub-300ms P95 latency.
Inference Server Latency-Oriented Batching Config Latency-Tuned Decoding Configuration Response Streamer Small Language Model Tier INT8 Quantized Engine GPU Node abstract
Alternative to: Throughput-optimized batch inference serving, Cost-optimized inference serving
Sources: Ch7.2
Layered hallucination defense (regulated domain)
When to choose. Choose for customer-facing agents in regulated or high-stakes domains (finance, healthcare, legal) where critical hallucinations require near-zero tolerance and human oversight.
Knowledge Base Auditor Knowledge Base Refresher Content Freshness Monitor Query Rewriter Dense-Sparse Hybrid Retriever Reranker abstract Retrieval Confidence Filter System Prompt Template Stepwise Reasoning Verifier Reasoning Consistency Checker Dual-Agent Critic Response Confidence Modulator Output Verifier Output Format Specification Token Uncertainty Scorer Citation Verifier LLM Judge abstract Human Evaluator Confidence Gate Human Approver Online Evaluator Evaluation Trace Sampler Quality Drift Detector Hallucination Threshold Policy Failure Taxonomy Evaluation Failure Analyzer Feedback Collector abstract
Sources: Ch3.10
Layered multi-environment benchmarking
When to choose. Choose when selecting models or reasoning architectures for heterogeneous workflows: screen on a general multi-environment benchmark, then domain benchmarks, then monitored pilots.
Benchmark Suite Manifest Benchmark Environment abstract Environment Snapshot Simulated User Agent Evaluation Harness State Outcome Scorer Trajectory Scorer Evaluation Score Aggregator Evaluation Protocol Statistical Comparator Evaluation Failure Analyzer
Sources: Ch3.2
Layered production RAG service
When to choose. Choose when a RAG prototype must meet production SLAs (e.g., P95 < 2 s, 99.9% uptime, < 0.1% errors, per-query cost targets) at millions of documents and concurrent load.
Ingestion Pipeline Orchestrator Source Change Detector Text Embedding Service Vector Index Store abstract Document Metadata Store Lexical Index Store Knowledge Base Backup Service RAG Query Orchestrator Dense-Sparse Hybrid Retriever Reranker abstract Context Assembler abstract Answer Synthesizer abstract Citation Extractor LLM Inference Service Embedding Cache Retrieval Result Cache Distributed Response Cache Cache Invalidator Circuit Breaker Retry Handler Graceful Degradation Manager API Gateway Proxy abstract REST Agent API Load Balancer abstract Autoscaler abstract Metrics Collector Log Aggregator Trace Collector Alert Manager
Sources: Ch6.5
Layered-CoT multi-agent research synthesis
When to choose. Choose when a task requires coordinated expertise across several specialised domains, synthesis of contradictory findings, and memory of prior related work; avoid for simple lookups or fixed procedural tasks.
Supervisor Agent Task Planner abstract Execution Plan abstract Agent Capability Registry Parallel Agent Coordinator Worker Agent abstract ReAct Agent Controller Reasoning Engine Tool Executor External Service API Findings Synthesis Agent Working Memory Buffer abstract Memory Consolidator Memory Retriever Episodic Memory Store Semantic Memory Store Procedural Memory Store
Sources: Ch5.1
Layered-resilience multi-tool agent
When to choose. Choose for production agents depending on rate-limited external APIs and multiple LLM providers (e.g., financial research during peak load) that must keep partial functionality and an audit trail of degradation events through transient errors, provider outages and multi-component failures.
Retry Handler Retry Policy Model Router Fallback Chain Policy LLM Inference Service Fallback LLM Inference Service Response Cache abstract Graceful Degradation Manager Capability Tier Map Reasoning Engine Rule-Based Analyzer Embedded Critical Checker Dependency Health Monitor Circuit Breaker Circuit Breaker Policy Tool Result Cache External Service API Error Presenter Audit Log Store Metrics Collector Alert Manager Platform Operator
Sources: Ch2.8
Logic-augmented reasoning (Logic Agent)
When to choose. Choose for formal domains with clear inference rules (legal, mathematical proofs, regulatory compliance, security analysis) where logical validity is non-negotiable.
Reasoning Engine Natural Language to Logic Translator Symbolic Logic Engine Logic Rule Library Logic Conclusion Verbalizer
Alternative to: Structured Chain-of-Thought reasoning with layered verification
Sources: Ch3.9
Long multi-turn assistant with hierarchical history compression
When to choose. Choose for support-style conversations extending to tens of turns in which early turns establish facts (account type, attempted fixes) referenced later.
Prompt Context Builder Working Memory Buffer abstract Hierarchical History Compressor Conversation State Store abstract Compression Fidelity Validator Context Budget Allocator Context Allocation Policy Retriever abstract Reasoning Engine LLM Inference Service Memory Consolidator Episodic Memory Store
Sources: Ch5.9
Long-context, variable-length LLM serving
When to choose. Choose when inference is attention-dominated on long inputs and generation lengths vary widely (e.g., contract analysis), to raise throughput without extra GPUs.
Inference Server Optimized Inference Engine abstract Fused Block-wise Attention Kernel Paged KV Cache Allocator KV Cache Store KV Cache Manager Engine Builder Engine Build Configuration Inference Engine Profiler GPU Node abstract
Sources: Ch4.4
Memory-augmented ReAct agent with episodic memory
When to choose. Choose when an agent handles recurring situations across sessions (support, multi-turn debugging) and should reason from past episodes and trajectories rather than generic procedures.
ReAct Agent Controller Memory Retriever Prompt Context Builder Episodic Memory Store Significance-Based Episode Encoder Memory Consolidator Episode Summarizer Episode Pattern Abstractor Multi-Signal Relevance Ranker Vector Index Store abstract Text Embedding Service Procedural Memory Store
Sources: Ch5.7
Message passing with selective context sharing
When to choose. Choose when agents hand off sequential subtasks and task dependencies are understood well enough to identify the minimal context each receiver needs.
Worker Agent abstract Agent Message Bus abstract
Alternative to: Multi-agent shared context pool, Hierarchical aggregation multi-agent
Sources: Ch5.9
Minimal single-cluster agent deployment
When to choose. Choose for early or small deployments (e.g., ~100 internal users) before demand variability or observability gaps justify autoscaling, service mesh or multi-region.
Container Orchestrator Stateless Workload Controller Deployment Manifest Layer-4 Load Balancer Container Health Prober
Alternative to: Production multi-agent Kubernetes deployment
Sources: Ch4.3
Mixed-initiative human-agent collaboration
When to choose. Choose when humans and agents share a task and initiative must transfer between them based on risk, confidence and expertise, requiring explicit responsibility allocation, synchronized handoff context and mutual performance monitoring.
Mixed-Initiative Controller abstract Negotiated Initiative Controller Fixed Subtask Initiative Controller Subdialogue Initiative Controller Responsibility Matrix Autonomy Scope Adjuster Handoff Context Packager Escalation Handoff Package State Checkpoint Store abstract Override Rationale Log Override Pattern Analyzer Reviewer Fatigue Monitor Confidence Basis Explainer Interaction Style Adapter Engagement Estimator Negotiation Dialogue Manager Consensus Decision Protocol Human Specialist
Sources: Ch10.2
Model sharding across GPUs
When to choose. Choose only when the model exceeds single-GPU memory (>80 GB) after quantization and its quality gain justifies interconnect cost and complexity.
LLM Inference Service Large Language Model Tier GPU Node abstract
Alternative to: Ensemble (complexity-routed) scaling
Sources: Ch1.8
Multi-agent coordination over single-agent subflows
When to choose. Choose for production systems where multi-agent coordination handles the high-level workflow while complex subtasks (multimodal document processing, retrieval) need the stronger state management, refinement loops and checkpointing of single-agent graph or sequential workflows.
Multi-Agent Coordinator abstract Worker Agent abstract State-Graph Orchestrator State Checkpoint Store abstract Knowledge Retrieval Agent
Sources: Ch2.4
Multi-hop question answering agent
When to choose. Choose when answers require integrating evidence across multiple documents or source systems with auditable, source-attributed reasoning (research synthesis, enterprise knowledge, medical literature).
Multi-Hop Retrieval Controller Question Decomposer Query Rewriter Knowledge Source Router Retriever abstract Vector Index Store abstract Knowledge Graph Store abstract Supporting Fact Extractor Cross-Source Consistency Verifier Multi-Hop Answer Synthesizer Confidence Estimator Working Memory Buffer abstract Action Policy Engine Explanation Presenter Web Navigation Agent
Sources: Ch3.3
Multi-layer caching
When to choose. Choose when cost optimization justifies complexity, access patterns are diverse, and 80-90% hit rates confer advantage at large scale.
In-Process Response Cache Distributed Response Cache Semantic Cache KV Cache Manager Cache Policy Cache Invalidator Cache Warmer Embedding Service abstract
Sources: Ch1.8
Multi-layered hybrid replanning agent
When to choose. Choose for embodied or long-running agents in dynamic environments facing a mix of safety-critical, predictable, localized and global failures across millisecond-to-hour timescales.
Plan-and-Execute Controller Optimal Heuristic Search Planner Geometric Distance Heuristic State-Space Graph Contingency Planner Conditional Execution Plan Plan Executor Plan Deviation Monitor Discrepancy Significance Evaluator Replanning Strategy Router Reflexive Safety Controller Contingency Branch Activator Incremental Search Replanner Persistent Search Tree Store Complete Replanner Replanning Layer Arbiter Execution Failure History Store
Alternative to: Single-strategy reactive replanning agent
Sources: Ch5.6
Multi-model images in development, model-specific images in production
When to choose. Choose when an organization needs both innovation velocity (research/dev) and production reliability (SLA-bound services).
Multi-Model Inference Image Model-Specific Inference Image Model Integrity Validator Dynamic Model Loader Inference Server
Sources: Ch7.1B
Multi-tenant SaaS on uniform GPU partitions
When to choose. Choose when many tenants run similar-sized models (e.g., fine-tuned 7B INT8) with similar throughput and strict per-tenant SLAs.
Uniform GPU Partition Layout GPU Partition GPU Partition Manager Accelerator Operator GPU Device Plugin Inference Server INT8 Quantized Engine GPU Telemetry Exporter
Alternative to: Multi-tier platform on mixed GPU partitions, Dedicated full-GPU single-tenant serving, Time-sliced shared GPU
Sources: Ch7.6
Multi-tier caching for RAG agents
When to choose. Choose for RAG agents with repeated or structurally similar queries where reasoning templates, query embeddings and idempotent tool results can be reused.
Reasoning Chain Cache Plan Template Extractor Plan Template Adapter Embedding Cache Embedding Service abstract Vector Index Store abstract Tool Result Cache Cache Invalidator Cache Dependency Index Cache Policy Tool Executor
Sources: Ch4.7
Multi-tier platform on mixed GPU partitions
When to choose. Choose for tiered offerings with diverse model sizes (7B-70B), or when SLA-critical agents share a cluster with batch jobs on separate partitions.
Mixed GPU Partition Layout GPU Partition GPU Partition Manager GPU Partition Reconfiguration Policy Accelerator Operator GPU Device Plugin Inference Server GPU Telemetry Exporter Alert Rule Set Platform Operator
Alternative to: Multi-tenant SaaS on uniform GPU partitions, Dedicated full-GPU single-tenant serving
Sources: Ch7.6
Multimodal sensor perception
When to choose. Choose for agents in rich physical or multimodal environments that must fuse vision, LiDAR, audio, sensor and text inputs into a coherent world model.
Sensor Input Adapter Signal Preprocessor Sensor Stream Synchronizer Multimodal Fusion Engine Perception Interpreter World Model State Agent Controller abstract
Sources: Ch1.4
Neural-guided MCTS (AlphaGo/AlphaZero style)
When to choose. Choose when expert data exists or self-play can generate it and upfront GPU training cost is acceptable, to reach stronger decisions with far fewer simulations.
MCTS Planner MCTS Search Tree Store Search Tree Pruner Value Network Evaluator Value Network Action Prior Estimator Policy Network Policy/Value Network Trainer Expert Demonstration Dataset Environment Simulator MCTS Search Configuration Search Budget Policy GPU Node abstract
Sources: Ch5.5
Offline reasoning evaluation pipeline
When to choose. Choose during development for comprehensive reasoning quality assessment across extensive test suites without production latency constraints.
Trace Collector Reasoning Trace Schema Reference Reasoning Dataset Evaluation Harness Reasoning Chain Decomposer Entailment Step Validator Reasoning Consistency Checker Information Gain Scorer Goal Alignment Checker Reasoning Quality Scorer abstract LLM Judge abstract Human Evaluator Evaluator Calibrator Evaluation Result Analyzer Tool Fault Injector Quality Improvement Backlog
Sources: Ch3.9
Offline-online evaluation flywheel
When to choose. Choose for any production agent needing continuous improvement: offline baselines and CI regression gating, online monitoring on sampled traffic, and production failures fed back into versioned test sets.
Evaluation Harness Evaluation Dataset Continuous Integration Runner Task Success Evaluator abstract Policy Adherence Evaluator Simulated Web Environment Simulated User Agent Synthetic Scenario Generator Synthetic Dataset Online Evaluator Evaluation Sampling Policy abstract LLM Judge abstract Judge Adversarial Tester Human Evaluator Evaluation Rubric Feedback Collector abstract Behavioral Signal Tracker Agent Behavior Anomaly Detector Agent Version Experimenter abstract Evaluation Failure Analyzer Failure Case Curator Trace Store
Sources: Ch3.3
Offline-only continuous evaluation
When to choose. Choose for low-traffic agents (e.g., ~50 queries/day would need ~33 days per A/B test) or frequent small changes; supplement with manual monitoring of production metrics.
Continuous Integration Runner Evaluation Harness Evaluation Dataset Response Scorer abstract Experiment Tracker Evaluation Result Store Evaluation Baseline Statistical Comparator Regression Gate Regression Threshold Policy Evaluation Report Publisher Metrics Collector
Alternative to: Progressive offline-to-online evaluation pipeline
Online reasoning quality monitoring
When to choose. Choose in production where latency budgets and cost preclude evaluating every inference; complements the offline pipeline.
Trace Collector Trace Sampling Policy Trace Store Ephemeral Trace Buffer Evaluation Trace Sampler Query Novelty Detector Feedback Collector abstract Behavioral Signal Tracker Reasoning Chain Decomposer Reasoning Quality Scorer abstract LLM Judge abstract Metrics Collector Quality Drift Detector Alert Manager Reasoning Quality Threshold Configuration
Sources: Ch3.9
Orchestrated core with event-driven periphery (multi-agent customer support)
When to choose. Choose when core processing needs predictable, centrally error-handled sequencing while non-critical notifications must fan out to many consumers without coupling.
Supervisor Agent Request Intake Agent Knowledge Retrieval Agent Intent Router Answer Synthesizer abstract Escalation Agent Static Delegation Interface Tool Protocol Server Embedding Service abstract Vector Index Store abstract LLM Inference Service Publish-Subscribe Bus Proactive Notifier Human Specialist Output Verifier Retrieval Result Cache Rule-Based Intent Classifier Template Response Generator Fallback Work Queue Trace Collector Metrics Collector
Sources: Ch1.3
Parallel paradigms with result fusion
When to choose. Choose when different paradigms see complementary aspects of the same input and redundancy/diversity justify multiplied compute cost (e.g., credit risk assessment).
Learned-Policy Decision Engine Utility-Based Decision Maker Rule-Based Decision Engine Decision Fusion Aggregator abstract Formal Rule Specification Utility Function Specification
Alternative to: Sequential neural-to-symbolic pipeline, Cooperative iterative neural-symbolic refinement, Embedded neural modules within a symbolic program
Sources: Ch5.13
Per-region full agent stacks with nearest-region routing
When to choose. Choose only when compliance requires geographic redundancy or a global user base needs low-latency regional endpoints.
DNS Load Balancer Load Balancer abstract Agent Controller abstract External Session State Store Vector Index Store abstract Distributed Response Cache Container Orchestrator
Alternative to: Single-region, multi-availability-zone deployment
Sources: Ch4.7
Performance proxy with selective API management
When to choose. Choose when most traffic needs maximum routing throughput but a subset of critical endpoints requires centralised authentication, rate limiting and policy enforcement.
High-Performance Reverse Proxy Policy Plugin Gateway Rate Limiter Authorization Policy Decision Point Agent API Gateway
Sources: Ch4.1
Plan-and-Execute structured workflow agent
When to choose. Choose for complex, mostly predictable multi-step workflows in stable environments where minimising expensive LLM calls matters.
Plan-and-Execute Controller Task Planner abstract Plan Executor Replanner abstract Execution Plan abstract Parallel Tool Dispatcher LLM Inference Service Foundation LLM abstract Tool Registry Tool Executor
Alternative to: Single-agent ReAct tool-using agent
Sources: Ch1.2
Plan-and-Execute with lean state
When to choose. Choose when the task prioritises efficiency; lean state (plan plus current step results) reduces context consumption.
State-Graph Orchestrator Task Planner abstract Plan Executor Replanner abstract Working Memory Buffer abstract Tool Executor
Alternative to: ReAct agent with rich state
Plugin-orchestrated enterprise agent
When to choose. Choose for enterprise agents integrating dozens of internal systems whose capability catalog grows over time, where plugins are reused across multiple agents and centralized observability, audit and credential control are required; avoid for focused agents with fewer than five capabilities or for strictly deterministic, graph-structured workflows.
Function-Calling Controller Function Choice Policy Tool Registry Tool Schema Tool Executor Parallel Tool Dispatcher Service Container Semantic Function Prompt Function Template REST API Adapter Database Connector LLM Inference Service Agent Workflow Configuration Secrets Vault Metrics Collector Token Cost Meter Trace Collector Audit Log Store
Sources: Ch2.5
Plugin-routed enterprise agent platform
When to choose. Choose for agent ecosystems with many specialized capabilities requiring dynamic routing, enterprise system integration and centralized observability, especially within an existing Azure/Microsoft estate.
Plugin Kernel Orchestrator Tool Registry LLM Inference Service Token Cost Meter
Sources: Ch2.1
Policy-governed human-over-the-loop
When to choose. Choose for enterprise agents making hundreds or thousands of decisions daily where real-time approval is impractical: humans author tiered policies and decision boundaries, the policy engine enforces them automatically, and only boundary crossings, low-confidence cases and irreversible actions reach humans.
Action Policy Engine Policy Context Aggregator Parameter Security Validator Organizational Baseline Policy Departmental Policy Team Policy Agent-Specific Policy Action Risk Tier Policy Approval Authority Matrix Policy Enforcement Mode Configuration abstract Policy Violation Responder Escalation Protocol abstract Confidence Estimator Confidence Gate Approval Gateway Human Approver Approval Timeout Fallback Policy abstract Trace Collector Audit Log Store Governance Decision Log Decision Factor Explainer Feedback Collector abstract Override Feedback Record Reward Model Trainer Override Rate Monitor Quality Drift Detector Fairness Monitor Human Oversight Protocol
Alternative to: Human-in-the-loop synchronous approval, Human-on-the-loop asynchronous monitoring
Sources: Ch10.5
PPO-based RLHF alignment pipeline
When to choose. Choose when alignment requires online learning, environmental interaction or integration with complex RL frameworks, and the organization can afford multi-model GPU memory and RL expertise.
Foundation LLM abstract Instruction Demonstration Dataset Fine-Tuning Pipeline abstract Reference Policy Model Alignment Prompt Dataset Candidate Response Sampler Annotation Task Router Preference Annotation Console Pairwise Comparison Format Preference Annotator abstract Annotation Guideline Annotation Quality Monitor Preference Label Aggregator Preference Agreement Filter Preference Dataset Reward Model Trainer Reward Model Preference Reward Scorer RLHF Policy Optimizer Value Network Preference Optimization Config Reward Hacking Monitor Fine-Tuned Agent Model Training Pipeline Orchestrator Distributed Model Executor abstract
Alternative to: DPO offline preference alignment
Sources: Ch10.3
Pre-optimized LLM microservice on Kubernetes
When to choose. Choose when agents need production LLM serving through OpenAI-compatible APIs with minimal inference-optimisation expertise, on a Kubernetes cluster with suitable GPUs.
LLM Inference Service OpenAI-Compatible Inference API Model Deployment Profile abstract In-Flight Batch Scheduler Inference Service Operator Deployment Manifest Container Orchestrator GPU Node abstract GPU Device Plugin Secrets Vault Container and Model Artifact Registry Layer-4 Load Balancer Agent API Gateway Liveness Endpoint Readiness Endpoint Container Health Prober Metric-Driven Autoscaler Autoscaling Policy Metrics Collector LLM Provider Adapter
Sources: Ch4.5
Preference-aligned (RLHF/DPO) agent
When to choose. Choose when desired behaviour involves context-dependent trade-offs (speed vs thoroughness, policy vs satisfaction, escalation vs autonomy) that no single demonstration captures.
Fine-Tuned Agent Model Candidate Response Sampler Preference Annotator abstract Preference Annotation Console Annotation Guideline Annotation Quality Monitor Preference Dataset Reward Model Trainer Reward Model Composite Reward Scorer Factuality Verifier abstract RLHF Policy Optimizer Direct Preference Optimizer Training Pipeline Orchestrator Training Hyperparameter Tuner Evaluation Harness Bias Evaluator Online Evaluator
Sources: Ch3.5
Proactive (push-based) assistant
When to choose. Choose when an agent must anticipate needs and initiate timely assistance before users ask (push-based), with graduated autonomy, granular consent and a feedback flywheel; avoid universal proactivity where interruptions cost more than they deliver.
Proactive Agent Situational Context Integrator Environmental Context Adapter User Affect Detector User Need Predictor User Trajectory Model Intervention Value Estimator Intervention Threshold Policy Intervention Timing Optimizer Notification Volume Governor Proactive Notifier Action Suggestion Engine User Activity History Store User Preference Profile Store User Exception Catalog Episodic Memory Store Confidence Gate Risk Gate Approval Gateway Post-Action Review Sampler Action Rollback Service Decision Factor Explainer Explanation Presenter Consent Manager Consent Registry Consent Enforcement Gate Audit Log Store Behavioral Signal Tracker Revealed Preference Learner Training Pipeline Orchestrator Value Alignment Criteria Human Approver
Sources: Ch10.2
Production centralized tracing
When to choose. Choose for production and enterprise deployments requiring centralized collection, long-term storage, cross-agent correlation, real-time alerting and compliance/audit integration.
Trace Collector Trace Schema Trace Context Propagator Trace Exporter Telemetry Gateway Trace Sampling Policy Trace Store Trace Visualizer Trace Error Classifier Agent Behavior Anomaly Detector Execution Profiler Token Cost Meter Alert Manager Platform Operator
Alternative to: Development direct-export tracing
Sources: Ch3.6
Production multi-agent Kubernetes deployment
When to choose. Choose when a multi-agent system faces variable demand, GPU inference workers and stateful memory stores requiring self-healing, autoscaling, isolation and monitoring.
Container Orchestrator Supervisor Agent Worker Agent abstract Stateless Workload Controller Stateful Workload Controller Deployment Manifest Metric-Driven Autoscaler Autoscaling Policy Container Health Prober Layer-4 Load Balancer Persistent Volume Model Artifact Cache GPU Node abstract Cluster Namespace Namespace Resource Quota Pod Network Policy Metrics Collector Metrics Dashboard abstract Alert Manager
Alternative to: Minimal single-cluster agent deployment
Sources: Ch4.3
Profile-optimize-benchmark agent performance lifecycle
When to choose. Choose for production agents that evolve continuously (prompt, model, framework changes) and must keep accuracy, latency and cost SLAs; baseline-profile, optimize, benchmark continuously and re-profile.
Trace Collector Execution Profiler Token Cost Meter Profile Report Performance Baseline Optimization Recommender Parallel Tool Dispatcher Tool Result Cache Evaluation Harness Evaluation Dataset Semantic Similarity Scorer Regression Threshold Policy Evaluation Result Store Performance Trend Analyzer Continuous Integration Runner Regression Gate Alert Manager Trace Exporter
Progressive deepening planning
When to choose. Choose when fast coarse plans are needed up front and world state becomes more certain during execution (e.g., route -> path -> motion planning in autonomous vehicles).
HTN Planner Task Network State Abstraction Mapper Plan Executor Replanner abstract World Model State
Sources: Ch5.4
Progressive offline-to-online evaluation pipeline
When to choose. Choose when production traffic is sufficient for A/B significance and changes are major (model swap, workflow restructuring): filter candidates offline, validate in staging, then A/B test before full rollout.
Continuous Integration Runner Evaluation Harness Evaluation Dataset Response Scorer abstract Experiment Tracker Evaluation Result Store Evaluation Baseline Statistical Comparator Regression Gate Regression Threshold Policy Staging Environment A/B Test Traffic Splitter A/B Test Configuration Experiment Guardrail Monitor Rollout Manager abstract Release Approver
Alternative to: Offline-only continuous evaluation
Prompt-chained tool workflow
When to choose. Choose when tool-chain steps are predictable and validation gates are required between steps.
Prompt Chain Orchestrator Sequential Tool Dispatcher Tool Executor Tool Result Transformer Progress Validator Retry Handler Working Memory Buffer abstract Tool Schema Native Function-Calling API abstract
Alternative to: Agent-driven tool chaining
Sources: Ch2.6
Prompt-engineered (zero-shot) agent
When to choose. Choose first: for well-understood tasks the model handles zero-shot and for problems stemming from unclear instructions, since prompt engineering costs minutes and needs no training infrastructure.
System Prompt Template Prompt Optimizer Evaluation Harness Evaluation Dataset Online Evaluator Inference Serving Configuration abstract LLM Inference Service Foundation LLM abstract Trace Collector
Alternative to: Few-shot in-context adaptation, Retrieval-grounded adaptation, Fine-tuned specialist agent
Sources: Ch3.5
Proportionate small/medium-organization AI governance
When to choose. Choose when a small or medium organization applies the same framework principles with simpler structures and fewer roles, quick-start focus on high-priority risks, proportional documentation and focused controls.
Accountable AI Executive AI Governance Policy AI Use Inventory AI Risk Tier Classifier Harm Risk Register Risk Scorer abstract Risk Treatment Plan Model Card Dataset Datasheet Audit Log Store Incident Manager
Alternative to: Integrated NIST AI RMF + ISO/IEC 42001 governance program
Sources: Ch9.8
Prototype RAG with embedded vector store
When to choose. Choose for first RAG prototypes, notebooks and local development with fewer than ~100K vectors, where simplicity and speed-to-market matter most; plan migration to a production-grade store.
Embedded Vector Index Store Hosted Embedding API Service Cost-Optimized Embedding Model Vector Retriever abstract
Alternative to: Single-node production vector store, Highly available replicated vector store cluster, Distributed billion-scale vector search
Pull-based metrics, dashboards, alerting and structured logging
When to choose. Choose to validate optimization effectiveness and SLA compliance for production inference deployments on a container orchestrator.
Inference Metrics Endpoint Metrics Scrape Configuration Metrics Collector Time-Series Metrics Store Metrics Dashboard abstract Alert Rule Set Alert Manager Platform Operator Log Aggregator Centralized Log Store Token Cost Meter
Pure learning-based decision agent
When to choose. Choose when training data is ample, reward signals are clearly defined, continuous adaptation is tolerated, and no explainability or verification mandate applies (e.g., self-play game agents).
Learned-Policy Decision Engine Policy Network Reinforcement Learning Policy Learner Environment Simulator Experience Replay Buffer Reward Function Specification Opponent Policy League
Alternative to: Three-layer hybrid decision agent (strategic utility / tactical rules / operational learning)
Quality-critical LLM serving
When to choose. Choose when response accuracy is the priority and higher precision, larger models and extra verification cost are acceptable.
Large Language Model Tier FP16 Inference Engine Model Ensemble Orchestrator Output Verifier LLM Inference Service
Sources: Ref7.15
Quality-optimized agent configuration
When to choose. Choose when accuracy is paramount and mistakes carry serious consequences (e.g., medical diagnosis, complex analysis) and latency/cost budgets allow ~6s P95 and ~$0.12/query.
Large Language Model Tier Inference Serving Configuration abstract Full Conversation Buffer Iteration Limit Policy ReAct Agent Controller
Alternative to: Balanced-performance agent configuration, Latency-optimized agent configuration, Cost-optimized agent configuration
Sources: Ch3.4
ReAct agent with rich state
When to choose. Choose when the task prioritises adaptation because solution paths are unpredictable and need dynamic adjustment.
Agent Controller abstract Reasoning Engine Tool Executor Working Memory Buffer abstract LLM Inference Service Prompt Context Builder Iteration Limit Policy Progress Validator State Cycle Detector Summarizing History Compressor
Alternative to: Plan-and-Execute with lean state
Read-only web research agent
When to choose. Choose for initial production deployment or READ tasks (product research, comparison, availability checking) that require no state modification, limiting risk while building confidence.
Web Navigation Agent Task Planner abstract Replanner abstract Browser Navigator Page Content Extractor Page State Observer External Website External Service API Retry Handler Working Memory Buffer abstract Conversation State Store abstract Trace Collector Trace Store
Alternative to: Transactional web agent
Sources: Ch3.3
Real-time voice customer-service agent
When to choose. Choose for live phone or voice-assistant interactions needing sub-second end-to-end turn latency.
Sensor Input Adapter Voice Turn Coordinator Streaming Speech Recognizer Speech Recognition Model ASR Word Boost List Agent Controller abstract LLM Inference Service Speech Synthesizer Speech Synthesis Model Voice Persona Profile Response Streamer Inference Server Queue-Depth Autoscaler
Sources: Ch7.5
Registry-gated GitOps progressive delivery
When to choose. Choose for enterprise agent releases needing reproducible artifacts, gated approvals, declarative deployment and automated canary rollback.
Continuous Integration Runner Experiment Tracker Model and Agent Release Registry Versioned Agent Release Stage Promotion Controller Regression Gate Release Approver Configuration Repository Deployment Manifest Environment Overlay GitOps Reconciler Rollout Manager abstract Rollout Analysis Template Experiment Guardrail Monitor Metrics Collector Layer-7 Load Balancer Container Orchestrator
Sources: Ch4.4
Regulated conversational claims processing with mandatory human checkpoints
When to choose. Choose when automated decisions in regulated domains (insurance, finance, healthcare) require mandatory human checkpoints for cases above complexity or value thresholds plus complete decision audit trails.
Conversational (Chat) Interface User Identity Verifier Workflow Orchestrator abstract Compliance Policy Rule Set Risk Gate Human Approver Audit Log Store Decision Explainer abstract Explanation Presenter Compliance Officer
Sources: Ch10.1
Reserved baseline + spot burst + on-demand buffer
When to choose. Choose to minimize cost when load has a predictable baseline and transient peaks, accepting spot interruptions covered by an on-demand buffer.
Reserved GPU Node Spot GPU Node On-Demand GPU Node Spot Interruption Handler Instance Group Autoscaler Load Balancer abstract
Sources: Ch4.7
Retrieval-grounded adaptation
When to choose. Choose when the agent lacks knowledge (recent events, proprietary information, specialized facts) that no prompt or demonstration can supply; leaves model weights unchanged.
System Prompt Template Retriever abstract Context Assembler abstract Vector Index Store abstract Factuality Verifier abstract LLM Inference Service Foundation LLM abstract
Alternative to: Fine-tuned specialist agent
Sources: Ch3.5
Risk-tiered human-in-the-loop oversight
When to choose. Choose when agents take actions with real-world consequences whose risk magnitude, confidence, and reversibility vary, so notification, approval, and monitoring must be applied selectively.
Oversight Gate abstract Confidence Gate Risk Gate Escalation Threshold Policy Adaptive Threshold Tuner Proactive Notifier Action Rollback Service Approval Gateway Approval Review Console Execution Monitor Console Human Approver Human Supervisor Audit Log Store
Role-based hierarchical crew
When to choose. Choose for multi-agent workflows with clear role specialization and well-defined task dependencies (e.g., research, write, edit content pipeline) executed sequentially or via manager delegation.
Role-Based Task Orchestrator Supervisor Agent Worker Agent abstract System Prompt Template
Alternative to: Plugin-routed enterprise agent platform
Sources: Ch2.1
Role-based hierarchical team with manager quality gates
When to choose. Choose when work maps to an organizational team with clear role specialization and needs manager-enforced validation gates and iterative refinement (content production, software pipelines, data analysis, compliance review).
Supervisor Agent Worker Agent abstract Task Definition System Prompt Template Critique Rubric Agent Delegation Interface abstract LLM Inference Service
Alternative to: Conversation-driven multi-agent collaboration, Role-based sequential team
Sources: Ch2.4
Role-based sequential team
When to choose. Choose when work maps to a team whose tasks have fixed, obvious progression and discrete deliverables at each stage.
Role-Based Task Orchestrator Worker Agent abstract Task Definition System Prompt Template LLM Inference Service
Alternative to: Conversation-driven multi-agent collaboration, Role-based hierarchical team with manager quality gates
Sources: Ch2.4
Rollout-based MCTS planning
When to choose. Choose when a fast simulator exists but no trained networks or training data are available, heuristics are hard to design, and near-optimal decisions within a real-time budget suffice.
MCTS Planner MCTS Search Tree Store Search Tree Pruner Rollout Simulator Environment Simulator Reward Function Specification MCTS Search Configuration Search Budget Policy Plan Executor
Alternative to: Neural-guided MCTS (AlphaGo/AlphaZero style)
Sources: Ch5.5
Rule-constrained utility optimisation (hybrid)
When to choose. Choose when hard safety, regulatory or ethical constraints coexist with continuous multi-objective trade-offs (autonomous driving lane changes, trading within risk limits, treatment recommendation).
Rule Constraint Filter Formal Rule Specification Decision Engine abstract Logic Conclusion Verbalizer Explanation Presenter
Sources: Ch5.11
Safety-separated reliability monitoring
When to choose. Choose when a guardrailed agent produces both safety blocks and infrastructure failures, so each category needs its own SLO, alert routing and dashboard panels to avoid conflated alerts and misattributed error budgets.
Guardrail Orchestrator Guardrail Policy Input Rail Output Rail Fact Checking Rail abstract Failure Category Classifier Metrics Collector Time-Series Metrics Store Service Level Objective Specification Alert Rule Set Alert Manager Metrics Dashboard abstract Security Analyst Platform Operator
Sources: Ch8.2B
Self-hosted accelerated function calling
When to choose. Choose for tool-heavy agents with data-residency, security, or cost constraints, or high throughput (thousands of requests per hour).
Self-Hosted Inference Endpoint LLM Inference Service Foundation LLM abstract Optimized Inference Engine abstract Engine Builder KV Cache Manager Tool Schema GPU Node abstract Container Orchestrator
Sources: Ch2.6
Self-hosted GPU embedding for regulated, long-document workloads
When to choose. Choose when data sovereignty or air-gapped operation is required, documents are long, and volume exceeds a few thousand queries per day.
Self-Hosted GPU Embedding Service Long-Context Embedding Model Inference Server Dynamic Batch Scheduler Engine Builder OpenAI-Compatible Inference API GPU Node abstract
Sources: Ch6.1
Separate modality stores with cross-modal reranking
When to choose. Choose for research environments experimenting with per-modality embedding models, or production systems with mature MLOps that accept the highest operational cost for best-in-class per-modality results.
Document Ingestor Modality-Specific Embedding Service Modality-Specific Vector Store Per-Modality Fan-Out Retriever Cross-Modal Reranker Cross-Modal Reranking Model Fine-Tuning Pipeline abstract Multimodal Context Assembler Multimodal Answer Synthesizer Inference Server
Alternative to: Unified embedding space multimodal RAG, Ground-to-text multimodal RAG with metadata
Sources: Ch2.7
Sequential neural-to-symbolic pipeline
When to choose. Choose when the task decomposes cleanly into perception then rule application and early-stage processing can extract all relevant information without later feedback; most common in production.
Perception Interpreter Neural Perception Model Neural-to-Symbolic Translator Paradigm Boundary Validator Rule-Based Decision Engine Formal Rule Specification Knowledge Graph Store abstract
Alternative to: Parallel paradigms with result fusion, Cooperative iterative neural-symbolic refinement, Embedded neural modules within a symbolic program
Sources: Ch5.13
Sequential tool-calling single agent
When to choose. Choose for single-agent workflows that follow a linear sequence (receive query, reason about tools, execute tools, synthesize answer) with state limited to conversation history, e.g., FAQ/knowledge-base chatbots and simple QA. Also fits rapid prototyping, conversation-based interfaces needing only message history, simple tool integration and unmodified ReAct reasoning (e.g., research assistants, database Q&A, calendar/email assistants).
ReAct Agent Controller Tool Executor Tool Schema LLM Inference Service Conversation State Store abstract Working Memory Buffer abstract Vector Retriever abstract Embedding Service abstract Agent Action Output Parser Full Conversation Buffer LLM Provider Adapter System Prompt Template Tool Integration Adapter abstract Trace Collector
Alternative to: Graph-based iterative workflow agent, Conversation-driven multi-agent collaboration, Role-based hierarchical crew, Plugin-routed enterprise agent platform
Serverless event-driven agents
When to choose. Choose for highly variable or bursty traffic with idle periods, discrete event workloads, rapid iteration (many deployments per day) and small teams without infrastructure expertise, when cold starts fit within the latency SLO.
Serverless Function Runtime On-Demand Consumption Plan Event-Triggered Agent Agent API Gateway Content-Based Event Router Publish-Subscribe Bus Message Queue Dead Letter Queue Event Stream Log Idempotency Store Schema Registry Object Store State Checkpoint Store abstract Stalled Workflow Resumer Durable State Machine Orchestrator Conversation State Store abstract Result Callback Webhook LLM Inference Service Trace Collector
Alternative to: Containerised microservices agent deployment
Sources: Ch4.2
Session-affinity replicas
When to choose. Choose when latency needs are extreme (sub-100 ms), scale is tens of instances, and session recreation on failure is acceptable.
Layer-4 Load Balancer Agent Controller abstract Instance-Local Session State In-Process Response Cache LLM Inference Service
Alternative to: Stateless replica horizontal scaling
Sources: Ch1.8
Single large instance (vertical scaling)
When to choose. Choose for computation or memory bottlenecks with predictable workloads and modest scale (10-50 req/s, bursts to ~100) in early production where failure tolerance is less critical.
Agent Controller abstract LLM Inference Service In-Process Response Cache Instance-Local Session State On-Demand GPU Node
Alternative to: Stateless replica horizontal scaling
Sources: Ch1.8
Single tool integration
When to choose. Choose when one primary capability extends the LLM (current data lookup, calculations, simple API calls) and each query triggers at most one independent tool call, e.g., FAQ or order-status agents.
Direct Tool-Calling Controller Reasoning Engine Native Function-Calling API abstract LLM Inference Service Tool Schema Tool Registry Tool Executor REST API Adapter External Service API Trace Collector
Sources: Ch2.6
Single-agent agentic RAG with self-hosted Nemotron
When to choose. Choose for knowledge-base assistants where the agent should decide per query whether to retrieve, using hybrid retrieval plus reranking and a self-hosted OpenAI-compatible endpoint.
ReAct Agent Controller System Prompt Template Tool Schema Iteration Limit Policy File Store Extractor Overlapping Window Chunker Text Embedding Service Retrieval-Optimized Embedding Model Embedded Vector Index Store Dense-Sparse Hybrid Retriever Vector Retriever abstract Keyword Retriever Reranker abstract OpenAI-Compatible Inference API Inference Server Small Language Model Tier In-Process Response Cache Trace Collector
Single-agent ReAct tool-using agent
When to choose. Choose when solution paths are unpredictable and each step depends on prior observations (research, debugging, non-standard support), and latency and cost overhead are acceptable.
ReAct Agent Controller Reasoning Engine Prompt Exemplar Set LLM Inference Service Foundation LLM abstract Working Memory Buffer abstract Context Window Manager abstract Tool Registry Tool Schema Tool Executor Retry Handler Circuit Breaker
Alternative to: Plan-and-Execute structured workflow agent
Sources: Ch1.2
Single-node production vector store
When to choose. Choose for production RAG with millions of documents where authentication, durability and monitoring are required but brief outages on node failure are tolerable (below 99.9% uptime needs).
Self-Managed Vector Index Store Persistent Volume Deployment Manifest API Key Authenticator Vector Store gRPC API Knowledge Chunk Metadata Schema Vector Index Build Configuration abstract Vector Search Configuration Vector Batch Ingestor Ingestion Pipeline Orchestrator Dense-Sparse Hybrid Retriever Fusion Weight Selector Metrics Collector Time-Series Metrics Store Metrics Dashboard abstract Alert Manager Alert Rule Set Vector Store Capacity Monitor
Alternative to: Highly available replicated vector store cluster
Sources: Ch6.2B
Single-region, multi-availability-zone deployment
When to choose. Choose by default for high availability (99.9%+) against individual zone failures without multi-region cost and complexity.
Load Balancer abstract Agent Controller abstract External Session State Store Vector Index Store abstract Container Orchestrator
Alternative to: Per-region full agent stacks with nearest-region routing
Sources: Ch4.7
Single-strategy reactive replanning agent
When to choose. Choose when failures are rare (<10% of executions) and unpredictable, time budgets tolerate multi-second replanning pauses, memory is constrained, or state spaces are small.
Plan-and-Execute Controller Optimal Heuristic Search Planner Geometric Distance Heuristic State-Space Graph Plan Executor Plan Deviation Monitor Discrepancy Significance Evaluator Complete Replanner
Sources: Ch5.6
SSE streaming RAG agent
When to choose. Choose for query-then-read agents (e.g., customer support RAG) where no mid-stream user interaction is needed and sub-second perceived responsiveness is required.
Conversational (Chat) Interface Server-Sent Events Stream Response Streamer Stream Connection Manager Agent Controller abstract Dense-Sparse Hybrid Retriever Vector Retriever abstract Keyword Retriever Vector Index Store abstract In-Process Response Cache Semantic Cache Context Assembler abstract Context Compressor Answer Synthesizer abstract Native Function-Calling API abstract LLM Inference Service Error Presenter Retry Handler Metrics Collector Trace Collector Layer-7 Load Balancer
Alternative to: Bidirectional interactive streaming agent
Sources: Ch2.9
Stage-gated data-quality pipeline for high-stakes RAG
When to choose. Choose when a RAG agent serves high-stakes or regulated domains (healthcare, finance, legal) where 97-99% 'good enough' quality, duplicate contradictions or embedded PII are unacceptable.
Knowledge Source System Source Change Detector Ingestion Pipeline Orchestrator Document Ingestor Data Quality Validator Source Document Schema Data Quality Rule Set Business Rule Validator Formal Rule Specification Document Structure Validator Data Quality Gate Data Quality SLA Specification Data Format Normalizer Canonical Format Specification Cascading Deduplicator Exact Hash Deduplicator Fuzzy Text Deduplicator Semantic Deduplicator Referential Integrity Validator Document Version Reconciler Multi-Strategy PII Detector Document PII Redactor Document Chunker abstract Text Embedding Service Vector Index Store abstract Index Integrity Validator Smoke Tester Content Freshness Monitor Content Freshness Policy Corpus Drift Detector Data Validation Result Store Data Quality Review Queue Data Quality Reviewer Metrics Dashboard abstract Alert Manager SLO Monitor Evaluation Dataset RAG Evaluator Audit Log Store
Sources: Ch6.4
Standard single-query RAG
When to choose. Choose for simple, single-need queries where decomposition adds latency and token cost without improving quality.
RAG Query Orchestrator Text Embedding Service Retriever abstract Context Assembler abstract Answer Synthesizer abstract LLM Inference Service
Alternative to: Decomposed RAG (query decomposition + parallel retrieval + synthesis)
Sources: Ch6.6
Stateless elastic agent fleet
When to choose. Choose for interactive agents with variable request durations and diurnal traffic, where every replica can serve any request because all state is externalized.
Agent Controller abstract Least-Connections Load Balancer Resource-Utilization Autoscaler Autoscaling Policy Container Orchestrator Deployment Manifest Container Health Prober Liveness Endpoint Readiness Endpoint External Session State Store Object Store Log Aggregator Metrics Collector Metrics Dashboard abstract
Alternative to: Sticky-session migration stage
Sources: Ch4.7
Stateless replica horizontal scaling
When to choose. Choose for most (~90%) production agent deployments: discrete or conversational requests with externalized state, unpredictable spikes, high availability (99.99%) or multi-tenant SaaS scale.
Layer-7 Load Balancer Layer-4 Load Balancer Agent API Gateway Agent Controller abstract LLM Inference Service Small Language Model Tier External Session State Store Distributed Response Cache KV Cache Manager Vector Index Store abstract Task Queue Metric-Driven Autoscaler Autoscaling Policy Deployment Manifest Container Health Prober Liveness Endpoint Readiness Endpoint Metrics Collector Alert Manager On-Demand GPU Node Container Orchestrator
Alternative to: Single large instance (vertical scaling), Session-affinity replicas
Sources: Ch1.8
Sticky-session migration stage
When to choose. Choose only temporarily while a partially stateful agent still caches conversation context locally, before state is externalized and routing switches to least connections.
Agent Controller abstract Session Affinity Load Balancer Instance-Local Session State External Session State Store Container Orchestrator Container Health Prober
Alternative to: Stateless elastic agent fleet
Sources: Ch4.7
Structured Chain-of-Thought reasoning with layered verification
When to choose. Choose for open-ended, contextual or judgment-heavy tasks where flexible pattern-based reasoning is more appropriate than formalisation.
Reasoning Engine Structured Reasoning Prompt Template abstract Stepwise Reasoning Verifier Citation Verifier Trace Collector
Sources: Ch3.9
Supervised physical autonomy (manufacturing / robotics)
When to choose. Choose for autonomous robots or manufacturing systems operating near people, where physical safety requires continuous human supervisory awareness, veto authority and immediate shutdown capability.
Agent Controller abstract Physical Safety Envelope Execution Monitor Console Human Supervisor Emergency Stop Controller Action Policy Engine
Sources: Ch10.5
Swarm intelligence multi-agent system
When to choose. Choose for distributed optimisation or coverage with simple agents where central control would be a bottleneck or single point of failure and 'good enough' solutions suffice.
Swarm Agent Swarm Anomaly Detector
Alternative to: Competitive (game-theoretic) multi-agent system
Sources: Ch1.3
Three-layer hybrid decision agent (strategic utility / tactical rules / operational learning)
When to choose. Choose for safety-critical open-world agents (e.g., autonomous vehicles) that must optimize goals, obey hard legal/safety rules, and adapt control behaviour at once.
Hybrid Decision Arbiter Decision Precedence Policy Utility-Based Decision Maker Utility Function Specification Rule-Based Decision Engine Rule Constraint Filter Formal Rule Specification Learned-Policy Decision Engine Policy Network Replanner abstract Execution Plan abstract Multimodal Fusion Engine World Model State Audit Log Store
Alternative to: Pure learning-based decision agent
Throughput-bound speculative serving
When to choose. Choose for batch or offline generation with predictable outputs (summarization, structured data) where GPU-hours matter more than time-to-first-token.
Inference Server Large Language Model Tier Speculative Draft Model abstract Speculative Decoder Speculative Decoding Configuration Throughput-Oriented Batching Config Token Predictability Analyzer Evaluation Harness
Sources: Ch4.4
Throughput-critical HITL approval (e.g., batch claims, bulk onboarding)
When to choose. Choose when processing volume matters more than immediate response: batch similar requests, consolidated review with bulk actions, predictive pre-approval, and risk-stratified autonomous handling of routine cases.
Oversight Gate abstract Risk Gate Pre-Approved Category Policy Approval Request Batcher Approval Request Store Approval Router Approval Review Console Oversight Performance Monitor Human Approver Human Specialist
Alternative to: Latency-critical HITL approval (e.g., real-time fraud review)
Sources: Ch10.4
Throughput-critical LLM serving
When to choose. Choose when maximum tokens per second is the priority (batch document analysis, summarization, embedding generation).
LLM Inference Service In-Flight Batch Scheduler Dynamic Batch Scheduler Throughput-Oriented Batching Config Throughput-Optimized Deployment Profile Paged KV Cache Allocator
Alternative to: Cost-sensitive LLM serving, Quality-critical LLM serving
Throughput-critical quantized LLM serving
When to choose. Choose for high-volume services (chatbots for millions of users, real-time translation, content moderation) that accept ~1-2% accuracy loss for 4-8x capacity.
Model Graph Exporter Model Quantizer Quantization Calibrator Quantization Calibration Dataset Engine Builder Engine Build Configuration FP8 Quantized Engine INT8 Quantized Engine Paged KV Cache Allocator In-Flight Batch Scheduler Fused Block-wise Attention Kernel Inference Server Evaluation Harness
Alternative to: Accuracy-critical FP16-baseline serving, Latency-critical low-batch serving
Throughput-optimized batch inference serving
When to choose. Choose for batch workloads (document summarization, moderation queues, bulk translation) where completion within a processing window matters more than per-request latency.
Concurrent Inference Request Dispatcher Connection Pool OpenAI-Compatible Inference API Inference Server Inference Batch Scheduler abstract Throughput-Oriented Batching Config GPU Node abstract
Alternative to: Latency-optimized interactive inference serving, Cost-optimized inference serving
Sources: Ch7.2
Tiered-autonomy conversational customer service assistant
When to choose. Choose for high-volume customer service where most requests have clear intent and straightforward resolution paths, autonomously resolving routine inquiries (Klarna 65%, Erica 98%) while escalating complex, sensitive or low-confidence cases to humans with full context.
Conversational (Chat) Interface Dialogue Flow Manager Intent Router Intent Taxonomy Clarification Manager Hierarchical History Compressor Working Memory Buffer abstract Session Summary Store Semantic Memory Store Escalation Agent Escalation Handler Escalation Handoff Package Human Specialist Feedback Collector abstract Behavioral Signal Tracker Active Learning Sampler Training Pipeline Orchestrator Metrics Dashboard abstract Alert Manager
Sources: Ch10.1
Time-indexed audio RAG
When to choose. Choose when critical knowledge exists only in recorded conversations (earnings calls, support calls, training sessions, design discussions) and answers must cite exact playback timestamps; skip when adequate transcripts or notes already exist.
Speech Transcriber Speech Recognition Model Time-Indexed Transcript Chunker Multimodal Chunk Metadata Schema Text Embedding Service Vector Index Store abstract Metadata-Filtered Retriever Source Media Store Timestamped Citation Presenter Answer Synthesizer abstract
Sources: Ch2.7
Top-down (rule/principle-based) value alignment
When to choose. Choose when explicit verifiability and auditability matter most and scenarios are clear and unambiguous; accept brittleness in novel situations outside the specified rules.
Constitution abstract Principle Priority Policy Operational Norm Set Formal Rule Specification Rule Constraint Filter Guardrail Policy Self-Reflection Critic Principle Adherence Evaluator Principle Adherence Test Suite
Alternative to: Bottom-up (learned) value alignment, Hybrid value alignment (hard constraints + learned nuance + continuous monitoring)
Sources: Ch9.6
ToT constraint satisfaction (DFS + sequential proposal + value checks)
When to choose. Choose for scheduling-style problems with interdependent hard and soft constraints where violations surface only at deeper levels and backtracking is essential.
Depth-First Thought Search Controller Sequential Proposal Thought Generator Value Thought Evaluator Problem Constraint Specification Thought Search Policy LLM Inference Service Human Evaluator External Service API
Alternative to: ToT creative generation (BFS + independent sampling + voting)
Sources: Ch5.2
ToT creative generation (BFS + independent sampling + voting)
When to choose. Choose for creative content whose quality depends on high-level structure and resists absolute scoring, with human review before publication.
Breadth-First Thought Search Controller Independent Sampling Thought Generator Vote Thought Evaluator Thought Decomposition Specification LLM Inference Service Human Evaluator External Service API
Alternative to: ToT constraint satisfaction (DFS + sequential proposal + value checks)
Sources: Ch5.2
Train, optimize, containerize, deploy pipeline
When to choose. Choose when taking a fine-tuned model to production serving through engine compilation and containerized inference microservices with guardrails and continuous monitoring.
Fine-Tuning Pipeline abstract Foundation LLM abstract Engine Builder Optimized Inference Engine abstract Container Image Builder Container Image LLM Inference Service Input Rail Output Rail Inference Server Dynamic Batch Scheduler Container Orchestrator Metrics Collector Time-Series Metrics Store Metrics Dashboard abstract
Transactional web agent
When to choose. Choose after read-only reliability is established, for WRITE tasks (login, cart modification, checkout, access requests) under explicit domain policies and approval hierarchies.
Web Navigation Agent Task Planner abstract Replanner abstract Browser Navigator Page Content Extractor Web Transaction Executor Page State Observer External Website Retry Handler Secrets Vault Action Policy Engine Compliance Policy Rule Set Approval Gateway Human Approver Human Supervisor Audit Log Store Working Memory Buffer abstract Conversation State Store abstract Trace Collector Trace Store
Alternative to: Read-only web research agent
Sources: Ch3.3
Transparent rule-based decision system
When to choose. Choose when every decision must trace to explicit, auditable logic (credit, medical devices, government benefits), expert knowledge is stable and articulable, behaviour must be deterministic and testable, or data is too scarce to learn.
Rule-Based Decision Engine Symbolic Logic Engine Rule Interpreter Production Rule Base Conflict Resolution Policy abstract Working Memory Fact Store Fact Contradiction Resolver Memory Lifecycle Manager Rule Firing Trace Logic Conclusion Verbalizer Counterfactual Explainer Explanation Presenter Audit Log Store Human Approver Domain Rule Expert Compliance Officer
Alternative to: Rule-constrained utility optimisation (hybrid)
Sources: Ch5.11
Tree-of-Thought deliberate reasoning module
When to choose. Choose when problems require exploratory search with backtracking, intermediate evaluation is feasible and informative, and the accuracy gap over CoT justifies 3-5x cost.
ReAct Agent Controller Tree Search Controller abstract Thought Generator abstract Thought State Evaluator abstract Thought Tree Store Thought Decomposition Specification Thought Search Policy Thought Evaluation Prompt Template LLM Inference Service Reasoning Chain Cache Token Budget Enforcer Model Router Working Memory Buffer abstract Episodic Memory Store Semantic Memory Store Procedural Memory Store
Alternative to: Graph-of-Thought synthesis reasoning module
Sources: Ch5.2
Unified embedding space multimodal RAG
When to choose. Choose for rapid deployment on an existing text RAG stack with mostly general imagery (product photos, simple diagrams), minimal budget, and prototyping timelines of days.
Document Ingestor Joint Multimodal Embedding Service Contrastive Image-Text Encoder Vector Index Store abstract Unified Embedding Retriever Context Assembler abstract Answer Synthesizer abstract Inference Server OpenAI-Compatible Inference API
Alternative to: Ground-to-text multimodal RAG with metadata, Separate modality stores with cross-modal reranking
Sources: Ch2.7
Unified multi-framework inference serving
When to choose. Choose when an agent system must serve heterogeneous models from multiple frameworks (LLMs, vision, classifiers, tree-based) on shared GPU infrastructure with one API and common monitoring and scaling.
Inference Server Tensor Inference API OpenAI-Compatible Inference API Inference Backend abstract Tensor Framework Backend LLM Generation Backend Dynamic Batch Scheduler In-Flight Batch Scheduler Inference Queue Policy Inference Serving Configuration abstract Model Repository Dynamic Model Loader Model Ensemble Orchestrator Model Ensemble Definition GPU Partition Inference Metrics Endpoint Metrics Collector Custom Metrics Adapter Metric-Driven Autoscaler Replica Placement Policy Metrics Dashboard abstract Alert Manager Trace Collector
Sources: Ch4.5
Validated transactional tool chain
When to choose. Choose when agents invoke chains of side-effecting tools (inventory reservation, payment, shipping) where out-of-order, duplicated or hallucinated calls cause concrete harm requiring manual remediation.
Tool Executor Tool Schema Tool Registry Tool Call Schema Validator Parameter Provenance Validator Parameter Context Grounding Validator Tool Precondition Checker Action Sequence Policy Tool Idempotency Guard Tool Response Plausibility Checker Tool Error Classifier Retry Handler Approval Gateway Trace Collector Trace Store Audit Log Store Tool Usage Analyzer
Sources: Ch3.7
Value-aligned multi-agent coordination
When to choose. Choose when several specialised agents optimize local objectives whose externalities can violate shared organizational values (e.g., supply chain procurement, logistics, finance, customer service).
Constitution abstract Operational Norm Set Worker Agent abstract Value Sentinel Agent Agent Message Bus abstract Preference Annotator abstract Preference Dataset Reward Model Decision Stakeholder
Sources: Ch9.6
Vector RAG only
When to choose. Choose for simple factual Q&A and semantic search (customer support, documentation search) where no relationship traversal is needed; lowest complexity (p95 ~120ms).
Document Chunker abstract Embedding Service abstract Vector Index Store abstract Vector Retriever abstract Context Assembler abstract Answer Synthesizer abstract LLM Inference Service
Alternative to: Knowledge graph only (text-to-graph-query QA), Hybrid RAG + knowledge graph
Vector-only memory
When to choose. Choose when semantic similarity is a good proxy for relevance over large heterogeneous corpora and speed and low engineering overhead matter more than logical precision (e.g., customer support over documentation); a common starting point before migrating to hybrid.
Semantic Memory Store Vector Index Store abstract Text Embedding Service Text Embedding Model abstract Semantic Boundary Chunker Vector Retriever abstract Retrieval Confidence Filter Answer Synthesizer abstract LLM Inference Service
Alternative to: Knowledge-graph memory, Hybrid vector + graph memory
Voice + vision multimodal agent (accessibility, visual troubleshooting)
When to choose. Choose when users combine a spoken query with a camera image or photo (scene description for blind users, product troubleshooting).
Sensor Input Adapter Multimodal Content Router Streaming Speech Recognizer Image Captioner Vision-Language Model abstract Agent Controller abstract Retriever abstract Tool Executor LLM Inference Service Speech Synthesizer Voice Persona Profile
Sources: Ch7.5
Weighted least-connections fleet on mixed GPU/CPU nodes
When to choose. Choose for highly heterogeneous independent workloads (e.g., 2-60 s document extraction) served by GPU and CPU replicas of measurably different capacity.
Agent Controller abstract Least-Connections Load Balancer Weighted Round-Robin Load Balancer Load Balancing Policy Resource-Utilization Autoscaler GPU Node abstract CPU Compute Node Container Orchestrator Metrics Collector
Sources: Ch4.7