Architecture profiles

Named, coherent configurations (ISO/IEC/IEEE 42010-style views) that select components for a class of use cases.

Accuracy-critical FP16-baseline serving

When to choose. Choose for medical, legal or financial applications; start at FP16 and evaluate FP8 only on Hopper with acceptable benchmark accuracy.

Engine Builder Engine Build Configuration FP16 Inference Engine Evaluation Harness Evaluation Baseline Benchmark Suite Manifest Paged KV Cache Allocator Inference Server

Alternative to: Throughput-critical quantized LLM serving

Sources: Ch7.4 Ref7.05

Adaptive RASC with confidence-based escalation

When to choose. Choose for cost-constrained, high-stakes production (medical coding, compliance review, customer support) where sample budgets must adapt to difficulty, reasoning quality should weight votes, and low-agreement cases must go to humans with auditable reasoning.

Reasoning Strategy Router Query Complexity Classifier Adaptive Sample Allocator Self-Consistency Sampling Policy Reasoning Path Sampler LLM Inference Service Final Answer Extractor Reasoning Path Quality Classifier Quality-Weighted Vote Aggregator Rationale Selector Explanation Presenter Confidence Estimator Confidence Gate Escalation Threshold Policy Human Specialist Audit Log Store

Alternative to: CoT + Self-Consistency (majority vote)

Sources: Ch5.3

Adaptive retrieval RAG (retrieve-vs-generate routing)

When to choose. Choose when a substantial share of queries (e.g., 20-40%) are answerable from parametric knowledge and retrieval cost or latency matters.

RAG Query Orchestrator Query-Type Retrieval Router Retrieval Routing Rule Set Parametric Answer Generator Retriever abstract Question Decomposer Parallel Sub-Query Retrieval Controller Multi-Hop Answer Synthesizer Answer Synthesizer abstract LLM Inference Service

Sources: Ch6.6

Adaptive rule learning (explainable adaptation)

When to choose. Choose when decisions must remain explainable but patterns evolve (e.g., fraud detection, clinical guideline drift, pricing), so learned rules are proposed, validated and expert-approved before entering the rule base.

Rule-Based Decision Engine Symbolic Logic Engine Production Rule Base Rule Firing Trace Audit Log Store Rule Outcome Monitor Misclassified Case Store Labeled Decision Case Dataset Rule Learner abstract Inductive Rule Learner Case-Based Rule Refiner Candidate Rule Queue Rule Validation Orchestrator Rule Consistency Checker Rule Quality Scorer Rule Acceptance Threshold Policy Holdout Evaluation Set A/B Test Traffic Splitter Domain Rule Expert Rule Base Updater Rejected Rule Log

Sources: Ch5.11

Adaptive utility-based recommender

When to choose. Choose when a personalization agent must balance relevance against diversity/novelty and adapt trade-off weights per user and context from observed behaviour.

Utility-Based Decision Maker Utility Function Specification Outcome Probability Estimator Contextual Weight Adapter Revealed Preference Learner User Feedback Store User Activity History Store Pareto Frontier Optimizer Explanation Presenter

Sources: Ch5.10

Agent CI/CD with progressive quality gates and canary release

When to choose. Choose for agent systems committed to multiple times daily that need fast feedback yet must block behavioural regressions, vulnerable images and unstable builds before full production exposure.

Continuous Integration Runner Evaluation Workflow Definition Static Code Analyzer Static Security Scanner Agent Test Runner Unit Test Suite Integration Test Suite Evaluation Harness Evaluation Dataset Chain-of-Thought Judge Regression Gate Regression Threshold Policy Container Image Builder Container and Model Artifact Registry Container Image Scanner Vulnerability Gate Policy Model and Agent Release Registry Staging Environment Smoke Tester SLO Monitor Release Approver Rollout Manager abstract Experiment Guardrail Monitor Online Evaluator Statistical Comparator Deployment Notifier

Sources: Ch4.2

Agent-driven tool chaining

When to choose. Choose for open-ended problems where the tool sequence depends on intermediate results and cannot be predicted in advance.

ReAct Agent Controller Reasoning Engine Native Function-Calling API abstract Tool Call Dispatcher abstract Sequential Tool Dispatcher Parallel Tool Dispatcher Tool Executor Tool Result Transformer Working Memory Buffer abstract Tool Schema Tool Registry

Sources: Ch2.6

Agent-local state

When to choose. Choose when individual workflows complete in seconds to minutes and need not be distributed across instances.

State-Graph Orchestrator Agent State Schema Working Memory Buffer abstract In-Memory State Store

Alternative to: Distributed state for horizontally scaled agents

Sources: Ch1.6

Balanced-performance agent configuration

When to choose. Choose as general-purpose production default: large model, temperature 0.3, compressed 5K context, 8 iterations (~91%, 3.5s, $0.06).

Large Language Model Tier Inference Serving Configuration abstract Summarizing History Compressor Context Compressor Iteration Limit Policy

Alternative to: Latency-optimized agent configuration, Cost-optimized agent configuration

Sources: Ch3.4

Batch incremental multi-source knowledge ETL

When to choose. Choose for enterprise knowledge bases needing multi-source integration, quality validation and deduplication where latency tolerance exceeds minutes; avoid for sub-second real-time freshness, small single-source applications, on-demand embedding of dynamic queries, or vector stores without batch operations.

Pipeline Scheduler Batch Extraction Policy Incremental Refresh Policy Ingestion Pipeline Orchestrator ETL Pipeline Configuration Incremental Watermark Store Source Change Detector Source Extractor Interface Source Record Extractor Paginated API Extractor File Store Extractor Document Quality Filter Text Normalizer Exact Hash Deduplicator Overlapping Window Chunker Semantic Boundary Chunker Chunk Metadata Extractor Text Embedding Service Chunk Schema Validator Vector Batch Ingestor Vector Index Builder IVF Index Configuration Self-Managed Vector Index Store CPU Dataframe Engine Dead Letter Queue Metrics Collector Alert Manager

Alternative to: GPU-accelerated web-scale data curation

Sources: Ch6.3A Ch6.3B

Bidirectional interactive streaming agent

When to choose. Choose for agents expecting mid-stream user feedback or interruption, collaborative multi-user sessions, or multi-agent outputs streamed over one connection.

Conversational (Chat) Interface WebSocket Channel Response Streamer Stream Connection Manager Stream Interrupt Handler Conversation State Store abstract Agent Controller abstract Worker Agent abstract Explanation Presenter Native Function-Calling API abstract LLM Inference Service Metrics Collector Trace Collector

Sources: Ch2.9

Bottom-up (learned) value alignment

When to choose. Choose when contextual nuance and adaptation to unanticipated scenarios matter more than verifiable guarantees and large volumes of high-quality feedback are available; values remain implicit and opaque.

Preference Annotator abstract Preference Dataset Reward Model Trainer Reward Model RLHF Policy Optimizer Inverse Reward Learner Revealed Preference Learner Fine-Tuned Agent Model

Alternative to: Top-down (rule/principle-based) value alignment, Hybrid value alignment (hard constraints + learned nuance + continuous monitoring)

Sources: Ch9.6

Centralized multi-agent orchestration

When to choose. Choose for well-defined sequential workflows with stable task structures where predictability and centralised state outweigh scalability (up to about a dozen workers).

Supervisor Agent Worker Agent abstract Static Delegation Interface Static Rule Task Router Task State Ledger Trace Collector

Alternative to: Decentralized (peer-to-peer) multi-agent orchestration, Hierarchical multi-agent orchestration, Federated multi-agent orchestration

Sources: Ch1.3

Centrally managed edge AI fleet

When to choose. Choose for large-scale (50+ sites), geographically distributed NVIDIA edge deployments without local technical staff, with frequent model updates, high-security or mission-critical availability needs; avoid for <10 locations, cloud-only workloads or rarely updated static models.

Edge Fleet Manager Edge Provisioning Service Edge Device Agent Edge Application Definition Edge Update Orchestrator Staged Rollout Policy Edge GPU Device Cluster Consensus Coordinator GPU Partition GPU Partition Layout abstract Container and Model Artifact Registry Certificate Authority Key Management Service Encrypted Model Volume Secure Boot Verifier Inference Server Optimized Inference Engine abstract Container Health Prober Alert Manager Alert Rule Set Metrics Collector

Alternative to: Pre-optimized LLM microservice on Kubernetes

Sources: Ch4.6

Clinical decision support output filtering

When to choose. Choose for healthcare or medical-device agents where outputs affect patient safety and HIPAA/FDA obligations apply: prioritise recall, fact-check clinical claims and require human approval for high-risk actions.

Output Risk Stratifier Output Rail PII Redactor Pattern PII Detector PII Pattern Library Fact Checking Rail abstract Bias Evaluator Toxicity Classifier Moderation Triage Router Content Moderator Action Policy Engine Action Risk Tier Policy Approval Gateway Human Approver Audit Log Store

Sources: Ch9.1

Competitive (game-theoretic) multi-agent system

When to choose. Choose for strategic scenarios with conflicting interests such as resource scheduling, marketplaces, adversarial security testing, or economic simulation.

Strategic Agent Auction Task Allocator Collusion Monitor

Alternative to: Swarm intelligence multi-agent system

Sources: Ch1.3

Compiled, quantized LLM engine behind a multi-framework server on Kubernetes

When to choose. Choose for long-running, high-throughput or latency-sensitive production serving of >7B-parameter models where inference cost dominates; avoid for rapid prototyping, research iteration, <1B models or low-traffic applications.

Engine Builder Model Quantizer Quantization Calibrator Quantization Calibration Dataset Optimized Inference Engine abstract KV Cache Manager Speculative Decoder In-Flight Batch Scheduler LLM Generation Backend Inference Server Model Repository Continuous Integration Runner Evaluation Harness Inference Performance Analyzer Inference Metrics Endpoint Metrics Collector Custom Metrics Adapter Metric-Driven Autoscaler GPU Node abstract

Sources: Ch4.6 Ref4.01 Ref4.03

Constitutional AI alignment (SL-CAI + RLAIF)

When to choose. Choose when alignment must scale without extensive human labeling of harmful content and values must be explicit and inspectable by regulators and stakeholders.

Constitution abstract Constitution Authoring Board Red-Team Prompt Dataset Critique-Revision Generator Self-Reflection Critic Critique Rubric Critique-Revision Dataset Fine-Tuning Pipeline abstract Candidate Response Sampler AI-Feedback Preference Labeler Preference Dataset Reward Model Trainer Reward Model Constitutional Reward Scorer RLHF Policy Optimizer Constitutionally Aligned Model Foundation LLM abstract LLM Inference Service Training Pipeline Orchestrator Principle Adherence Evaluator

Alternative to: Human-feedback RLHF alignment, Hybrid Constitutional AI then RLHF refinement

Sources: Ch9.5

Constitutional defense-in-depth runtime

When to choose. Choose for production deployments of aligned models, since no single alignment method eliminates harmful outputs; combine training-time alignment with runtime rails, moderation, human escalation and monitoring.

Constitutionally Aligned Model LLM Inference Service Agent Controller abstract Guardrail Orchestrator Guardrail Policy Canonical Form Matcher Input Rail Dialog Rail Retrieval Rail Execution Rail Output Rail Third-Party Moderation Service Violation Response Policy abstract Escalation Handler Human Specialist Guardrail Violation Monitor Adversarial Robustness Evaluator Alignment Drift Monitor Audit Log Store

Sources: Ch9.5 Ref9.01

Containerised microservices agent deployment

When to choose. Choose when components have fundamentally different scaling needs (e.g., CPU retrieval vs GPU generation), multiple teams deploy independently, latency must be consistent with warm instances, and the organisation already operates Kubernetes.

Agent API Gateway Layer-7 Load Balancer Identity Provider Rate Limiter Container Orchestrator Metric-Driven Autoscaler Autoscaling Policy Container Health Prober Liveness Endpoint Intent Router Knowledge Retrieval Agent Answer Synthesizer abstract Escalation Handler Model Router Model Routing Policy Message Queue Circuit Breaker Retry Handler Trace Collector Log Aggregator Metrics Collector Alert Manager GPU Node abstract

Alternative to: Serverless event-driven agents

Sources: Ch4.2

Continual learning with experience replay

When to choose. Choose when a learned policy must acquire new skills in production without catastrophic forgetting of earlier ones.

Experience Replay Buffer Continual Learning Trainer Policy Network Event-Based Episode Encoder

Sources: Ch5.7

Continuous compliance automation

When to choose. Choose when compliance must be a continuous, embedded process with automated monitoring, enforcement, alerting and reporting, keeping humans for judgment.

Metrics Collector Log Aggregator Continuous Compliance Monitor Action Policy Engine Bias Evaluator Alert Manager Incident Manager Compliance Dashboard Compliance Report Generator Compliance Evidence Repository Compliance Gate Continuous Integration Runner

Sources: Ref9.09 Ref9.10

Continuous delivery with human release gate

When to choose. Choose when a human must approve promotion from staging to production after automated validation.

Continuous Integration Runner Static Code Analyzer Evaluation Harness Load Test Runner Regression Gate Container Image Builder Container Image Scanner Container and Model Artifact Registry Configuration Repository Staging Environment Release Approver Rollout Manager abstract

Alternative to: Fully automated continuous deployment

Sources: Ch4.1

Conversation-driven multi-agent collaboration

When to choose. Choose for multi-agent workflows where flexible dialogue-driven collaboration and readable transcripts aid development, exploration and human oversight; not for deterministic, auditable production decisions.

Conversational Agent Coordinator Worker Agent abstract Human Proxy Agent Code Execution Runner Conversation State Store abstract Dual-Agent Critic Shared Message History Termination Checker Iteration Limit Policy Execution Sandbox abstract LLM Inference Service Human Approver

Alternative to: Role-based hierarchical crew, Plugin-routed enterprise agent platform, Role-based sequential team, Role-based hierarchical team with manager quality gates

Sources: Ch2.1 Ch2.4

Cooperative iterative neural-symbolic refinement

When to choose. Choose when symbolic knowledge must direct further neural analysis to resolve ambiguity and iterative refinement justifies extra cost and latency (e.g., circuit-board defect diagnosis).

Iterative Refinement Coordinator Perception Interpreter Neural Perception Model Symbolic Logic Engine Iteration Limit Policy Termination Checker

Alternative to: Sequential neural-to-symbolic pipeline, Parallel paradigms with result fusion, Embedded neural modules within a symbolic program

Sources: Ch5.13

Coordinator-managed working-memory budgets

When to choose. Choose when multiple worker agents compete for limited context capacity and global allocation should override individual over-consumption, accepting coordinator overhead and a potential single point of failure.

Agent Context Budget Coordinator Worker Agent abstract

Sources: Ch5.9

Cost-optimised routed agent (router-first, slim context)

When to choose. Choose for high-volume workloads with power-law query complexity (most queries simple) where token and API costs threaten viability, e.g., loan pre-screening or omnichannel support.

Model Router Query Complexity Classifier Small Language Model Tier Large Language Model Tier LLM Inference Service KV Cache Manager Dynamic Tool Loader Tool Schema Trajectory Pruner Summarizing History Compressor Tool Result Cache Request Batcher Parallel Agent Coordinator Token Budget Enforcer Token Budget Policy Token Cost Meter Execution Profiler Evaluation Baseline Quality Drift Detector Alert Manager

Sources: Ch3.10

Cost-optimized agent configuration

When to choose. Choose for high-volume simple queries where ~84% accuracy suffices and per-query cost must be minimal (~$0.01).

Small Language Model Tier Inference Serving Configuration abstract Context Compressor Iteration Limit Policy

Sources: Ch3.4

Cost-optimized conversational agent (token economics)

When to choose. Choose for high-volume, multi-turn conversational agents with large static context (system prompt, tool definitions, user profile) whose token cost must fall without degrading quality; applied progressively (caching, retrieval, output limits, routing) with quality validation after each step.

Token Cost Meter Cost Rate Card Cost Attribution Aggregator Cost Reporting Dashboard Cacheable Prefix Prompt Layout KV Cache Manager Prompt Context Builder Vector Retriever abstract Output Token Limit Policy Output Format Specification Model Router Rule-Based Complexity Classifier Small Language Model Tier Large Language Model Tier Regression Gate

Sources: Ch8.3

Cost-optimized inference serving

When to choose. Choose under budget constraints or at large concurrent scale, accepting ~5% quality loss for 40-60% infrastructure savings.

Evaluation Harness Evaluation Dataset Small Language Model Tier Model Quantizer INT8 Quantized Engine Inference Server Metric-Driven Autoscaler Autoscaling Policy

Alternative to: Throughput-optimized batch inference serving, Latency-optimized interactive inference serving

Sources: Ch7.2

Cost-sensitive LLM serving

When to choose. Choose when cost per inference is the priority and some quality and latency can be traded away.

Small Language Model Tier Model Quantizer INT4 Quantized Engine Request Batcher Response Cache abstract Spot GPU Node Metric-Driven Autoscaler Scheduled Scaler Token Cost Meter

Alternative to: Quality-critical LLM serving

Sources: Ref7.04 Ref7.15 Ref7.17

CoT + Self-Consistency (majority vote)

When to choose. Choose for complex multi-step reasoning with definable correct answers (math, legal analysis, medical diagnosis) where base accuracy is insufficient and errors are path-specific; offers the best cost-accuracy trade-off without tree search or graph refinement.

Chain-of-Thought Prompt abstract Reasoning Path Sampler Self-Consistency Sampling Policy LLM Inference Service Final Answer Extractor Majority Vote Aggregator Confidence Estimator

Alternative to: Adaptive RASC with confidence-based escalation

Sources: Ch5.3

Curated multi-pass long-source analysis

When to choose. Choose when source material (large codebases, long contracts, multi-paper corpora) approaches or exceeds the context window or suffers lost-in-the-middle effects, and sequential-pass latency is acceptable.

Prompt Context Builder Working Memory Buffer abstract Vector Retriever abstract Reranker abstract Context Compressor Context Assembler abstract Reasoning Consistency Checker Reasoning Engine Episodic Memory Store Memory Retriever Context Budget Allocator Code Dependency Analyzer LLM Inference Service

Sources: Ch5.9

Customer-service transcript curation for support-agent training

When to choose. Choose when curating ASR-produced call transcripts to train a support agent: tolerate higher perplexity, use higher deduplication thresholds, and classify out sales/survey calls.

Speech Transcriber Partitioned Dataset Reader Document Quality Filter Perplexity Filter Exact Hash Deduplicator Near-Duplicate Detector Domain Relevance Classifier Document PII Redactor Data Curator Curated Training Corpus Fine-Tuning Pipeline abstract

Alternative to: GPU-accelerated web-corpus curation for a domain agent

Sources: Ch7.5

Decentralized (peer-to-peer) multi-agent orchestration

When to choose. Choose for large-scale systems where central coordination is infeasible, agent populations change frequently, and resilience to individual agent failure is critical.

Worker Agent abstract Agent Capability Registry Capability-Matching Allocator Discoverable Delegation Interface Agent Message Bus abstract Trace Collector

Alternative to: Centralized multi-agent orchestration, Hierarchical multi-agent orchestration, Federated multi-agent orchestration

Sources: Ch1.3

Decomposed RAG (query decomposition + parallel retrieval + synthesis)

When to choose. Choose for multi-part or comparison questions with several distinct information needs, when measured accuracy gains justify two extra LLM calls.

RAG Query Orchestrator Query Complexity Classifier Question Decomposer Decomposition Prompt Template Parallel Sub-Query Retrieval Controller Retriever abstract Multi-Hop Answer Synthesizer Synthesis Prompt Template LLM Inference Service

Alternative to: Standard single-query RAG

Sources: Ch6.6

Dedicated full-GPU single-tenant serving

When to choose. Choose for 70B+ models or one high-concurrency, latency-critical agent needing full memory bandwidth and large continuous batches.

Unpartitioned GPU Layout Dedicated GPU Device Inference Server In-Flight Batch Scheduler Paged KV Cache Allocator

Alternative to: Multi-tenant SaaS on uniform GPU partitions, Multi-tier platform on mixed GPU partitions

Sources: Ch7.6

Defense-in-depth reasoning verification

When to choose. Choose for high-stakes domains (finance, healthcare, legal) where plausible but wrong reasoning is unacceptable and automated verification must be layered with human review.

Reasoning Verifier abstract Fine-Tuned Step Verifier Business Rule Logic Verifier Citation Verifier Reasoning Consistency Checker Verifier Ensemble Aggregator Self-Reflection Critic Confidence Estimator Confidence Gate Human Specialist Trace Annotation Console

Sources: Ch3.6

Dependency-governed concurrent multi-agent workflow

When to choose. Choose when several specialist agents share workflow state and run many concurrent sessions in production (e.g., 50+), where deadlock, out-of-order execution and lost updates emerge that low-concurrency testing does not reveal.

Agent Dependency Graph Dependency Cycle Validator Phase Barrier Executor Worker Agent abstract Shared Blackboard Store Optimistic Lock Controller State Merge Policy Trace Context Propagator Trace Exporter Telemetry Gateway Trace Store Evaluation Failure Analyzer Bottleneck Analyzer

Sources: Ch8.2A

Design-time Pareto analysis with runtime scalarized utility

When to choose. Choose when objectives are incommensurable or stakeholders dispute weights: compute Pareto frontiers at design time, let stakeholders choose trade-offs, then deploy a scalarized utility for real-time operation.

Pareto Frontier Optimizer Pareto Frontier Set Explanation Presenter Decision Stakeholder Decision Sensitivity Analyzer Utility Function Specification Utility-Based Decision Maker Evaluation Harness Decision Scenario Simulator

Sources: Ch5.10

Deterministic auditable function-calling agent

When to choose. Choose when regulation requires identical decisions for identical inputs and clear decision trails, e.g., loan application processing calling credit, income and fraud services.

Direct Tool-Calling Controller Tool Schema Tool Executor External Service API Audit Log Store Native Function-Calling API abstract

Alternative to: Conversation-driven multi-agent collaboration

Sources: Ch2.1 Ch2.3

Development direct-export tracing

When to choose. Choose for development or small teams: agents export traces directly to a lightweight backend without a collector, inspected ad hoc.

Trace Collector Trace Schema Trace Exporter Trace Store Trace Visualizer Workflow Execution Debugger Agent Developer

Alternative to: Production centralized tracing

Sources: Ch3.6

Distributed billion-scale vector search

When to choose. Choose for enterprise deployments with billions of vectors needing horizontal scaling, high availability and GPU acceleration, operated by teams experienced with distributed systems and container orchestration.

Distributed Vector Index Store Cluster Consensus Coordinator Object Store Event Stream Log Container Orchestrator GPU Node abstract Vector Index Build Configuration abstract Approximate Vector Search Retriever

Alternative to: Highly available replicated vector store cluster

Sources: Ch6.2A

Distributed state for horizontally scaled agents

When to choose. Choose for long-running workflows (hours or days) or high-throughput systems with hundreds of concurrent requests that need multiple agent instances.

State-Graph Orchestrator Agent State Schema Working Memory Buffer abstract Distributed Cache State Store Database State Store State Concurrency Controller abstract State Retention Policy

Alternative to: Agent-local state

Sources: Ch1.6

Document + chart analysis agent

When to choose. Choose when critical quantitative data resides in charts within uploaded or indexed documents (e.g., earnings reports).

Document Ingestor Chart Data Extractor Vision-Language Model abstract Inference Server Agent Controller abstract LLM Inference Service

Sources: Ch7.5

Documented content moderation at scale

When to choose. Choose when an AI system makes millions of daily moderation decisions across languages and jurisdictions and must remain transparent, contestable and auditable.

Toxicity Classifier Regional Moderation Policy Confidence Gate Human Specialist Audit Log Store Moderation Appeal Tracker Bias Evaluator Model Card System Card Transparency Report Incident Manager

Sources: Ch9.8

DPO offline preference alignment

When to choose. Choose for pure preference learning from static datasets, particularly for organizations with limited RL expertise or computational resources.

Foundation LLM abstract Instruction Demonstration Dataset Fine-Tuning Pipeline abstract Reference Policy Model Alignment Prompt Dataset Candidate Response Sampler Preference Annotation Console Pairwise Comparison Format Preference Annotator abstract Annotation Quality Monitor Preference Dataset Direct Preference Optimizer Preference Optimization Config Fine-Tuned Agent Model

Alternative to: PPO-based RLHF alignment pipeline

Sources: Ch10.3

Embedded neural modules within a symbolic program

When to choose. Choose when deep integration (e.g., knowledge-grounded question answering) provides capabilities that justify reduced modularity and dual-paradigm expertise.

Symbolic Program Executor Natural Language to Logic Translator Vector Retriever abstract Symbolic Logic Engine Logic Conclusion Verbalizer Domain Ontology Knowledge Graph Store abstract

Alternative to: Sequential neural-to-symbolic pipeline, Parallel paradigms with result fusion, Cooperative iterative neural-symbolic refinement

Sources: Ch5.13

Ensemble (complexity-routed) scaling

When to choose. Choose when query complexity varies and most queries can be answered by smaller models, to cut cost 40-60% while escalating complex queries.

Model Router Query Complexity Classifier Embedding Service abstract Small Language Model Tier Standard Language Model Tier Large Language Model Tier LLM Inference Service

Alternative to: Model sharding across GPUs

Sources: Ch1.8

Explainable decision support (agent as assistant)

When to choose. Choose when an agent pre-screens or recommends high-stakes decisions (diagnostic imaging, loan screening, legal case prioritization, borderline moderation) while a human expert makes the final call; low-risk integration that is easy to turn off (Ref10.06 assistant pattern).

Decision Engine abstract Confidence Estimator Confidence Calibrator Trace Collector Trace Store Explanation Presenter Layered Explanation View abstract Proactive Explanation Expander Explanation Method Selector Attribution Analyzer Counterfactual Explainer Precedent Case Retriever Confidence Indicator Decision Override Control Acknowledgment Friction Gate Feedback Collector abstract Audit Log Store Override Pattern Analyzer Human Approver

Sources: Ch10.5 Ref10.06

Fact-checking-emphasis guardrails

When to choose. Choose for educational apps that emphasize factual accuracy while relaxing dialog restrictions.

Guardrail Policy Input Rail Output Rail Fact Checking Rail abstract LLM Inference Service

Alternative to: Full six-rail defense-in-depth

Sources: Ch7.1A

Fairness-assured high-stakes decision agent

When to choose. Choose for agents making consequential decisions about individuals (healthcare access, lending, hiring, criminal justice, education) where disparate impact creates legal and ethical exposure; combines layered fairness rails, group-wise auditing, continuous fairness monitoring with SLOs, active debiasing, human override and consent-governed demographic data.

Agent Controller abstract Predictive Decision Model Inference Server Guardrail Orchestrator Input Rail Retrieval Rail Output Rail Execution Rail Output Bias Detector abstract Bias Evaluator Demographic Audit Dataset Fairness Threshold Policy Fairness Monitor Alert Manager Bias Mitigator abstract Risk Gate Human Approver Decision Appeal Service Audit Log Store Fairness Auditor Demographic Data Store Purpose-Based Access Controller Consent Registry Model Card Oversight Governance Committee

Sources: Ch9.4 Ref9.02

Federated multi-agent orchestration

When to choose. Choose for regulated or multi-stakeholder settings where autonomous domains (organisations, partners) must keep local governance and coordinate only through contracts.

Supervisor Agent Worker Agent abstract Agent Service API abstract Agent Message Contract Schema Registry

Alternative to: Centralized multi-agent orchestration, Decentralized (peer-to-peer) multi-agent orchestration, Hierarchical multi-agent orchestration

Sources: Ch1.3

Feedback-driven continuous improvement flywheel

When to choose. Choose for production agents with real users where automated metrics miss satisfaction, comprehension and efficiency: collect explicit and implicit feedback, analyse themes and root causes, augment tests, refine, validate offline and via A/B, redeploy.

Feedback Collector abstract Behavioral Signal Tracker User Feedback Store Sentiment Classifier Feedback Theme Clusterer Feedback Prioritizer Evaluation Failure Analyzer Evaluation Dataset Evaluation Harness Regression Gate A/B Test Traffic Splitter Experiment Guardrail Monitor Rollout Manager abstract

Sources: Ch3.2

Few-shot in-context adaptation

When to choose. Choose for specialized domains or novel formats where zero-shot fails, high-quality demonstrations are available, and tasks are pattern-recognizable (classification, structured extraction, format transformation); not when knowledge is missing or deep behavioural change is needed.

System Prompt Template Prompt Exemplar Set Similarity Exemplar Selector Prompt Context Builder Trajectory Harvester Data Curator LLM Judge abstract LLM Inference Service Foundation LLM abstract Evaluation Harness

Alternative to: Retrieval-grounded adaptation, Fine-tuned specialist agent

Sources: Ch3.5

Financial services compliance output filtering

When to choose. Choose for banking/fintech agents subject to FINRA, FCRA and fair lending rules: layer deny-list, semantic compliance, PII filtering, disclaimers, bias detection and full audit trails.

Guardrail Orchestrator Guardrail Policy LLM Self-Check Output Rail Domain Compliance Rail abstract Pattern Compliance Checker Semantic Compliance Classifier Disclaimer Injector PII Redactor Deny-List Content Filter Output Bias Detector abstract Template Response Generator Audit Log Store Trace Collector Compliance Officer

Sources: Ch9.1

Fine-tuned specialist agent

When to choose. Choose only when behavioural-consistency problems resist prompt engineering and RAG, the domain is well structured, and roughly 1,000+ (ideally 5,000-10,000) quality trajectories are available.

Agent Trajectory Dataset Synthetic Data Generator Domain Expert Annotator Data Curator Continued Pretrainer Domain Text Corpus LoRA Fine-Tuner abstract LoRA Adapter Fine-Tuned Agent Model Training Pipeline Orchestrator Evaluation Harness Bias Evaluator Online Evaluator Rollout Manager abstract Inference Server System Prompt Template

Sources: Ch3.5

Full agentic system sandboxing

When to choose. Choose for high-risk deployments needing maximum containment even if the agent is compromised via prompt injection; requires secret injection and mediated production APIs.

Agent Controller abstract Tool Executor Code Execution Runner Execution Sandbox abstract Just-in-Time Credential Broker Secrets Vault Pod Network Policy Runtime Security Policy Enforcer Sandbox Anomaly Detector

Alternative to: Individual code snippet sandboxing

Sources: Ch9.3

Full six-rail defense-in-depth

When to choose. Choose for high-security environments (healthcare, finance) that accept 50-150ms latency as compliance cost.

Guardrail Policy Input Rail Dialog Rail Retrieval Rail Execution Rail Output Rail Fact Checking Rail abstract LLM Inference Service

Alternative to: Input/output rails only, Fact-checking-emphasis guardrails

Sources: Ch7.1A

Fully automated continuous deployment

When to choose. Choose when automated validation is trusted to promote passing builds from staging to production without a human gate.

Continuous Integration Runner Static Code Analyzer Evaluation Harness Load Test Runner Regression Gate Container Image Builder Container Image Scanner Container and Model Artifact Registry Configuration Repository Staging Environment Canary Rollout Controller

Alternative to: Continuous delivery with human release gate

Sources: Ch4.1

GDPR Art. 32 layered security for personal data and AI

When to choose. Choose when handling sensitive personal data (financial, health); protection scales with data sensitivity and processing risk.

Data Classification Policy Data Classifier Field-Level Encryptor Key Management Service Authorization Policy Decision Point Identity Provider Approval Gateway Audit Log Store Access Anomaly Detector Security Analyst Penetration Tester Incident Manager Incident Escalation Policy Personal Data Breach Notifier Personal Data Breach Register Input Rail PII Redactor Training Data Lineage Store Model and Agent Release Registry Agent API Gateway

Sources: Ch9.7 Ref9.05

GDPR data-subject rights and consent management

When to choose. Choose for any system processing personal data of EU residents, regardless of organization size or location: granular consent, rights-request fulfilment and retention enforcement.

Data Subject Privacy Portal Privacy Notice Consent Manager Consent Registry Consent Enforcement Gate Lawful Basis Record abstract Records of Processing Register Data Subject Request Handler Data Subject Identity Verifier Data Subject Request Queue Erasure Orchestrator Personal Data Exporter Personal Data Store Backup Archive Store Personal Data Retention Policy Data Retention Enforcer Data Processing Agreement External Service API Audit Log Store Data Protection Officer

Sources: Ch9.7 Ref9.05

GDPR-compliant automated decision-making

When to choose. Choose when an AI model makes decisions with legal or similarly significant effects on individuals (credit, insurance claims, hiring): DPIA, minimized inputs, fairness monitoring, borderline human review, explanations and appeals.

DPIA Record Purpose-Bound Data Schema Predictive Decision Model Decision Engine abstract Confidence Gate Human Approver Bias Evaluator Fairness Threshold Policy Proxy Feature Detector Decision Factor Explainer Decision Appeal Service Human Specialist Authorization Policy Decision Point Audit Log Store Data Protection Officer

Sources: Ch9.7 Ref9.03

Governed clinical decision-support AI (e.g., patient deterioration prediction)

When to choose. Choose when an AI system augments high-stakes clinical judgment under healthcare regulation: clinicians retain final authority, predictions are explained, performance and fairness are monitored continuously, and the system is withdrawn on threshold breach.

Oversight Governance Committee Accountable AI Executive AI Governance Policy AI Impact Assessment Record Harm Risk Register Compliance Requirement Crosswalk Bias Evaluator Fairness Threshold Policy Data Drift Detector Continual Learning Trainer Regression Gate Holdout Evaluation Set Approval Gateway Human Approver Feedback Collector abstract Audit Log Store Incident Manager AI Incident Classification Policy Feature Flag Service

Sources: Ch9.8

GPU-accelerated multimodal production stack

When to choose. Choose for production multimodal agents needing high throughput and low latency on self-hosted GPUs, with model-agnostic APIs, GPU vector search, health probing, metrics and zero-downtime model updates.

Ingestion Pipeline Orchestrator Inference Server OpenAI-Compatible Inference API Engine Builder Model Quantizer Optimized Inference Engine abstract Foundation LLM abstract Vision-Language Model abstract Chart-to-Table Model Joint Multimodal Embedding Service GPU-Accelerated Vector Index Store Metadata-Filtered Retriever GPU Node abstract Container Orchestrator Container Health Prober Load Balancer abstract Metrics Collector Rollout Manager abstract

Sources: Ch2.7 Ref2.07

GPU-accelerated web-corpus curation for a domain agent

When to choose. Choose when building a pretraining or domain corpus from large noisy web scrapes (billions of documents) for a single-language domain agent.

Partitioned Dataset Reader Language Identification Filter Document Quality Filter Perplexity Filter Exact Hash Deduplicator Near-Duplicate Detector Domain Relevance Classifier Document PII Redactor Synthetic Data Generator Data Curator GPU-Accelerated Dataframe Engine Curated Training Corpus

Alternative to: Customer-service transcript curation for support-agent training

Sources: Ch7.5

GPU-accelerated web-scale data curation

When to choose. Choose when processing billions of documents or terabytes where CPU pipelines take days, for regular full reprocessing and fuzzy deduplication at scale, especially with existing GPU clusters; stay on CPU below ~10 million documents.

Ingestion Pipeline Orchestrator Full Refresh Policy GPU-Accelerated Dataframe Engine GPU Node abstract Document Quality Filter Text Normalizer Exact Hash Deduplicator Near-Duplicate Detector Text Embedding Service Vector Batch Ingestor

Alternative to: Batch incremental multi-source knowledge ETL

Sources: Ch6.3B

GPU-enabled container cluster with telemetry

When to choose. Choose for production container clusters running GPU-accelerated workloads that need automated GPU enablement and GPU observability.

Container Orchestrator Accelerator Operator GPU Device Plugin Node Capability Labeler Accelerator Container Runtime GPU Telemetry Exporter Metrics Collector Metrics Dashboard abstract Alert Manager GPU Node abstract

Sources: Ref4.05 Ref4.07

Graph-based iterative workflow agent

When to choose. Choose when the workflow combines sequential stages with iterative refinement (generate-test-regenerate, feedback loops), conditional branching, and complex structured state requiring custom merge logic.

State-Graph Orchestrator Workflow State Graph Agent State Schema Rule-Based Transition Router State Checkpoint Store abstract Reasoning Engine Output Verifier Code Execution Runner Feedback Collector abstract Failure Analyzer Iteration Limit Policy Workflow Execution Debugger Execution Sandbox abstract Working Memory Buffer abstract

Alternative to: Conversation-driven multi-agent collaboration, Role-based hierarchical crew, Plugin-routed enterprise agent platform

Sources: Ch2.1 Ch2.2 Ch2.3

Graph-of-Thought synthesis reasoning module

When to choose. Choose when problems decompose into independent subproblems whose solutions must be merged and synthesis adds value beyond the best individual exploration.

Graph-of-Thought Controller Graph of Operations Thought Generator abstract Thought Aggregator Thought Refiner Thought State Evaluator abstract Graph Reasoning State Store Reasoning Graph Analyzer Procedural Memory Store Plan Template Extractor Reasoning Strategy Refiner State Checkpoint Store abstract LLM Inference Service

Alternative to: Tree-of-Thought deliberate reasoning module

Sources: Ch5.2

Graph-routed customer support agent

When to choose. Choose when inquiries follow a decision tree: classification determines which specialised branch handles the query, with multi-turn state and escalation to humans under sub-second latency.

State-Graph Orchestrator Workflow State Graph Agent State Schema Working Memory Buffer abstract Intent Router Rule-Based Transition Router Worker Agent abstract Escalation Handler Human Specialist System Prompt Template LLM Inference Service Foundation LLM abstract Optimized Inference Engine abstract Inference Serving Configuration abstract Trace Collector

Sources: Ch1.5B

Ground-to-text multimodal RAG with metadata

When to choose. Choose for information-dense visuals (financial reports, scientific charts, technical diagrams) requiring precise retrieval, accepting one-time preprocessing cost while keeping text retrieval infrastructure unchanged.

Ingestion Pipeline Orchestrator Document Ingestor Multimodal Content Router Heuristic Image Type Classifier VLM-based Image Type Classifier Image Captioner Chart Data Extractor Vision-Language Model abstract Chart-to-Table Model Semantic Boundary Chunker Text Embedding Service Vector Index Store abstract Source Media Store Multimodal Chunk Metadata Schema Grounded Text Retriever Multimodal Context Assembler Multimodal Answer Synthesizer Inference Server OpenAI-Compatible Inference API Throughput-Oriented Batching Config Latency-Oriented Batching Config

Alternative to: Unified embedding space multimodal RAG, Separate modality stores with cross-modal reranking

Sources: Ch2.7 Ref2.07

Guardrail-sandwiched ReAct RAG agent

When to choose. Choose when a retrieval-augmented agent must enforce safety and compliance on both user input and agent output.

Input Rail ReAct Agent Controller Retriever abstract LLM Inference Service OpenAI-Compatible Inference API Output Rail Tool Protocol Server

Sources: Ref7.14

Guardrails-wrapped self-hosted inference microservice

When to choose. Choose for production LLM applications needing optimized self-hosted inference plus layered runtime safety, with safety policies and inference engines updated independently.

Guardrail Orchestrator Guardrail Policy Input Rail Dialog Rail Retrieval Rail Output Rail Fact Checking Rail abstract OpenAI-Compatible Inference API Inference Server Inference Engine Selector Pre-compiled Engine Backend Portable LLM Runtime Backend Model-Specific Inference Image Metric-Driven Autoscaler Container Orchestrator

Sources: Ch7.1B

Heuristic fast path with rule-based escalation

When to choose. Choose when most cases are routine and must be decided in milliseconds while borderline, high-stakes cases still need careful multi-rule evaluation (e.g., loan approval).

Heuristic Decision Engine Heuristic Rule Set Rule-Based Decision Engine Symbolic Logic Engine Production Rule Base

Sources: Ch5.11

Hierarchical aggregation multi-agent

When to choose. Choose for very large decomposable tasks (e.g., multi-document synthesis) where no single agent should hold all raw inputs, accepting coordination overhead that grows with hierarchy depth.

Supervisor Agent Worker Agent abstract

Alternative to: Multi-agent shared context pool, Message passing with selective context sharing

Sources: Ch5.9

Hierarchical backbone with flat leaves

When to choose. Choose when high-level workflow structure benefits from HTN scaling but primitive-level operations need sophisticated local optimization (e.g., manufacturing batch sequencing plus machine scheduling).

HTN Planner Decomposition Method Library Task Network Local Constraint Scheduler Flat Planner Plan Executor

Alternative to: Progressive deepening planning

Sources: Ch5.4

Hierarchical multi-agent orchestration

When to choose. Choose for large systems that partition naturally into subsystems or mirror team structures, balancing central oversight with distributed execution.

Supervisor Agent Worker Agent abstract Agent Delegation Interface abstract Task State Ledger Trace Collector

Alternative to: Centralized multi-agent orchestration, Decentralized (peer-to-peer) multi-agent orchestration, Federated multi-agent orchestration

Sources: Ch1.3

Hierarchical multi-agent pool scaling

When to choose. Choose for complex workflows where specialised sub-agents with different load patterns justify independent pool scaling despite orchestration and monitoring overhead.

Supervisor Agent Worker Agent abstract Task Queue State Checkpoint Store abstract Metric-Driven Autoscaler Layer-7 Load Balancer Metrics Collector

Sources: Ch1.8

Hierarchical replanning

When to choose. Choose when execution outcomes can deviate from predictions (failures, unexpected effects, resource conflicts) and valid higher-level decisions should be preserved during recovery.

HTN Planner Task Network Plan Executor Plan Deviation Monitor Tool Error Classifier Retry Handler Replanner abstract

Sources: Ch5.4

Highly available replicated vector store cluster

When to choose. Choose when 99.9%+ uptime is required so single-node failures must not interrupt service; add shards and nodes when capacity, not only availability, must scale.

Self-Managed Vector Index Store Persistent Volume API Key Authenticator Vector Store gRPC API Gossip Membership Service Shard Replication Manager Shard Query Router Replication and Sharding Configuration Load Balancer abstract Readiness Endpoint Pod Network Policy Container Orchestrator Metrics Collector Alert Manager Vector Store Capacity Monitor

Sources: Ch6.2B

Human-feedback RLHF alignment

When to choose. Choose when capturing contextual nuance and diverse, empirical human preferences outweighs annotation cost, annotator exposure to harmful content and value opacity.

Preference Annotator abstract Preference Annotation Console Candidate Response Sampler Preference Dataset Reward Model Trainer Reward Model Preference Reward Scorer RLHF Policy Optimizer Foundation LLM abstract LLM Inference Service

Alternative to: Constitutional AI alignment (SL-CAI + RLAIF), Hybrid Constitutional AI then RLHF refinement

Sources: Ch9.5

Human-in-the-loop pre-execution approval

When to choose. Choose when consequential, often regulated decisions (coverage, large payments, production changes) must not execute without explicit human authorization, and request volumes fit planned reviewer capacity.

Agent Controller abstract Approval Proposal Builder Confidence Estimator Oversight Gate abstract Confidence Gate Risk Gate Exception Gate Escalation Threshold Policy Approval Gateway State-Graph Orchestrator State Checkpoint Store abstract Approval Request Store Approval Router Approval Authority Matrix Approver Notifier Approval Review Console Explanation Presenter Approval Outcome Router Approval Escalation Scheduler Escalation Chain Policy Approval SLA Policy Approval Timeout Fallback Policy abstract Action Policy Engine Human Approver Human Specialist Tool Executor Audit Log Store Oversight Performance Monitor Approval Pattern Auditor Adaptive Threshold Tuner Override Pattern Analyzer

Alternative to: Human-on-the-loop monitoring with reactive intervention

Sources: Ch10.4 Ref10.01

Human-in-the-loop synchronous approval

When to choose. Choose for high-consequence, irreversible, binary decisions (production deletions, large financial transfers, treatment changes, legal notices) where decision latency of hours is acceptable and volume is low.

Risk Gate Approval Gateway Approval Request Store Approval Router Approval Authority Matrix Approver Notifier Approval Queue API Approval Review Console Approval Escalation Scheduler Escalation Chain Policy Human Approver Audit Log Store

Alternative to: Human-on-the-loop asynchronous monitoring, Human-over-the-loop strategic governance

Sources: Ch9.2

Human-on-the-loop asynchronous monitoring

When to choose. Choose for medium-consequence reversible actions (content moderation, access changes, automated emails) at volumes where per-decision approval is impractical; unsuitable for irreversible actions.

Agent Controller abstract Audit Log Store Execution Monitor Console Human Supervisor Action Rollback Service

Alternative to: Human-in-the-loop synchronous approval, Human-over-the-loop strategic governance

Sources: Ch9.2

Human-on-the-loop monitoring with reactive intervention

When to choose. Choose for high-volume operations (thousands of transactions per hour) where per-action pre-approval is impractical: actions execute immediately under continuous observation and humans intervene after the fact when anomalies emerge.

Agent Controller abstract Decision Telemetry Event Schema Trace Collector Metrics Collector Log Aggregator Trace Store Time-Series Metrics Store Centralized Log Store Behavioral Baseline Builder Behavioral Baseline Agent Behavior Anomaly Detector Telemetry Change-Point Detector Agent Interaction Graph Analyzer Alert Manager Alert Rule Set Alert Consolidator Alert Suppression Rule Set Incident Manager Execution Monitor Console Trace Visualizer Change Impact Attributor Agent Intervention Controller Intervention Authority Policy Agent Circuit Breaker Guardrail Orchestrator Guardrail Violation Monitor Oversight Intensity Policy Human Supervisor Override Pattern Analyzer Audit Log Store

Alternative to: Human-in-the-loop pre-execution approval

Sources: Ch10.4

Human-over-the-loop strategic governance

When to choose. Choose for low-consequence routine actions at very high scale (report generation, retrieval, log analysis) where only aggregate patterns matter and detection delay until periodic review is acceptable.

Agent Controller abstract Audit Log Store Metrics Dashboard abstract Policy Adherence Evaluator Oversight Governance Committee

Alternative to: Human-in-the-loop synchronous approval, Human-on-the-loop asynchronous monitoring

Sources: Ch9.2

Hybrid 3D parallelism

When to choose. Choose for scaling very large models to thousands of GPUs: data parallel across nodes, tensor parallel within NVLink nodes, pipeline parallel for massive models.

Data Parallel Executor Tensor Parallel Executor Pipeline Parallel Executor Collective Communication Library GPU Switch Fabric Inter-Node Network Fabric GPU Node abstract

Sources: Ch7.1A

Hybrid Constitutional AI then RLHF refinement

When to choose. Choose (the text's recommended practice) to establish principled, scalable foundations via critique-revision and RLAIF, then refine contextual nuance with human-preference RLHF, alternating iteratively.

Constitution abstract Critique-Revision Generator Critique-Revision Dataset Fine-Tuning Pipeline abstract Candidate Response Sampler AI-Feedback Preference Labeler Preference Annotator abstract Preference Annotation Console Preference Dataset Reward Model Trainer Reward Model Constitutional Reward Scorer Preference Reward Scorer RLHF Policy Optimizer Constitutionally Aligned Model Training Pipeline Orchestrator Principle Adherence Evaluator Alignment Drift Monitor Direct Preference Optimizer Fine-Tuned Agent Model Content Safety Filter abstract

Alternative to: Constitutional AI alignment (SL-CAI + RLAIF), Human-feedback RLHF alignment

Sources: Ch9.5 Ch10.3

Hybrid cost-optimized customer service stack

When to choose. Choose for high-volume customer service with repetitive queries and stable prompts: combine prompt caching, complexity-based model routing, FAQ response caching and TTL tool-result caching.

KV Cache Manager Model Router Query Complexity Classifier Small Language Model Tier Large Language Model Tier Response Cache abstract Tool Result Cache Cache Policy

Sources: Ch3.4

Hybrid edge inference with cloud retraining

When to choose. Choose when time-critical decisions need edge inference (sub-50 ms, offline, privacy, bandwidth) while periodic cloud synchronization supports retraining, analytics and version control.

Edge Device abstract Edge Inference Runtime Edge-Optimized Model Edge Device Agent Edge Update Orchestrator Device Inventory Registry Model Quantizer Model Pruner Engine Builder Data Curator Training Pipeline Orchestrator Quality Drift Detector

Sources: Ch4.3

Hybrid RAG + fine-tuned model

When to choose. Choose when both current and domain knowledge are needed, highest accuracy is the priority and budget allows both fine-tuning and retrieval infrastructure (Ref6.01).

Domain Text Corpus Data Curator Fine-Tuning Pipeline abstract Fine-Tuned Agent Model Inference Server Retriever abstract Vector Index Store abstract Context Assembler abstract Answer Synthesizer abstract Knowledge Base Refresher Evaluation Harness Holdout Evaluation Set RAG Evaluator

Alternative to: Retrieval-grounded adaptation, Fine-tuned specialist agent

Sources: Ref6.01

Hybrid RAG + knowledge graph

When to choose. Choose when queries consistently need both semantic understanding and precise relationship tracking (financial compliance, fraud detection, research analysis); highest complexity (p95 ~180ms).

Document Chunker abstract Embedding Service abstract Vector Index Store abstract Vector Retriever abstract Entity Recognizer Entity Linker abstract Relation Extractor abstract Knowledge Graph Loader Property Graph Store Graph Retriever Hybrid Retriever abstract Context Assembler abstract Synthesis Prompt Template Answer Synthesizer abstract LLM Inference Service Subgraph Cache Circuit Breaker

Alternative to: Vector RAG only, Knowledge graph only (text-to-graph-query QA)

Sources: Ch1.7B

Hybrid task queue + event stream messaging

When to choose. Choose when a system needs both reliable task orchestration/request-response (broker queue) and high-throughput replayable event streaming for data pipelines.

Message Queue Event Stream Log Dead Letter Queue Event Retention Policy Worker Agent abstract

Sources: Ch4.1

Hybrid value alignment (hard constraints + learned nuance + continuous monitoring)

When to choose. Choose for consequential domains (healthcare, lending, autonomous vehicles): explicit rules for non-negotiable constraints combined with RLHF-learned contextual judgment, continuous drift monitoring and stakeholder feedback; the chapter's recommended approach.

Constitution abstract Principle Priority Policy Principle Conflict Register Operational Norm Set Formal Rule Specification Rule Constraint Filter Guardrail Policy Preference Annotator abstract Preference Dataset Reward Model Trainer Reward Model RLHF Policy Optimizer Fine-Tuned Agent Model Alignment Drift Monitor Principle Adherence Evaluator Red Team Tester User Complaint Intake Oversight Governance Committee Decision Stakeholder Bias Evaluator External Auditor

Alternative to: Top-down (rule/principle-based) value alignment, Bottom-up (learned) value alignment

Sources: Ch9.6

Hybrid vector + graph memory

When to choose. Choose when both broad semantic retrieval and precise structured reasoning are needed and justify maintaining two stores (e.g., medical diagnosis, legal research, literature synthesis, enterprise entity unification, episode timelines).

Semantic Memory Store Episodic Memory Store Vector Index Store abstract Knowledge Graph Store abstract Retrieval-Augmented Graph Retriever Graph-Constrained Vector Retriever Entity Recognizer Relation Extractor abstract Multi-Signal Entity Resolver Text Embedding Service

Alternative to: Vector-only memory, Knowledge-graph memory

Sources: Ch5.7 Ch5.8

Hybrid vector-plus-graph RAG

When to choose. Choose when answers need both rich semantic context from documents and precise quantitative or relational facts from a knowledge graph, grounding generation to reduce hallucination.

Text Embedding Service Parallel Fusion Retriever Vector Retriever abstract Vector Index Store abstract Text-to-Graph-Query Translator Graph Retriever Knowledge Graph Store abstract Context Assembler abstract Answer Synthesizer abstract LLM Inference Service

Sources: Ch5.13

Individual code snippet sandboxing

When to choose. Choose when generated code is the primary attack surface and the agent system itself is protected by input filtering and prompt isolation; simplest setup since only the sandbox service is configured.

Agent Controller abstract Sandbox Execution API Code Execution Runner Execution Sandbox abstract Static Security Scanner

Alternative to: Full agentic system sandboxing

Sources: Ch9.3

Input/output rails only

When to choose. Choose for consumer apps with non-state-modifying agents, where execution rails can be disabled.

Guardrail Policy Input Rail Output Rail LLM Inference Service

Alternative to: Full six-rail defense-in-depth

Sources: Ch7.1A

Integrated memory-perception agent

When to choose. Choose for stateful, personalised interactions where the agent must resolve implicit references to past interactions, learn from successful resolutions, and combine current perception with separate episodic, semantic and procedural memory.

Agent Controller abstract Perception Interpreter Memory Retriever Memory Consolidator Memory Write Validator Memory Lifecycle Manager Working Memory Buffer abstract State Checkpoint Store abstract Context Window Manager abstract Semantic Memory Store Episodic Memory Store Procedural Memory Store LLM Inference Service Embedding Service abstract

Sources: Ch1.4

Integrated NIST AI RMF + ISO/IEC 42001 governance program

When to choose. Choose when an organization needs both operationally agile AI risk management and a formal, certifiable management system with accountability and external validation, integrating sector regulations into one program.

Oversight Governance Committee Accountable AI Executive AI Governance Policy Risk Tolerance Policy AI Use Inventory AI Risk Tier Classifier AI System Context Record AI Impact Assessment Record Harm Risk Register Risk Scorer abstract Risk Monitor Risk Treatment Plan AI Control Catalog Compliance Requirement Crosswalk Control Effectiveness Evaluator Continuous Compliance Monitor Compliance Evidence Collector Compliance Evidence Repository Internal Auditor External Auditor Remediation Tracker Model Card Dataset Datasheet System Card Transparency Report Audit Log Store Incident Manager AI Incident Classification Policy

Alternative to: Proportionate small/medium-organization AI governance

Sources: Ch9.8 Ref9.06 Ref9.07

Knowledge graph only (text-to-graph-query QA)

When to choose. Choose for relationship queries and multi-hop traversal over well-modeled entities (p95 ~80ms, medium complexity).

Document Chunker abstract Entity Recognizer Entity Linker abstract Relation Extractor abstract Knowledge Graph Loader Property Graph Store Graph Schema Introspector Graph Schema Text-to-Graph-Query Translator Graph Retriever Answer Synthesizer abstract LLM Inference Service

Alternative to: Vector RAG only, Hybrid RAG + knowledge graph

Sources: Ch1.7A Ch1.7B

Knowledge-graph memory

When to choose. Choose when reasoning requires explicit relationships, multi-hop inference or guaranteed consistency and knowledge is already structured (e.g., supply chain planning, clinical decision support).

Semantic Memory Store Knowledge Graph Store abstract Knowledge Graph Rule Set Graph Rule Inferencer Graph Consistency Validator Graph Retriever Knowledge Graph Loader Subgraph Cache

Alternative to: Vector-only memory, Hybrid vector + graph memory

Sources: Ch5.7 Ch5.8

Latency-critical HITL approval (e.g., real-time fraud review)

When to choose. Choose when approvals must complete within seconds (real-time trading, emergency response, fraud prevention): dedicated reviewer pools, simplified criteria, fast-track escalation, pre-approved categories and one-click interfaces.

Confidence Gate Approval Gateway Pre-Approved Category Policy Approval Router Approver Notifier Approval Review Console Approval SLA Policy Oversight Performance Monitor Approval Escalation Scheduler Human Approver

Alternative to: Throughput-critical HITL approval (e.g., batch claims, bulk onboarding)

Sources: Ch10.4

Latency-critical LLM serving

When to choose. Choose when time to first token is the priority (interactive chat, code completion).

LLM Inference Service Speculative Decoder Speculative Draft Model abstract Latency-Oriented Batching Config Inference Queue Policy Latency-Optimized Deployment Profile

Alternative to: Throughput-critical LLM serving, Cost-sensitive LLM serving, Quality-critical LLM serving

Sources: Ref7.02 Ref7.15

Latency-critical low-batch serving

When to choose. Choose for premium single users with batch size 1-4, predictable sequence lengths (e.g., <512-token translation) and strict per-token latency SLAs, when GPU memory is abundant.

Engine Builder Engine Build Configuration Contiguous KV Cache Allocator Fused Block-wise Attention Kernel Inference Server

Alternative to: Throughput-critical quantized LLM serving

Sources: Ch7.4

Latency-critical templated agent

When to choose. Choose when latency is a hard workflow constraint (e.g., sub-2-second clinical documentation) and task types are recurring enough for pre-compiled context.

Model Router Small Language Model Tier Large Language Model Tier LLM Inference Service Precompiled Context Template Prompt Context Builder Conversation State Store abstract Response Streamer Token Cost Meter Metrics Collector

Sources: Ch3.10

Latency-optimized agent configuration

When to choose. Choose for interactive, real-time applications where sub-2-second response is mandatory (small model, compressed context, 5 iterations).

Small Language Model Tier Inference Serving Configuration abstract Context Compressor Iteration Limit Policy

Alternative to: Cost-optimized agent configuration

Sources: Ch3.4

Latency-optimized interactive inference serving

When to choose. Choose for interactive applications (chat, code completion, real-time moderation) needing sub-200ms to sub-300ms P95 latency.

Inference Server Latency-Oriented Batching Config Latency-Tuned Decoding Configuration Response Streamer Small Language Model Tier INT8 Quantized Engine GPU Node abstract

Alternative to: Throughput-optimized batch inference serving, Cost-optimized inference serving

Sources: Ch7.2

Layered hallucination defense (regulated domain)

When to choose. Choose for customer-facing agents in regulated or high-stakes domains (finance, healthcare, legal) where critical hallucinations require near-zero tolerance and human oversight.

Knowledge Base Auditor Knowledge Base Refresher Content Freshness Monitor Query Rewriter Dense-Sparse Hybrid Retriever Reranker abstract Retrieval Confidence Filter System Prompt Template Stepwise Reasoning Verifier Reasoning Consistency Checker Dual-Agent Critic Response Confidence Modulator Output Verifier Output Format Specification Token Uncertainty Scorer Citation Verifier LLM Judge abstract Human Evaluator Confidence Gate Human Approver Online Evaluator Evaluation Trace Sampler Quality Drift Detector Hallucination Threshold Policy Failure Taxonomy Evaluation Failure Analyzer Feedback Collector abstract

Sources: Ch3.10

Layered multi-environment benchmarking

When to choose. Choose when selecting models or reasoning architectures for heterogeneous workflows: screen on a general multi-environment benchmark, then domain benchmarks, then monitored pilots.

Benchmark Suite Manifest Benchmark Environment abstract Environment Snapshot Simulated User Agent Evaluation Harness State Outcome Scorer Trajectory Scorer Evaluation Score Aggregator Evaluation Protocol Statistical Comparator Evaluation Failure Analyzer

Sources: Ch3.2

Layered production RAG service

When to choose. Choose when a RAG prototype must meet production SLAs (e.g., P95 < 2 s, 99.9% uptime, < 0.1% errors, per-query cost targets) at millions of documents and concurrent load.

Ingestion Pipeline Orchestrator Source Change Detector Text Embedding Service Vector Index Store abstract Document Metadata Store Lexical Index Store Knowledge Base Backup Service RAG Query Orchestrator Dense-Sparse Hybrid Retriever Reranker abstract Context Assembler abstract Answer Synthesizer abstract Citation Extractor LLM Inference Service Embedding Cache Retrieval Result Cache Distributed Response Cache Cache Invalidator Circuit Breaker Retry Handler Graceful Degradation Manager API Gateway Proxy abstract REST Agent API Load Balancer abstract Autoscaler abstract Metrics Collector Log Aggregator Trace Collector Alert Manager

Sources: Ch6.5

Layered-CoT multi-agent research synthesis

When to choose. Choose when a task requires coordinated expertise across several specialised domains, synthesis of contradictory findings, and memory of prior related work; avoid for simple lookups or fixed procedural tasks.

Supervisor Agent Task Planner abstract Execution Plan abstract Agent Capability Registry Parallel Agent Coordinator Worker Agent abstract ReAct Agent Controller Reasoning Engine Tool Executor External Service API Findings Synthesis Agent Working Memory Buffer abstract Memory Consolidator Memory Retriever Episodic Memory Store Semantic Memory Store Procedural Memory Store

Sources: Ch5.1

Layered-resilience multi-tool agent

When to choose. Choose for production agents depending on rate-limited external APIs and multiple LLM providers (e.g., financial research during peak load) that must keep partial functionality and an audit trail of degradation events through transient errors, provider outages and multi-component failures.

Retry Handler Retry Policy Model Router Fallback Chain Policy LLM Inference Service Fallback LLM Inference Service Response Cache abstract Graceful Degradation Manager Capability Tier Map Reasoning Engine Rule-Based Analyzer Embedded Critical Checker Dependency Health Monitor Circuit Breaker Circuit Breaker Policy Tool Result Cache External Service API Error Presenter Audit Log Store Metrics Collector Alert Manager Platform Operator

Sources: Ch2.8

Logic-augmented reasoning (Logic Agent)

When to choose. Choose for formal domains with clear inference rules (legal, mathematical proofs, regulatory compliance, security analysis) where logical validity is non-negotiable.

Reasoning Engine Natural Language to Logic Translator Symbolic Logic Engine Logic Rule Library Logic Conclusion Verbalizer

Alternative to: Structured Chain-of-Thought reasoning with layered verification

Sources: Ch3.9

Long multi-turn assistant with hierarchical history compression

When to choose. Choose for support-style conversations extending to tens of turns in which early turns establish facts (account type, attempted fixes) referenced later.

Prompt Context Builder Working Memory Buffer abstract Hierarchical History Compressor Conversation State Store abstract Compression Fidelity Validator Context Budget Allocator Context Allocation Policy Retriever abstract Reasoning Engine LLM Inference Service Memory Consolidator Episodic Memory Store

Sources: Ch5.9

Long-context, variable-length LLM serving

When to choose. Choose when inference is attention-dominated on long inputs and generation lengths vary widely (e.g., contract analysis), to raise throughput without extra GPUs.

Inference Server Optimized Inference Engine abstract Fused Block-wise Attention Kernel Paged KV Cache Allocator KV Cache Store KV Cache Manager Engine Builder Engine Build Configuration Inference Engine Profiler GPU Node abstract

Sources: Ch4.4

Memory-augmented ReAct agent with episodic memory

When to choose. Choose when an agent handles recurring situations across sessions (support, multi-turn debugging) and should reason from past episodes and trajectories rather than generic procedures.

ReAct Agent Controller Memory Retriever Prompt Context Builder Episodic Memory Store Significance-Based Episode Encoder Memory Consolidator Episode Summarizer Episode Pattern Abstractor Multi-Signal Relevance Ranker Vector Index Store abstract Text Embedding Service Procedural Memory Store

Sources: Ch5.7

Message passing with selective context sharing

When to choose. Choose when agents hand off sequential subtasks and task dependencies are understood well enough to identify the minimal context each receiver needs.

Worker Agent abstract Agent Message Bus abstract

Alternative to: Multi-agent shared context pool, Hierarchical aggregation multi-agent

Sources: Ch5.9

Minimal single-cluster agent deployment

When to choose. Choose for early or small deployments (e.g., ~100 internal users) before demand variability or observability gaps justify autoscaling, service mesh or multi-region.

Container Orchestrator Stateless Workload Controller Deployment Manifest Layer-4 Load Balancer Container Health Prober

Alternative to: Production multi-agent Kubernetes deployment

Sources: Ch4.3

Mixed-initiative human-agent collaboration

When to choose. Choose when humans and agents share a task and initiative must transfer between them based on risk, confidence and expertise, requiring explicit responsibility allocation, synchronized handoff context and mutual performance monitoring.

Mixed-Initiative Controller abstract Negotiated Initiative Controller Fixed Subtask Initiative Controller Subdialogue Initiative Controller Responsibility Matrix Autonomy Scope Adjuster Handoff Context Packager Escalation Handoff Package State Checkpoint Store abstract Override Rationale Log Override Pattern Analyzer Reviewer Fatigue Monitor Confidence Basis Explainer Interaction Style Adapter Engagement Estimator Negotiation Dialogue Manager Consensus Decision Protocol Human Specialist

Sources: Ch10.2

Model sharding across GPUs

When to choose. Choose only when the model exceeds single-GPU memory (>80 GB) after quantization and its quality gain justifies interconnect cost and complexity.

LLM Inference Service Large Language Model Tier GPU Node abstract

Alternative to: Ensemble (complexity-routed) scaling

Sources: Ch1.8

Multi-agent coordination over single-agent subflows

When to choose. Choose for production systems where multi-agent coordination handles the high-level workflow while complex subtasks (multimodal document processing, retrieval) need the stronger state management, refinement loops and checkpointing of single-agent graph or sequential workflows.

Multi-Agent Coordinator abstract Worker Agent abstract State-Graph Orchestrator State Checkpoint Store abstract Knowledge Retrieval Agent

Sources: Ch2.4

Multi-agent shared context pool

When to choose. Choose when collaborating agents must reason over identical, consistently updated information and redundant storage of common context should be eliminated.

Shared Blackboard Store Worker Agent abstract State Concurrency Controller abstract

Alternative to: Hierarchical aggregation multi-agent, Message passing with selective context sharing

Sources: Ch5.9

Multi-hop question answering agent

When to choose. Choose when answers require integrating evidence across multiple documents or source systems with auditable, source-attributed reasoning (research synthesis, enterprise knowledge, medical literature).

Multi-Hop Retrieval Controller Question Decomposer Query Rewriter Knowledge Source Router Retriever abstract Vector Index Store abstract Knowledge Graph Store abstract Supporting Fact Extractor Cross-Source Consistency Verifier Multi-Hop Answer Synthesizer Confidence Estimator Working Memory Buffer abstract Action Policy Engine Explanation Presenter Web Navigation Agent

Sources: Ch3.3

Multi-layer caching

When to choose. Choose when cost optimization justifies complexity, access patterns are diverse, and 80-90% hit rates confer advantage at large scale.

In-Process Response Cache Distributed Response Cache Semantic Cache KV Cache Manager Cache Policy Cache Invalidator Cache Warmer Embedding Service abstract

Sources: Ch1.8

Multi-layered hybrid replanning agent

When to choose. Choose for embodied or long-running agents in dynamic environments facing a mix of safety-critical, predictable, localized and global failures across millisecond-to-hour timescales.

Plan-and-Execute Controller Optimal Heuristic Search Planner Geometric Distance Heuristic State-Space Graph Contingency Planner Conditional Execution Plan Plan Executor Plan Deviation Monitor Discrepancy Significance Evaluator Replanning Strategy Router Reflexive Safety Controller Contingency Branch Activator Incremental Search Replanner Persistent Search Tree Store Complete Replanner Replanning Layer Arbiter Execution Failure History Store

Alternative to: Single-strategy reactive replanning agent

Sources: Ch5.6

Multi-model images in development, model-specific images in production

When to choose. Choose when an organization needs both innovation velocity (research/dev) and production reliability (SLA-bound services).

Multi-Model Inference Image Model-Specific Inference Image Model Integrity Validator Dynamic Model Loader Inference Server

Sources: Ch7.1B

Multi-tenant SaaS on uniform GPU partitions

When to choose. Choose when many tenants run similar-sized models (e.g., fine-tuned 7B INT8) with similar throughput and strict per-tenant SLAs.

Uniform GPU Partition Layout GPU Partition GPU Partition Manager Accelerator Operator GPU Device Plugin Inference Server INT8 Quantized Engine GPU Telemetry Exporter

Alternative to: Multi-tier platform on mixed GPU partitions, Dedicated full-GPU single-tenant serving, Time-sliced shared GPU

Sources: Ch7.6

Multi-tier caching for RAG agents

When to choose. Choose for RAG agents with repeated or structurally similar queries where reasoning templates, query embeddings and idempotent tool results can be reused.

Reasoning Chain Cache Plan Template Extractor Plan Template Adapter Embedding Cache Embedding Service abstract Vector Index Store abstract Tool Result Cache Cache Invalidator Cache Dependency Index Cache Policy Tool Executor

Sources: Ch4.7

Multi-tier platform on mixed GPU partitions

When to choose. Choose for tiered offerings with diverse model sizes (7B-70B), or when SLA-critical agents share a cluster with batch jobs on separate partitions.

Mixed GPU Partition Layout GPU Partition GPU Partition Manager GPU Partition Reconfiguration Policy Accelerator Operator GPU Device Plugin Inference Server GPU Telemetry Exporter Alert Rule Set Platform Operator

Alternative to: Multi-tenant SaaS on uniform GPU partitions, Dedicated full-GPU single-tenant serving

Sources: Ch7.6

Multimodal sensor perception

When to choose. Choose for agents in rich physical or multimodal environments that must fuse vision, LiDAR, audio, sensor and text inputs into a coherent world model.

Sensor Input Adapter Signal Preprocessor Sensor Stream Synchronizer Multimodal Fusion Engine Perception Interpreter World Model State Agent Controller abstract

Sources: Ch1.4

Neural-guided MCTS (AlphaGo/AlphaZero style)

When to choose. Choose when expert data exists or self-play can generate it and upfront GPU training cost is acceptable, to reach stronger decisions with far fewer simulations.

MCTS Planner MCTS Search Tree Store Search Tree Pruner Value Network Evaluator Value Network Action Prior Estimator Policy Network Policy/Value Network Trainer Expert Demonstration Dataset Environment Simulator MCTS Search Configuration Search Budget Policy GPU Node abstract

Sources: Ch5.5

Offline reasoning evaluation pipeline

When to choose. Choose during development for comprehensive reasoning quality assessment across extensive test suites without production latency constraints.

Trace Collector Reasoning Trace Schema Reference Reasoning Dataset Evaluation Harness Reasoning Chain Decomposer Entailment Step Validator Reasoning Consistency Checker Information Gain Scorer Goal Alignment Checker Reasoning Quality Scorer abstract LLM Judge abstract Human Evaluator Evaluator Calibrator Evaluation Result Analyzer Tool Fault Injector Quality Improvement Backlog

Sources: Ch3.9

Offline-online evaluation flywheel

When to choose. Choose for any production agent needing continuous improvement: offline baselines and CI regression gating, online monitoring on sampled traffic, and production failures fed back into versioned test sets.

Evaluation Harness Evaluation Dataset Continuous Integration Runner Task Success Evaluator abstract Policy Adherence Evaluator Simulated Web Environment Simulated User Agent Synthetic Scenario Generator Synthetic Dataset Online Evaluator Evaluation Sampling Policy abstract LLM Judge abstract Judge Adversarial Tester Human Evaluator Evaluation Rubric Feedback Collector abstract Behavioral Signal Tracker Agent Behavior Anomaly Detector Agent Version Experimenter abstract Evaluation Failure Analyzer Failure Case Curator Trace Store

Sources: Ch3.3

Offline-only continuous evaluation

When to choose. Choose for low-traffic agents (e.g., ~50 queries/day would need ~33 days per A/B test) or frequent small changes; supplement with manual monitoring of production metrics.

Continuous Integration Runner Evaluation Harness Evaluation Dataset Response Scorer abstract Experiment Tracker Evaluation Result Store Evaluation Baseline Statistical Comparator Regression Gate Regression Threshold Policy Evaluation Report Publisher Metrics Collector

Alternative to: Progressive offline-to-online evaluation pipeline

Sources: Ch3.1A Ch3.1B

Online reasoning quality monitoring

When to choose. Choose in production where latency budgets and cost preclude evaluating every inference; complements the offline pipeline.

Trace Collector Trace Sampling Policy Trace Store Ephemeral Trace Buffer Evaluation Trace Sampler Query Novelty Detector Feedback Collector abstract Behavioral Signal Tracker Reasoning Chain Decomposer Reasoning Quality Scorer abstract LLM Judge abstract Metrics Collector Quality Drift Detector Alert Manager Reasoning Quality Threshold Configuration

Sources: Ch3.9

Orchestrated core with event-driven periphery (multi-agent customer support)

When to choose. Choose when core processing needs predictable, centrally error-handled sequencing while non-critical notifications must fan out to many consumers without coupling.

Supervisor Agent Request Intake Agent Knowledge Retrieval Agent Intent Router Answer Synthesizer abstract Escalation Agent Static Delegation Interface Tool Protocol Server Embedding Service abstract Vector Index Store abstract LLM Inference Service Publish-Subscribe Bus Proactive Notifier Human Specialist Output Verifier Retrieval Result Cache Rule-Based Intent Classifier Template Response Generator Fallback Work Queue Trace Collector Metrics Collector

Sources: Ch1.3

Parallel paradigms with result fusion

When to choose. Choose when different paradigms see complementary aspects of the same input and redundancy/diversity justify multiplied compute cost (e.g., credit risk assessment).

Learned-Policy Decision Engine Utility-Based Decision Maker Rule-Based Decision Engine Decision Fusion Aggregator abstract Formal Rule Specification Utility Function Specification

Alternative to: Sequential neural-to-symbolic pipeline, Cooperative iterative neural-symbolic refinement, Embedded neural modules within a symbolic program

Sources: Ch5.13

Per-region full agent stacks with nearest-region routing

When to choose. Choose only when compliance requires geographic redundancy or a global user base needs low-latency regional endpoints.

DNS Load Balancer Load Balancer abstract Agent Controller abstract External Session State Store Vector Index Store abstract Distributed Response Cache Container Orchestrator

Alternative to: Single-region, multi-availability-zone deployment

Sources: Ch4.7

Performance proxy with selective API management

When to choose. Choose when most traffic needs maximum routing throughput but a subset of critical endpoints requires centralised authentication, rate limiting and policy enforcement.

High-Performance Reverse Proxy Policy Plugin Gateway Rate Limiter Authorization Policy Decision Point Agent API Gateway

Sources: Ch4.1

Plan-and-Execute structured workflow agent

When to choose. Choose for complex, mostly predictable multi-step workflows in stable environments where minimising expensive LLM calls matters.

Plan-and-Execute Controller Task Planner abstract Plan Executor Replanner abstract Execution Plan abstract Parallel Tool Dispatcher LLM Inference Service Foundation LLM abstract Tool Registry Tool Executor

Alternative to: Single-agent ReAct tool-using agent

Sources: Ch1.2

Plan-and-Execute with lean state

When to choose. Choose when the task prioritises efficiency; lean state (plan plus current step results) reduces context consumption.

State-Graph Orchestrator Task Planner abstract Plan Executor Replanner abstract Working Memory Buffer abstract Tool Executor

Alternative to: ReAct agent with rich state

Sources: Ch1.5A Ch1.6

Plugin-orchestrated enterprise agent

When to choose. Choose for enterprise agents integrating dozens of internal systems whose capability catalog grows over time, where plugins are reused across multiple agents and centralized observability, audit and credential control are required; avoid for focused agents with fewer than five capabilities or for strictly deterministic, graph-structured workflows.

Function-Calling Controller Function Choice Policy Tool Registry Tool Schema Tool Executor Parallel Tool Dispatcher Service Container Semantic Function Prompt Function Template REST API Adapter Database Connector LLM Inference Service Agent Workflow Configuration Secrets Vault Metrics Collector Token Cost Meter Trace Collector Audit Log Store

Sources: Ch2.5

Plugin-routed enterprise agent platform

When to choose. Choose for agent ecosystems with many specialized capabilities requiring dynamic routing, enterprise system integration and centralized observability, especially within an existing Azure/Microsoft estate.

Plugin Kernel Orchestrator Tool Registry LLM Inference Service Token Cost Meter

Sources: Ch2.1

Policy-governed human-over-the-loop

When to choose. Choose for enterprise agents making hundreds or thousands of decisions daily where real-time approval is impractical: humans author tiered policies and decision boundaries, the policy engine enforces them automatically, and only boundary crossings, low-confidence cases and irreversible actions reach humans.

Action Policy Engine Policy Context Aggregator Parameter Security Validator Organizational Baseline Policy Departmental Policy Team Policy Agent-Specific Policy Action Risk Tier Policy Approval Authority Matrix Policy Enforcement Mode Configuration abstract Policy Violation Responder Escalation Protocol abstract Confidence Estimator Confidence Gate Approval Gateway Human Approver Approval Timeout Fallback Policy abstract Trace Collector Audit Log Store Governance Decision Log Decision Factor Explainer Feedback Collector abstract Override Feedback Record Reward Model Trainer Override Rate Monitor Quality Drift Detector Fairness Monitor Human Oversight Protocol

Alternative to: Human-in-the-loop synchronous approval, Human-on-the-loop asynchronous monitoring

Sources: Ch10.5

PPO-based RLHF alignment pipeline

When to choose. Choose when alignment requires online learning, environmental interaction or integration with complex RL frameworks, and the organization can afford multi-model GPU memory and RL expertise.

Foundation LLM abstract Instruction Demonstration Dataset Fine-Tuning Pipeline abstract Reference Policy Model Alignment Prompt Dataset Candidate Response Sampler Annotation Task Router Preference Annotation Console Pairwise Comparison Format Preference Annotator abstract Annotation Guideline Annotation Quality Monitor Preference Label Aggregator Preference Agreement Filter Preference Dataset Reward Model Trainer Reward Model Preference Reward Scorer RLHF Policy Optimizer Value Network Preference Optimization Config Reward Hacking Monitor Fine-Tuned Agent Model Training Pipeline Orchestrator Distributed Model Executor abstract

Alternative to: DPO offline preference alignment

Sources: Ch10.3

Pre-optimized LLM microservice on Kubernetes

When to choose. Choose when agents need production LLM serving through OpenAI-compatible APIs with minimal inference-optimisation expertise, on a Kubernetes cluster with suitable GPUs.

LLM Inference Service OpenAI-Compatible Inference API Model Deployment Profile abstract In-Flight Batch Scheduler Inference Service Operator Deployment Manifest Container Orchestrator GPU Node abstract GPU Device Plugin Secrets Vault Container and Model Artifact Registry Layer-4 Load Balancer Agent API Gateway Liveness Endpoint Readiness Endpoint Container Health Prober Metric-Driven Autoscaler Autoscaling Policy Metrics Collector LLM Provider Adapter

Sources: Ch4.5

Preference-aligned (RLHF/DPO) agent

When to choose. Choose when desired behaviour involves context-dependent trade-offs (speed vs thoroughness, policy vs satisfaction, escalation vs autonomy) that no single demonstration captures.

Fine-Tuned Agent Model Candidate Response Sampler Preference Annotator abstract Preference Annotation Console Annotation Guideline Annotation Quality Monitor Preference Dataset Reward Model Trainer Reward Model Composite Reward Scorer Factuality Verifier abstract RLHF Policy Optimizer Direct Preference Optimizer Training Pipeline Orchestrator Training Hyperparameter Tuner Evaluation Harness Bias Evaluator Online Evaluator

Sources: Ch3.5

Proactive (push-based) assistant

When to choose. Choose when an agent must anticipate needs and initiate timely assistance before users ask (push-based), with graduated autonomy, granular consent and a feedback flywheel; avoid universal proactivity where interruptions cost more than they deliver.

Proactive Agent Situational Context Integrator Environmental Context Adapter User Affect Detector User Need Predictor User Trajectory Model Intervention Value Estimator Intervention Threshold Policy Intervention Timing Optimizer Notification Volume Governor Proactive Notifier Action Suggestion Engine User Activity History Store User Preference Profile Store User Exception Catalog Episodic Memory Store Confidence Gate Risk Gate Approval Gateway Post-Action Review Sampler Action Rollback Service Decision Factor Explainer Explanation Presenter Consent Manager Consent Registry Consent Enforcement Gate Audit Log Store Behavioral Signal Tracker Revealed Preference Learner Training Pipeline Orchestrator Value Alignment Criteria Human Approver

Sources: Ch10.2

Production centralized tracing

When to choose. Choose for production and enterprise deployments requiring centralized collection, long-term storage, cross-agent correlation, real-time alerting and compliance/audit integration.

Trace Collector Trace Schema Trace Context Propagator Trace Exporter Telemetry Gateway Trace Sampling Policy Trace Store Trace Visualizer Trace Error Classifier Agent Behavior Anomaly Detector Execution Profiler Token Cost Meter Alert Manager Platform Operator

Alternative to: Development direct-export tracing

Sources: Ch3.6

Production multi-agent Kubernetes deployment

When to choose. Choose when a multi-agent system faces variable demand, GPU inference workers and stateful memory stores requiring self-healing, autoscaling, isolation and monitoring.

Container Orchestrator Supervisor Agent Worker Agent abstract Stateless Workload Controller Stateful Workload Controller Deployment Manifest Metric-Driven Autoscaler Autoscaling Policy Container Health Prober Layer-4 Load Balancer Persistent Volume Model Artifact Cache GPU Node abstract Cluster Namespace Namespace Resource Quota Pod Network Policy Metrics Collector Metrics Dashboard abstract Alert Manager

Alternative to: Minimal single-cluster agent deployment

Sources: Ch4.3

Profile-optimize-benchmark agent performance lifecycle

When to choose. Choose for production agents that evolve continuously (prompt, model, framework changes) and must keep accuracy, latency and cost SLAs; baseline-profile, optimize, benchmark continuously and re-profile.

Trace Collector Execution Profiler Token Cost Meter Profile Report Performance Baseline Optimization Recommender Parallel Tool Dispatcher Tool Result Cache Evaluation Harness Evaluation Dataset Semantic Similarity Scorer Regression Threshold Policy Evaluation Result Store Performance Trend Analyzer Continuous Integration Runner Regression Gate Alert Manager Trace Exporter

Sources: Ch7.3 Ref1.01 Ref3.07

Progressive deepening planning

When to choose. Choose when fast coarse plans are needed up front and world state becomes more certain during execution (e.g., route -> path -> motion planning in autonomous vehicles).

HTN Planner Task Network State Abstraction Mapper Plan Executor Replanner abstract World Model State

Sources: Ch5.4

Progressive offline-to-online evaluation pipeline

When to choose. Choose when production traffic is sufficient for A/B significance and changes are major (model swap, workflow restructuring): filter candidates offline, validate in staging, then A/B test before full rollout.

Continuous Integration Runner Evaluation Harness Evaluation Dataset Response Scorer abstract Experiment Tracker Evaluation Result Store Evaluation Baseline Statistical Comparator Regression Gate Regression Threshold Policy Staging Environment A/B Test Traffic Splitter A/B Test Configuration Experiment Guardrail Monitor Rollout Manager abstract Release Approver

Alternative to: Offline-only continuous evaluation

Sources: Ch3.1A Ch3.1B

Prompt-chained tool workflow

When to choose. Choose when tool-chain steps are predictable and validation gates are required between steps.

Prompt Chain Orchestrator Sequential Tool Dispatcher Tool Executor Tool Result Transformer Progress Validator Retry Handler Working Memory Buffer abstract Tool Schema Native Function-Calling API abstract

Alternative to: Agent-driven tool chaining

Sources: Ch2.6

Prompt-engineered (zero-shot) agent

When to choose. Choose first: for well-understood tasks the model handles zero-shot and for problems stemming from unclear instructions, since prompt engineering costs minutes and needs no training infrastructure.

System Prompt Template Prompt Optimizer Evaluation Harness Evaluation Dataset Online Evaluator Inference Serving Configuration abstract LLM Inference Service Foundation LLM abstract Trace Collector

Alternative to: Few-shot in-context adaptation, Retrieval-grounded adaptation, Fine-tuned specialist agent

Sources: Ch3.5

Proportionate small/medium-organization AI governance

When to choose. Choose when a small or medium organization applies the same framework principles with simpler structures and fewer roles, quick-start focus on high-priority risks, proportional documentation and focused controls.

Accountable AI Executive AI Governance Policy AI Use Inventory AI Risk Tier Classifier Harm Risk Register Risk Scorer abstract Risk Treatment Plan Model Card Dataset Datasheet Audit Log Store Incident Manager

Alternative to: Integrated NIST AI RMF + ISO/IEC 42001 governance program

Sources: Ch9.8

Prototype RAG with embedded vector store

When to choose. Choose for first RAG prototypes, notebooks and local development with fewer than ~100K vectors, where simplicity and speed-to-market matter most; plan migration to a production-grade store.

Embedded Vector Index Store Hosted Embedding API Service Cost-Optimized Embedding Model Vector Retriever abstract

Alternative to: Single-node production vector store, Highly available replicated vector store cluster, Distributed billion-scale vector search

Sources: Ch6.1 Ch6.2A

Pull-based metrics, dashboards, alerting and structured logging

When to choose. Choose to validate optimization effectiveness and SLA compliance for production inference deployments on a container orchestrator.

Inference Metrics Endpoint Metrics Scrape Configuration Metrics Collector Time-Series Metrics Store Metrics Dashboard abstract Alert Rule Set Alert Manager Platform Operator Log Aggregator Centralized Log Store Token Cost Meter

Sources: Ch7.2 Ref7.16

Pure learning-based decision agent

When to choose. Choose when training data is ample, reward signals are clearly defined, continuous adaptation is tolerated, and no explainability or verification mandate applies (e.g., self-play game agents).

Learned-Policy Decision Engine Policy Network Reinforcement Learning Policy Learner Environment Simulator Experience Replay Buffer Reward Function Specification Opponent Policy League

Alternative to: Three-layer hybrid decision agent (strategic utility / tactical rules / operational learning)

Sources: Ch5.12 Ch5.13

Quality-critical LLM serving

When to choose. Choose when response accuracy is the priority and higher precision, larger models and extra verification cost are acceptable.

Large Language Model Tier FP16 Inference Engine Model Ensemble Orchestrator Output Verifier LLM Inference Service

Sources: Ref7.15

Quality-optimized agent configuration

When to choose. Choose when accuracy is paramount and mistakes carry serious consequences (e.g., medical diagnosis, complex analysis) and latency/cost budgets allow ~6s P95 and ~$0.12/query.

Large Language Model Tier Inference Serving Configuration abstract Full Conversation Buffer Iteration Limit Policy ReAct Agent Controller

Alternative to: Balanced-performance agent configuration, Latency-optimized agent configuration, Cost-optimized agent configuration

Sources: Ch3.4

ReAct agent with rich state

When to choose. Choose when the task prioritises adaptation because solution paths are unpredictable and need dynamic adjustment.

Agent Controller abstract Reasoning Engine Tool Executor Working Memory Buffer abstract LLM Inference Service Prompt Context Builder Iteration Limit Policy Progress Validator State Cycle Detector Summarizing History Compressor

Alternative to: Plan-and-Execute with lean state

Sources: Ch1.5A Ch1.6

Read-only web research agent

When to choose. Choose for initial production deployment or READ tasks (product research, comparison, availability checking) that require no state modification, limiting risk while building confidence.

Web Navigation Agent Task Planner abstract Replanner abstract Browser Navigator Page Content Extractor Page State Observer External Website External Service API Retry Handler Working Memory Buffer abstract Conversation State Store abstract Trace Collector Trace Store

Alternative to: Transactional web agent

Sources: Ch3.3

Real-time voice customer-service agent

When to choose. Choose for live phone or voice-assistant interactions needing sub-second end-to-end turn latency.

Sensor Input Adapter Voice Turn Coordinator Streaming Speech Recognizer Speech Recognition Model ASR Word Boost List Agent Controller abstract LLM Inference Service Speech Synthesizer Speech Synthesis Model Voice Persona Profile Response Streamer Inference Server Queue-Depth Autoscaler

Sources: Ch7.5

Registry-gated GitOps progressive delivery

When to choose. Choose for enterprise agent releases needing reproducible artifacts, gated approvals, declarative deployment and automated canary rollback.

Continuous Integration Runner Experiment Tracker Model and Agent Release Registry Versioned Agent Release Stage Promotion Controller Regression Gate Release Approver Configuration Repository Deployment Manifest Environment Overlay GitOps Reconciler Rollout Manager abstract Rollout Analysis Template Experiment Guardrail Monitor Metrics Collector Layer-7 Load Balancer Container Orchestrator

Sources: Ch4.4

Regulated conversational claims processing with mandatory human checkpoints

When to choose. Choose when automated decisions in regulated domains (insurance, finance, healthcare) require mandatory human checkpoints for cases above complexity or value thresholds plus complete decision audit trails.

Conversational (Chat) Interface User Identity Verifier Workflow Orchestrator abstract Compliance Policy Rule Set Risk Gate Human Approver Audit Log Store Decision Explainer abstract Explanation Presenter Compliance Officer

Sources: Ch10.1

Reserved baseline + spot burst + on-demand buffer

When to choose. Choose to minimize cost when load has a predictable baseline and transient peaks, accepting spot interruptions covered by an on-demand buffer.

Reserved GPU Node Spot GPU Node On-Demand GPU Node Spot Interruption Handler Instance Group Autoscaler Load Balancer abstract

Sources: Ch4.7

Retrieval-grounded adaptation

When to choose. Choose when the agent lacks knowledge (recent events, proprietary information, specialized facts) that no prompt or demonstration can supply; leaves model weights unchanged.

System Prompt Template Retriever abstract Context Assembler abstract Vector Index Store abstract Factuality Verifier abstract LLM Inference Service Foundation LLM abstract

Alternative to: Fine-tuned specialist agent

Sources: Ch3.5

Risk-tiered human-in-the-loop oversight

When to choose. Choose when agents take actions with real-world consequences whose risk magnitude, confidence, and reversibility vary, so notification, approval, and monitoring must be applied selectively.

Oversight Gate abstract Confidence Gate Risk Gate Escalation Threshold Policy Adaptive Threshold Tuner Proactive Notifier Action Rollback Service Approval Gateway Approval Review Console Execution Monitor Console Human Approver Human Supervisor Audit Log Store

Sources: Ch1.1A Ch1.1B

Role-based hierarchical crew

When to choose. Choose for multi-agent workflows with clear role specialization and well-defined task dependencies (e.g., research, write, edit content pipeline) executed sequentially or via manager delegation.

Role-Based Task Orchestrator Supervisor Agent Worker Agent abstract System Prompt Template

Alternative to: Plugin-routed enterprise agent platform

Sources: Ch2.1

Role-based hierarchical team with manager quality gates

When to choose. Choose when work maps to an organizational team with clear role specialization and needs manager-enforced validation gates and iterative refinement (content production, software pipelines, data analysis, compliance review).

Supervisor Agent Worker Agent abstract Task Definition System Prompt Template Critique Rubric Agent Delegation Interface abstract LLM Inference Service

Alternative to: Conversation-driven multi-agent collaboration, Role-based sequential team

Sources: Ch2.4

Role-based sequential team

When to choose. Choose when work maps to a team whose tasks have fixed, obvious progression and discrete deliverables at each stage.

Role-Based Task Orchestrator Worker Agent abstract Task Definition System Prompt Template LLM Inference Service

Alternative to: Conversation-driven multi-agent collaboration, Role-based hierarchical team with manager quality gates

Sources: Ch2.4

Rollout-based MCTS planning

When to choose. Choose when a fast simulator exists but no trained networks or training data are available, heuristics are hard to design, and near-optimal decisions within a real-time budget suffice.

MCTS Planner MCTS Search Tree Store Search Tree Pruner Rollout Simulator Environment Simulator Reward Function Specification MCTS Search Configuration Search Budget Policy Plan Executor

Alternative to: Neural-guided MCTS (AlphaGo/AlphaZero style)

Sources: Ch5.5

Rule-constrained utility optimisation (hybrid)

When to choose. Choose when hard safety, regulatory or ethical constraints coexist with continuous multi-objective trade-offs (autonomous driving lane changes, trading within risk limits, treatment recommendation).

Rule Constraint Filter Formal Rule Specification Decision Engine abstract Logic Conclusion Verbalizer Explanation Presenter

Sources: Ch5.11

Safety-separated reliability monitoring

When to choose. Choose when a guardrailed agent produces both safety blocks and infrastructure failures, so each category needs its own SLO, alert routing and dashboard panels to avoid conflated alerts and misattributed error budgets.

Guardrail Orchestrator Guardrail Policy Input Rail Output Rail Fact Checking Rail abstract Failure Category Classifier Metrics Collector Time-Series Metrics Store Service Level Objective Specification Alert Rule Set Alert Manager Metrics Dashboard abstract Security Analyst Platform Operator

Sources: Ch8.2B

Self-hosted accelerated function calling

When to choose. Choose for tool-heavy agents with data-residency, security, or cost constraints, or high throughput (thousands of requests per hour).

Self-Hosted Inference Endpoint LLM Inference Service Foundation LLM abstract Optimized Inference Engine abstract Engine Builder KV Cache Manager Tool Schema GPU Node abstract Container Orchestrator

Sources: Ch2.6

Self-hosted GPU embedding for regulated, long-document workloads

When to choose. Choose when data sovereignty or air-gapped operation is required, documents are long, and volume exceeds a few thousand queries per day.

Self-Hosted GPU Embedding Service Long-Context Embedding Model Inference Server Dynamic Batch Scheduler Engine Builder OpenAI-Compatible Inference API GPU Node abstract

Sources: Ch6.1

Separate modality stores with cross-modal reranking

When to choose. Choose for research environments experimenting with per-modality embedding models, or production systems with mature MLOps that accept the highest operational cost for best-in-class per-modality results.

Document Ingestor Modality-Specific Embedding Service Modality-Specific Vector Store Per-Modality Fan-Out Retriever Cross-Modal Reranker Cross-Modal Reranking Model Fine-Tuning Pipeline abstract Multimodal Context Assembler Multimodal Answer Synthesizer Inference Server

Alternative to: Unified embedding space multimodal RAG, Ground-to-text multimodal RAG with metadata

Sources: Ch2.7

Sequential neural-to-symbolic pipeline

When to choose. Choose when the task decomposes cleanly into perception then rule application and early-stage processing can extract all relevant information without later feedback; most common in production.

Perception Interpreter Neural Perception Model Neural-to-Symbolic Translator Paradigm Boundary Validator Rule-Based Decision Engine Formal Rule Specification Knowledge Graph Store abstract

Alternative to: Parallel paradigms with result fusion, Cooperative iterative neural-symbolic refinement, Embedded neural modules within a symbolic program

Sources: Ch5.13

Sequential tool-calling single agent

When to choose. Choose for single-agent workflows that follow a linear sequence (receive query, reason about tools, execute tools, synthesize answer) with state limited to conversation history, e.g., FAQ/knowledge-base chatbots and simple QA. Also fits rapid prototyping, conversation-based interfaces needing only message history, simple tool integration and unmodified ReAct reasoning (e.g., research assistants, database Q&A, calendar/email assistants).

ReAct Agent Controller Tool Executor Tool Schema LLM Inference Service Conversation State Store abstract Working Memory Buffer abstract Vector Retriever abstract Embedding Service abstract Agent Action Output Parser Full Conversation Buffer LLM Provider Adapter System Prompt Template Tool Integration Adapter abstract Trace Collector

Alternative to: Graph-based iterative workflow agent, Conversation-driven multi-agent collaboration, Role-based hierarchical crew, Plugin-routed enterprise agent platform

Sources: Ch2.1 Ch2.2 Ch2.3

Serverless event-driven agents

When to choose. Choose for highly variable or bursty traffic with idle periods, discrete event workloads, rapid iteration (many deployments per day) and small teams without infrastructure expertise, when cold starts fit within the latency SLO.

Serverless Function Runtime On-Demand Consumption Plan Event-Triggered Agent Agent API Gateway Content-Based Event Router Publish-Subscribe Bus Message Queue Dead Letter Queue Event Stream Log Idempotency Store Schema Registry Object Store State Checkpoint Store abstract Stalled Workflow Resumer Durable State Machine Orchestrator Conversation State Store abstract Result Callback Webhook LLM Inference Service Trace Collector

Alternative to: Containerised microservices agent deployment

Sources: Ch4.2

Session-affinity replicas

When to choose. Choose when latency needs are extreme (sub-100 ms), scale is tens of instances, and session recreation on failure is acceptable.

Layer-4 Load Balancer Agent Controller abstract Instance-Local Session State In-Process Response Cache LLM Inference Service

Alternative to: Stateless replica horizontal scaling

Sources: Ch1.8

Shared multi-agent episodic memory

When to choose. Choose when multiple agents operate in a common environment and should learn vicariously from teammates' experience, with episodes updated as conditions change.

Episodic Memory Store Vector Index Store abstract Metadata-Filtered Retriever Event-Based Episode Encoder Memory Retriever

Sources: Ch5.7

Shared-state multi-agent workflow

When to choose. Choose when specialised agents with sequential dependencies must coordinate; shared state gives single-structure observability, simpler recovery and easy onboarding of new agents compared with message passing or publish-subscribe.

State-Graph Orchestrator Workflow State Graph Agent State Schema Working Memory Buffer abstract Rule-Based Transition Router Worker Agent abstract

Sources: Ch1.5A Ch1.6

Single large instance (vertical scaling)

When to choose. Choose for computation or memory bottlenecks with predictable workloads and modest scale (10-50 req/s, bursts to ~100) in early production where failure tolerance is less critical.

Agent Controller abstract LLM Inference Service In-Process Response Cache Instance-Local Session State On-Demand GPU Node

Alternative to: Stateless replica horizontal scaling

Sources: Ch1.8

Single tool integration

When to choose. Choose when one primary capability extends the LLM (current data lookup, calculations, simple API calls) and each query triggers at most one independent tool call, e.g., FAQ or order-status agents.

Direct Tool-Calling Controller Reasoning Engine Native Function-Calling API abstract LLM Inference Service Tool Schema Tool Registry Tool Executor REST API Adapter External Service API Trace Collector

Sources: Ch2.6

Single-agent agentic RAG with self-hosted Nemotron

When to choose. Choose for knowledge-base assistants where the agent should decide per query whether to retrieve, using hybrid retrieval plus reranking and a self-hosted OpenAI-compatible endpoint.

ReAct Agent Controller System Prompt Template Tool Schema Iteration Limit Policy File Store Extractor Overlapping Window Chunker Text Embedding Service Retrieval-Optimized Embedding Model Embedded Vector Index Store Dense-Sparse Hybrid Retriever Vector Retriever abstract Keyword Retriever Reranker abstract OpenAI-Compatible Inference API Inference Server Small Language Model Tier In-Process Response Cache Trace Collector

Sources: Ref7.07 Ref7.13

Single-agent ReAct tool-using agent

When to choose. Choose when solution paths are unpredictable and each step depends on prior observations (research, debugging, non-standard support), and latency and cost overhead are acceptable.

ReAct Agent Controller Reasoning Engine Prompt Exemplar Set LLM Inference Service Foundation LLM abstract Working Memory Buffer abstract Context Window Manager abstract Tool Registry Tool Schema Tool Executor Retry Handler Circuit Breaker

Alternative to: Plan-and-Execute structured workflow agent

Sources: Ch1.2

Single-node production vector store

When to choose. Choose for production RAG with millions of documents where authentication, durability and monitoring are required but brief outages on node failure are tolerable (below 99.9% uptime needs).

Self-Managed Vector Index Store Persistent Volume Deployment Manifest API Key Authenticator Vector Store gRPC API Knowledge Chunk Metadata Schema Vector Index Build Configuration abstract Vector Search Configuration Vector Batch Ingestor Ingestion Pipeline Orchestrator Dense-Sparse Hybrid Retriever Fusion Weight Selector Metrics Collector Time-Series Metrics Store Metrics Dashboard abstract Alert Manager Alert Rule Set Vector Store Capacity Monitor

Alternative to: Highly available replicated vector store cluster

Sources: Ch6.2B

Single-region, multi-availability-zone deployment

When to choose. Choose by default for high availability (99.9%+) against individual zone failures without multi-region cost and complexity.

Load Balancer abstract Agent Controller abstract External Session State Store Vector Index Store abstract Container Orchestrator

Alternative to: Per-region full agent stacks with nearest-region routing

Sources: Ch4.7

Single-strategy reactive replanning agent

When to choose. Choose when failures are rare (<10% of executions) and unpredictable, time budgets tolerate multi-second replanning pauses, memory is constrained, or state spaces are small.

Plan-and-Execute Controller Optimal Heuristic Search Planner Geometric Distance Heuristic State-Space Graph Plan Executor Plan Deviation Monitor Discrepancy Significance Evaluator Complete Replanner

Sources: Ch5.6

SSE streaming RAG agent

When to choose. Choose for query-then-read agents (e.g., customer support RAG) where no mid-stream user interaction is needed and sub-second perceived responsiveness is required.

Conversational (Chat) Interface Server-Sent Events Stream Response Streamer Stream Connection Manager Agent Controller abstract Dense-Sparse Hybrid Retriever Vector Retriever abstract Keyword Retriever Vector Index Store abstract In-Process Response Cache Semantic Cache Context Assembler abstract Context Compressor Answer Synthesizer abstract Native Function-Calling API abstract LLM Inference Service Error Presenter Retry Handler Metrics Collector Trace Collector Layer-7 Load Balancer

Alternative to: Bidirectional interactive streaming agent

Sources: Ch2.9

Stage-gated data-quality pipeline for high-stakes RAG

When to choose. Choose when a RAG agent serves high-stakes or regulated domains (healthcare, finance, legal) where 97-99% 'good enough' quality, duplicate contradictions or embedded PII are unacceptable.

Knowledge Source System Source Change Detector Ingestion Pipeline Orchestrator Document Ingestor Data Quality Validator Source Document Schema Data Quality Rule Set Business Rule Validator Formal Rule Specification Document Structure Validator Data Quality Gate Data Quality SLA Specification Data Format Normalizer Canonical Format Specification Cascading Deduplicator Exact Hash Deduplicator Fuzzy Text Deduplicator Semantic Deduplicator Referential Integrity Validator Document Version Reconciler Multi-Strategy PII Detector Document PII Redactor Document Chunker abstract Text Embedding Service Vector Index Store abstract Index Integrity Validator Smoke Tester Content Freshness Monitor Content Freshness Policy Corpus Drift Detector Data Validation Result Store Data Quality Review Queue Data Quality Reviewer Metrics Dashboard abstract Alert Manager SLO Monitor Evaluation Dataset RAG Evaluator Audit Log Store

Sources: Ch6.4

Standard single-query RAG

When to choose. Choose for simple, single-need queries where decomposition adds latency and token cost without improving quality.

RAG Query Orchestrator Text Embedding Service Retriever abstract Context Assembler abstract Answer Synthesizer abstract LLM Inference Service

Alternative to: Decomposed RAG (query decomposition + parallel retrieval + synthesis)

Sources: Ch6.6

Stateless elastic agent fleet

When to choose. Choose for interactive agents with variable request durations and diurnal traffic, where every replica can serve any request because all state is externalized.

Agent Controller abstract Least-Connections Load Balancer Resource-Utilization Autoscaler Autoscaling Policy Container Orchestrator Deployment Manifest Container Health Prober Liveness Endpoint Readiness Endpoint External Session State Store Object Store Log Aggregator Metrics Collector Metrics Dashboard abstract

Alternative to: Sticky-session migration stage

Sources: Ch4.7

Stateless replica horizontal scaling

When to choose. Choose for most (~90%) production agent deployments: discrete or conversational requests with externalized state, unpredictable spikes, high availability (99.99%) or multi-tenant SaaS scale.

Layer-7 Load Balancer Layer-4 Load Balancer Agent API Gateway Agent Controller abstract LLM Inference Service Small Language Model Tier External Session State Store Distributed Response Cache KV Cache Manager Vector Index Store abstract Task Queue Metric-Driven Autoscaler Autoscaling Policy Deployment Manifest Container Health Prober Liveness Endpoint Readiness Endpoint Metrics Collector Alert Manager On-Demand GPU Node Container Orchestrator

Alternative to: Single large instance (vertical scaling), Session-affinity replicas

Sources: Ch1.8

Sticky-session migration stage

When to choose. Choose only temporarily while a partially stateful agent still caches conversation context locally, before state is externalized and routing switches to least connections.

Agent Controller abstract Session Affinity Load Balancer Instance-Local Session State External Session State Store Container Orchestrator Container Health Prober

Alternative to: Stateless elastic agent fleet

Sources: Ch4.7

Structured Chain-of-Thought reasoning with layered verification

When to choose. Choose for open-ended, contextual or judgment-heavy tasks where flexible pattern-based reasoning is more appropriate than formalisation.

Reasoning Engine Structured Reasoning Prompt Template abstract Stepwise Reasoning Verifier Citation Verifier Trace Collector

Sources: Ch3.9

Supervised physical autonomy (manufacturing / robotics)

When to choose. Choose for autonomous robots or manufacturing systems operating near people, where physical safety requires continuous human supervisory awareness, veto authority and immediate shutdown capability.

Agent Controller abstract Physical Safety Envelope Execution Monitor Console Human Supervisor Emergency Stop Controller Action Policy Engine

Sources: Ch10.5

Swarm intelligence multi-agent system

When to choose. Choose for distributed optimisation or coverage with simple agents where central control would be a bottleneck or single point of failure and 'good enough' solutions suffice.

Swarm Agent Swarm Anomaly Detector

Alternative to: Competitive (game-theoretic) multi-agent system

Sources: Ch1.3

Three-layer hybrid decision agent (strategic utility / tactical rules / operational learning)

When to choose. Choose for safety-critical open-world agents (e.g., autonomous vehicles) that must optimize goals, obey hard legal/safety rules, and adapt control behaviour at once.

Hybrid Decision Arbiter Decision Precedence Policy Utility-Based Decision Maker Utility Function Specification Rule-Based Decision Engine Rule Constraint Filter Formal Rule Specification Learned-Policy Decision Engine Policy Network Replanner abstract Execution Plan abstract Multimodal Fusion Engine World Model State Audit Log Store

Alternative to: Pure learning-based decision agent

Sources: Ch5.12 Ch5.13

Throughput-bound speculative serving

When to choose. Choose for batch or offline generation with predictable outputs (summarization, structured data) where GPU-hours matter more than time-to-first-token.

Inference Server Large Language Model Tier Speculative Draft Model abstract Speculative Decoder Speculative Decoding Configuration Throughput-Oriented Batching Config Token Predictability Analyzer Evaluation Harness

Sources: Ch4.4

Throughput-critical HITL approval (e.g., batch claims, bulk onboarding)

When to choose. Choose when processing volume matters more than immediate response: batch similar requests, consolidated review with bulk actions, predictive pre-approval, and risk-stratified autonomous handling of routine cases.

Oversight Gate abstract Risk Gate Pre-Approved Category Policy Approval Request Batcher Approval Request Store Approval Router Approval Review Console Oversight Performance Monitor Human Approver Human Specialist

Alternative to: Latency-critical HITL approval (e.g., real-time fraud review)

Sources: Ch10.4

Throughput-critical LLM serving

When to choose. Choose when maximum tokens per second is the priority (batch document analysis, summarization, embedding generation).

LLM Inference Service In-Flight Batch Scheduler Dynamic Batch Scheduler Throughput-Oriented Batching Config Throughput-Optimized Deployment Profile Paged KV Cache Allocator

Alternative to: Cost-sensitive LLM serving, Quality-critical LLM serving

Sources: Ref7.02 Ref7.04 Ref7.15

Throughput-critical quantized LLM serving

When to choose. Choose for high-volume services (chatbots for millions of users, real-time translation, content moderation) that accept ~1-2% accuracy loss for 4-8x capacity.

Model Graph Exporter Model Quantizer Quantization Calibrator Quantization Calibration Dataset Engine Builder Engine Build Configuration FP8 Quantized Engine INT8 Quantized Engine Paged KV Cache Allocator In-Flight Batch Scheduler Fused Block-wise Attention Kernel Inference Server Evaluation Harness

Alternative to: Accuracy-critical FP16-baseline serving, Latency-critical low-batch serving

Sources: Ch7.4 Ref7.05

Throughput-optimized batch inference serving

When to choose. Choose for batch workloads (document summarization, moderation queues, bulk translation) where completion within a processing window matters more than per-request latency.

Concurrent Inference Request Dispatcher Connection Pool OpenAI-Compatible Inference API Inference Server Inference Batch Scheduler abstract Throughput-Oriented Batching Config GPU Node abstract

Alternative to: Latency-optimized interactive inference serving, Cost-optimized inference serving

Sources: Ch7.2

Tiered-autonomy conversational customer service assistant

When to choose. Choose for high-volume customer service where most requests have clear intent and straightforward resolution paths, autonomously resolving routine inquiries (Klarna 65%, Erica 98%) while escalating complex, sensitive or low-confidence cases to humans with full context.

Conversational (Chat) Interface Dialogue Flow Manager Intent Router Intent Taxonomy Clarification Manager Hierarchical History Compressor Working Memory Buffer abstract Session Summary Store Semantic Memory Store Escalation Agent Escalation Handler Escalation Handoff Package Human Specialist Feedback Collector abstract Behavioral Signal Tracker Active Learning Sampler Training Pipeline Orchestrator Metrics Dashboard abstract Alert Manager

Sources: Ch10.1

Time-indexed audio RAG

When to choose. Choose when critical knowledge exists only in recorded conversations (earnings calls, support calls, training sessions, design discussions) and answers must cite exact playback timestamps; skip when adequate transcripts or notes already exist.

Speech Transcriber Speech Recognition Model Time-Indexed Transcript Chunker Multimodal Chunk Metadata Schema Text Embedding Service Vector Index Store abstract Metadata-Filtered Retriever Source Media Store Timestamped Citation Presenter Answer Synthesizer abstract

Sources: Ch2.7

Time-sliced shared GPU

When to choose. Choose for internal development, trusted co-located or batch workloads, or GPUs without partitioning support, where latency variance is acceptable.

Time-Sliced GPU Share Inference Server

Alternative to: Multi-tenant SaaS on uniform GPU partitions

Sources: Ch7.6

Top-down (rule/principle-based) value alignment

When to choose. Choose when explicit verifiability and auditability matter most and scenarios are clear and unambiguous; accept brittleness in novel situations outside the specified rules.

Constitution abstract Principle Priority Policy Operational Norm Set Formal Rule Specification Rule Constraint Filter Guardrail Policy Self-Reflection Critic Principle Adherence Evaluator Principle Adherence Test Suite

Alternative to: Bottom-up (learned) value alignment, Hybrid value alignment (hard constraints + learned nuance + continuous monitoring)

Sources: Ch9.6

ToT constraint satisfaction (DFS + sequential proposal + value checks)

When to choose. Choose for scheduling-style problems with interdependent hard and soft constraints where violations surface only at deeper levels and backtracking is essential.

Depth-First Thought Search Controller Sequential Proposal Thought Generator Value Thought Evaluator Problem Constraint Specification Thought Search Policy LLM Inference Service Human Evaluator External Service API

Alternative to: ToT creative generation (BFS + independent sampling + voting)

Sources: Ch5.2

ToT creative generation (BFS + independent sampling + voting)

When to choose. Choose for creative content whose quality depends on high-level structure and resists absolute scoring, with human review before publication.

Breadth-First Thought Search Controller Independent Sampling Thought Generator Vote Thought Evaluator Thought Decomposition Specification LLM Inference Service Human Evaluator External Service API

Alternative to: ToT constraint satisfaction (DFS + sequential proposal + value checks)

Sources: Ch5.2

Train, optimize, containerize, deploy pipeline

When to choose. Choose when taking a fine-tuned model to production serving through engine compilation and containerized inference microservices with guardrails and continuous monitoring.

Fine-Tuning Pipeline abstract Foundation LLM abstract Engine Builder Optimized Inference Engine abstract Container Image Builder Container Image LLM Inference Service Input Rail Output Rail Inference Server Dynamic Batch Scheduler Container Orchestrator Metrics Collector Time-Series Metrics Store Metrics Dashboard abstract

Sources: Ref7.14 Ref7.18

Transactional web agent

When to choose. Choose after read-only reliability is established, for WRITE tasks (login, cart modification, checkout, access requests) under explicit domain policies and approval hierarchies.

Web Navigation Agent Task Planner abstract Replanner abstract Browser Navigator Page Content Extractor Web Transaction Executor Page State Observer External Website Retry Handler Secrets Vault Action Policy Engine Compliance Policy Rule Set Approval Gateway Human Approver Human Supervisor Audit Log Store Working Memory Buffer abstract Conversation State Store abstract Trace Collector Trace Store

Alternative to: Read-only web research agent

Sources: Ch3.3

Transparent rule-based decision system

When to choose. Choose when every decision must trace to explicit, auditable logic (credit, medical devices, government benefits), expert knowledge is stable and articulable, behaviour must be deterministic and testable, or data is too scarce to learn.

Rule-Based Decision Engine Symbolic Logic Engine Rule Interpreter Production Rule Base Conflict Resolution Policy abstract Working Memory Fact Store Fact Contradiction Resolver Memory Lifecycle Manager Rule Firing Trace Logic Conclusion Verbalizer Counterfactual Explainer Explanation Presenter Audit Log Store Human Approver Domain Rule Expert Compliance Officer

Alternative to: Rule-constrained utility optimisation (hybrid)

Sources: Ch5.11

Tree-of-Thought deliberate reasoning module

When to choose. Choose when problems require exploratory search with backtracking, intermediate evaluation is feasible and informative, and the accuracy gap over CoT justifies 3-5x cost.

ReAct Agent Controller Tree Search Controller abstract Thought Generator abstract Thought State Evaluator abstract Thought Tree Store Thought Decomposition Specification Thought Search Policy Thought Evaluation Prompt Template LLM Inference Service Reasoning Chain Cache Token Budget Enforcer Model Router Working Memory Buffer abstract Episodic Memory Store Semantic Memory Store Procedural Memory Store

Alternative to: Graph-of-Thought synthesis reasoning module

Sources: Ch5.2

Unified embedding space multimodal RAG

When to choose. Choose for rapid deployment on an existing text RAG stack with mostly general imagery (product photos, simple diagrams), minimal budget, and prototyping timelines of days.

Document Ingestor Joint Multimodal Embedding Service Contrastive Image-Text Encoder Vector Index Store abstract Unified Embedding Retriever Context Assembler abstract Answer Synthesizer abstract Inference Server OpenAI-Compatible Inference API

Alternative to: Ground-to-text multimodal RAG with metadata, Separate modality stores with cross-modal reranking

Sources: Ch2.7

Unified multi-framework inference serving

When to choose. Choose when an agent system must serve heterogeneous models from multiple frameworks (LLMs, vision, classifiers, tree-based) on shared GPU infrastructure with one API and common monitoring and scaling.

Inference Server Tensor Inference API OpenAI-Compatible Inference API Inference Backend abstract Tensor Framework Backend LLM Generation Backend Dynamic Batch Scheduler In-Flight Batch Scheduler Inference Queue Policy Inference Serving Configuration abstract Model Repository Dynamic Model Loader Model Ensemble Orchestrator Model Ensemble Definition GPU Partition Inference Metrics Endpoint Metrics Collector Custom Metrics Adapter Metric-Driven Autoscaler Replica Placement Policy Metrics Dashboard abstract Alert Manager Trace Collector

Sources: Ch4.5

Validated transactional tool chain

When to choose. Choose when agents invoke chains of side-effecting tools (inventory reservation, payment, shipping) where out-of-order, duplicated or hallucinated calls cause concrete harm requiring manual remediation.

Tool Executor Tool Schema Tool Registry Tool Call Schema Validator Parameter Provenance Validator Parameter Context Grounding Validator Tool Precondition Checker Action Sequence Policy Tool Idempotency Guard Tool Response Plausibility Checker Tool Error Classifier Retry Handler Approval Gateway Trace Collector Trace Store Audit Log Store Tool Usage Analyzer

Sources: Ch3.7

Value-aligned multi-agent coordination

When to choose. Choose when several specialised agents optimize local objectives whose externalities can violate shared organizational values (e.g., supply chain procurement, logistics, finance, customer service).

Constitution abstract Operational Norm Set Worker Agent abstract Value Sentinel Agent Agent Message Bus abstract Preference Annotator abstract Preference Dataset Reward Model Decision Stakeholder

Sources: Ch9.6

Vector RAG only

When to choose. Choose for simple factual Q&A and semantic search (customer support, documentation search) where no relationship traversal is needed; lowest complexity (p95 ~120ms).

Document Chunker abstract Embedding Service abstract Vector Index Store abstract Vector Retriever abstract Context Assembler abstract Answer Synthesizer abstract LLM Inference Service

Alternative to: Knowledge graph only (text-to-graph-query QA), Hybrid RAG + knowledge graph

Sources: Ch1.7A Ch1.7B

Vector-only memory

When to choose. Choose when semantic similarity is a good proxy for relevance over large heterogeneous corpora and speed and low engineering overhead matter more than logical precision (e.g., customer support over documentation); a common starting point before migrating to hybrid.

Semantic Memory Store Vector Index Store abstract Text Embedding Service Text Embedding Model abstract Semantic Boundary Chunker Vector Retriever abstract Retrieval Confidence Filter Answer Synthesizer abstract LLM Inference Service

Alternative to: Knowledge-graph memory, Hybrid vector + graph memory

Sources: Ch5.7 Ch5.8

Voice + vision multimodal agent (accessibility, visual troubleshooting)

When to choose. Choose when users combine a spoken query with a camera image or photo (scene description for blind users, product troubleshooting).

Sensor Input Adapter Multimodal Content Router Streaming Speech Recognizer Image Captioner Vision-Language Model abstract Agent Controller abstract Retriever abstract Tool Executor LLM Inference Service Speech Synthesizer Voice Persona Profile

Sources: Ch7.5

Weighted least-connections fleet on mixed GPU/CPU nodes

When to choose. Choose for highly heterogeneous independent workloads (e.g., 2-60 s document extraction) served by GPU and CPU replicas of measurably different capacity.

Agent Controller abstract Least-Connections Load Balancer Weighted Round-Robin Load Balancer Load Balancing Policy Resource-Utilization Autoscaler GPU Node abstract CPU Compute Node Container Orchestrator Metrics Collector

Sources: Ch4.7