Variation points
149 abstract roles and their interchangeable variants, with the book's selection criteria. Alternatives must share a common ancestor (enforced by SHACL).
Agent Controller abstract
An abstract agent runtime that runs the control loop deciding, for one agent, how reasoning, tool actions, and observations are sequenced toward a goal.
- Direct Tool-Calling ControllerChoose for simple, deterministic tasks where a straightforward tool call suffices and latency or budget constraints cannot absorb multi-step reasoning or planning overhead.
- Function-Calling ControllerChoose when an agent must integrate an extensive, growing catalog of capabilities (tens of functions across many domains) whose composition cannot be pre-programmed and new capabilities must be added without changing orchestration code. Avoid for focused agents with fewer than five capabilities, or for workflows needing iterative refinement loops, conditional branching on intermediate results, or strictly deterministic execution (prefer explicit state-graph orchestration).
- Plan-and-Execute ControllerChoose for complex workflows with mostly predictable steps in stable environments, resource-constrained settings, or supervisor-worker coordination; avoid for highly dynamic, real-time, simple, or open-ended exploratory tasks.
- ReAct Agent ControllerChoose when solution paths are unpredictable and require adaptive investigation (research, debugging, non-standard support); avoid for simple deterministic, latency-critical, budget-constrained, or formally verified tasks.
Agent Delegation Interface abstract
An inter-agent interface through which one agent hands a task, with structured context, to another agent.
- Discoverable Delegation InterfaceChoose in heterogeneous systems where agent availability fluctuates or several agents offer similar capabilities with different trade-offs, requiring runtime discovery.
- Static Delegation InterfaceChoose when task delegation boundaries are well-defined and agent capabilities are static, so no runtime capability discovery is needed.
Agent Message Bus abstract
A communication medium that transports messages or events between agents in a multi-agent system.
- Event Broker abstractChoose when high throughput is critical, agents run asynchronously at variable rates, resilience to temporary failures is essential, state changes must reach many subscribers, and eventual consistency is acceptable.
- Content-Based Event RouterChoose when events in one stream must reach different specialised agents depending on classification metadata in the payload (e.g., financial vs legal documents).
- Event Stream LogChoose when events must be replayed to rebuild state or train models, analysed temporally, or retained as a permanent audit record.
- Message QueueChoose for FIFO buffering and temporal decoupling where queues must absorb traffic spikes and ordering matters.
- Publish-Subscribe BusChoose when a single event must reach multiple subscribers simultaneously by topic for parallel downstream processing.
- Content-Based Event Router
- Synchronous Message ChannelChoose when immediate feedback or strong consistency is essential, interactions are brief and synchronous, the agent network is small, and explicit communication trails aid compliance.
Deadlock Controller abstract
An abstract coordination component that keeps multi-agent workflows free of circular waits in which each agent blocks on another's output.
- Dependency Cycle ValidatorChoose when agent dependencies can be declared at workflow-configuration time; static validation makes runtime deadlock structurally impossible and fails fast, with no scanning cost or detection latency.
- Runtime Deadlock DetectorChoose only when dependencies cannot be declared in advance; it requires periodic scanning of active agents and lets a deadlock persist until the next scan detects it.
Extraction Cadence Policy abstract
An abstract configuration choosing whether a pipeline extracts source data in periodic batches or continuously as it arrives.
- Batch Extraction PolicyChoose for relatively static knowledge bases where daily updates suffice, and for structured sources with efficient timestamp-based change detection.
- Streaming Extraction PolicyChoose for rapidly changing domains (e.g., news, stock trading) and semi-structured event sources where near-real-time freshness justifies added complexity.
Knowledge Refresh Policy abstract
An abstract configuration choosing whether a pipeline run processes only the delta since the last successful run or reprocesses the full corpus.
- Full Refresh PolicyChoose for initial population, after chunking or schema parameter changes, or to remediate quality issues from earlier permissive filters.
- Incremental Refresh PolicyChoose as the default for scheduled (hourly or daily) runs that keep the knowledge base fresh.
Mixed-Initiative Controller abstract
An abstract controller that determines which party, agent or human, holds initiative at each point of a shared task and how control transfers between them.
- Fixed Subtask Initiative ControllerChoose when workflow stages can be pre-partitioned by clear decision criteria, e.g., routine verification to the agent and judgment-heavy decisions to humans.
- Negotiated Initiative ControllerChoose when no fixed ownership is appropriate and initiative should flow dynamically to whichever party has the relevant capability, capacity and confidence at each stage.
- Subdialogue Initiative ControllerChoose when execution hits a specific ambiguity or misunderstanding that a brief clarifying exchange can resolve before control returns to the primary party.
- Unsolicited Reporting ControllerChoose when one party holds information the other needs but should not control when or how it is acted upon; the recipient keeps full initiative.
State Checkpoint Store abstract
A database to which a checkpointer persists agent/thread state so conversations can resume across sessions, restarts and idle periods.
- Database State StoreChoose for long-running processes (hours or days) whose state must survive instance failure and be resumable by any instance.
- Distributed Cache State StoreChoose for multi-agent or multi-instance systems needing shared, low-latency state for horizontal scaling.
- Graph State StoreChoose when state needs both persistence and relational context (e.g., ticket, related tickets, consulted articles).
- In-Memory State StoreChoose for short-lived, agent-local workflows that complete in seconds to minutes on a single instance; state is lost if the instance fails.
- Serialized File State StoreChoose when persistence is needed but relational queries over state are not.
State Concurrency Controller abstract
A coordination mechanism that keeps shared workflow state consistent when multiple agent instances update it concurrently.
- CRDT State MergerChoose when state must converge regardless of the order of concurrent updates.
- Optimistic Lock ControllerChoose when instances can detect conflicting updates and retry them.
- Pessimistic Lock ControllerChoose when instances must hold exclusive access before modifying state.
Task Allocator abstract
A coordination component that decides which agent receives each task or resource in a multi-agent system.
- Auction Task AllocatorChoose when allocation should adapt dynamically to agents' workload, expertise and confidence, or when agents compete for shared resources.
- Capability-Matching AllocatorChoose when collaborators must be discovered at runtime by advertised capability rather than hardcoded.
- Static Rule Task RouterChoose when task types and agent capabilities are stable and predictable routing matters more than dynamic allocation.
Transition Router abstract
A decision component that evaluates current workflow state against the logic tree to select the next node, tool, or sub-agent to execute.
- Hybrid Transition RouterChoose when LLM reasoning about the next step must be constrained by explicit rules.
- LLM Transition RouterChoose when deciding what to do next requires reasoning, e.g., tasks without predetermined solution paths.
- Rule-Based Transition RouterChoose when decision logic can be expressed as explicit rules over state fields.
Worker Agent abstract
A specialised agent that executes an assigned subtask, often owning a focused tool domain such as database or external-API tools.
Workflow Orchestrator abstract
An execution engine that carries a multi-step agent workflow forward by performing selected operations and writing their results back into explicit workflow state.
- Dependency-Ordered Executor abstractChoose over spawning all agents concurrently whenever agents consume each other's outputs, since probabilistic start ordering violates dependencies under load.
- Asynchronous Dependency-Chaining ExecutorChoose when phase-barrier latency proves unacceptable and maximal parallelism is needed; accepts higher complexity and cascading cancellation of dependents when an agent fails.
- Phase Barrier ExecutorChoose when simple, robust failure semantics matter more than maximal parallelism: a failure in a phase stops the workflow immediately, at the cost of waiting for the slowest agent of each phase.
- Asynchronous Dependency-Chaining Executor
- Durable State Machine OrchestratorChoose when multi-step agent workflows need branching, error handling or durations of hours or days that exceed individual function timeout limits.
- Ingestion Pipeline Orchestrator
- Multi-Agent Coordinator abstract
- Conversational Agent CoordinatorChoose when collaboration benefits from flexible dialogue (agents challenging outputs, requesting clarification, negotiating), for exploratory workflows, prototyping and human-overseen scenarios where readable transcripts aid transparency; avoid where determinism and structured audit trails are required.
- Role-Based Task OrchestratorChoose when the workflow mirrors an organisational structure with clear roles, predictable task dependencies and sequential or hierarchical delegation; avoid for dynamic workflows needing conditional branching or iterative self-correction loops.
- Supervisor AgentChoose when a workflow maps to an organizational team with clear role specialization and needs validation gates and iterative refinement until outputs meet quality standards. Avoid when agents need frequent ad-hoc back-and-forth interaction or when roles overlap or capabilities are vaguely defined.
- Conversational Agent Coordinator
- Parallel Agent CoordinatorChoose when agents can analyse the same source information independently; reduces both tokens and latency.
- Plugin Kernel OrchestratorChoose for enterprise agent ecosystems with dozens of specialized capabilities, evolving service catalogs and existing-system integration needing central orchestration; avoid for focused single-purpose agents where plugin abstraction is over-engineering.
- Prompt Chain OrchestratorChoose when a task decomposes into a fixed sequence of processing stages, each needing a different prompting strategy, with intermediate outputs inspected between stages; for simple linear workflows without branching, simple function composition suffices.
- State-Graph OrchestratorChoose when tasks lack predetermined solutions and need conditional branching (logic trees), crash recovery and traceable transitions; when workflows exhibit decision-tree routing on classification or intermediate results, when the graph is valuable documentation, when observability matters, or when workflows change frequently. Avoid for simple linear workflows or when the learning-curve cost is unjustified for small teams or short-lived projects.
Agent Service API abstract
A network API through which an agent exposes its capabilities as a distributed service to agents across platform, language or organisational boundaries.
- REST Agent APIChoose when human readability matters, clients are diverse (any HTTP client), performance requirements are moderate, and REST infrastructure can be reused.
- gRPC Agent APIChoose when throughput and latency are critical, type safety matters, bidirectional streaming is needed, and both client and server implementations are controlled.
Cache Invalidation Policy abstract
An abstract configuration deciding when cached tool results become stale and must be refreshed, balancing result freshness against latency and cost savings.
- Event-Based Cache Invalidation PolicyChoose when the upstream source can signal publication of updates, so freshness need not wait for a TTL.
- Time-Based Cache Invalidation PolicyChoose when the source updates on a known schedule and no change notifications are available.
Tool Call Dispatcher abstract
An abstract execution component that schedules multiple model-proposed tool calls, either in dependency order or concurrently, and collects their results for the agent.
- Asynchronous Tool DispatcherChoose when profiling shows GPU idle gaps during I/O-bound tool calls in multi-step agent loops (e.g., ReAct) under concurrent load.
- Parallel Tool DispatcherChoose for independent data fetching from multiple sources where latency matters and tools share no mutable state and have no coordination requirements.
- Sequential Tool DispatcherChoose when tools have dependencies (Step N+1 needs Step N's output), when guaranteed execution ordering is needed, when tools modify shared data, or when orchestration logic depends on seeing results sequentially.
Tool Integration Adapter abstract
An abstract connector that binds a tool to an external system through a specific integration mechanism and returns structured results.
- Database ConnectorChoose when the agent needs direct querying or updating of structured data stores.
- Message Queue AdapterChoose for distributed systems needing reliable, buffered delivery that survives temporary component failure.
- REST API AdapterChoose for synchronous CRUD operations and cloud-service integration; the most common mechanism.
- Webhook ReceiverChoose for event-driven, asynchronous workflows where external systems should notify the agent instead of being polled.
Chain-of-Thought Prompt abstract
A prompt artifact that elicits explicit, step-by-step intermediate reasoning from a model before it states its final answer, with or without worked demonstrations.
- Auto-CoT PromptChoose when many distinct problem types need structured reasoning but resources are insufficient to hand-craft examples for each, or when the question distribution is unknown or evolving; accept added pipeline complexity and slightly lower quality than hand-crafted few-shot.
- Few-Shot CoT PromptChoose when reasoning structure and consistency matter more than deployment speed (e.g., medical diagnosis, legal analysis, financial modelling) and expert time can be invested in two to three exemplary demonstrations per problem type.
- Zero-Shot CoT PromptChoose when immediate deployment is needed across diverse, unpredictable problem types without time or budget to craft examples, accepting variable reasoning quality on ambiguous problems.
Conflict Resolution Policy abstract
An abstract configuration that determines which rule fires when several rules match the current facts simultaneously.
- Layered Conflict Resolution PolicyChoose when fine-grained control is needed at scale: explicit priorities for business-critical rules, automatic specificity for most rules, and recency so updated rules take effect without re-prioritising.
- Order-of-Entry Conflict Resolution PolicyChoose when simplicity and determinism matter most and the rule base can be maintained as a carefully sequenced decision list with specific rules placed before general fallbacks.
- Recency Conflict Resolution PolicyChoose when knowledge evolves rapidly and newer rules encode more current understanding (e.g., new fraud patterns); requires rule modification timestamps and clear documentation.
- Salience (Priority) Conflict Resolution PolicyChoose when explicit control is needed, e.g., safety-critical systems where emergency rules must pre-empt optimisation or routine rules; requires careful priority assignment and maintenance.
- Specificity Conflict Resolution PolicyChoose when rules form natural general-to-specific hierarchies where nuanced policies should override broad defaults; needs a separate policy for equal-specificity conflicts.
Decision Engine abstract
An abstract cognition component that selects which action an agent takes from the currently available options given its model of the current state.
- Goal-Based Decision EngineChoose when binary goals suffice and no trade-offs exist; offers faster development and easier debugging.
- Heuristic Decision EngineChoose when real-time decisions must be made in milliseconds and a 'good enough' answer is acceptable, and when the heuristic's systematic bias causes minimal harm in the context (e.g., routine, clear-cut cases).
- Hybrid Decision ArbiterChoose when the problem mixes decision types with incompatible requirements, needs both adaptation and transparency, is safety-critical in an open world, or combines structured knowledge with unstructured perception (high decomposability, hard constraints, mixed data availability).
- Learned-Policy Decision EngineChoose when the strategy space is too large to specify manually, the environment is dynamic or partially unknown, and sufficient training data and computation are available.
- MDP Policy SolverChoose for sequential decisions where current actions affect future options and the state is (or has been augmented to be) Markovian.
- POMDP Policy SolverChoose when the problem is non-Markovian or history-dependent (e.g., treatment response depending on prior treatments) and state augmentation is insufficient.
- Pareto Frontier OptimizerChoose when multiple stakeholders hold diverse or disputed preferences, when trade-offs must be explored before committing, or when objectives are incommensurable and resist a common scale.
- Rule-Based Decision EngineChoose when rules capture the domain completely and consistently produce correct decisions, and transparency and predictability are paramount.
- Utility-Based Decision MakerChoose when decisions involve complex trade-offs among multiple objectives with no clear priority ordering, probabilistic outcomes, or continuous optimization where the degree of success matters more than binary goal achievement; use scalarized weights when a single decision-maker has clear preferences.
Decision Explainer abstract
An abstract explanation component that generates a justification for a specific agent or model decision in terms understandable to its audience.
- Counterfactual ExplainerChoose when users need to know the minimal changes that would flip the decision (why-not / recourse).
- Decision Factor ExplainerChoose when users need the key factors behind a decision and their relative importance.
- Example-Based ExplainerChoose when comparing the case to k similar past decisions conveys the rationale better than factor weights.
- Local Surrogate ExplainerChoose when the decision model is opaque and a model-agnostic local approximation of top feature contributions is needed.
Decision Fusion Aggregator abstract
An abstract cognition component that combines conclusions produced independently and in parallel by different decision paradigms into one unified decision.
- Constraint-Gated Decision FusionChoose when regulatory or safety rules must never be overridden by learned or utility scores.
- Learned Meta Decision FusionChoose when training data showing which combinations produce the best outcomes is available.
- Weighted Voting Decision FusionChoose when component accuracies on similar cases are known and redundancy should catch errors through disagreement.
Execution Plan abstract
An explicit, ordered or dependency-graph representation of the steps an agent will execute to achieve a goal.
Exemplar Selector abstract
An abstract component that chooses which demonstration examples from a demonstration pool are placed into a prompt, and in what order, for in-context learning.
- Cluster-Diverse Exemplar SelectorChoose when demonstrations must cover diverse problem types and reasoning chains are to be generated automatically rather than written manually.
- Coreset Exemplar Pre-selectorChoose when the demonstration pool is large or noisy and a compact core subset satisfying sufficiency and necessity is needed.
- Similarity Exemplar SelectorChoose when accuracy gains justify per-query retrieval cost; computational cost becomes high on large demonstration pools.
Factuality Verifier abstract
A semantic validation component that fact-checks generated claims against knowledge bases or retrieved evidence, producing correctness signals for runtime validation or reward computation.
- Fact Checking Rail abstract
- Alignment-Score Fact CheckerChoose for high-throughput scenarios such as customer service chatbots and Q&A systems.
- Cascaded Fact CheckerChoose for production deployments needing the lowest hallucination rate at mostly-minimal added latency.
- NLI Fact CheckerChoose for technical support and tutoring where moderate latency (100-150ms) is acceptable.
- Self-Check LLM Fact CheckerChoose for reasoning-heavy domains such as medical diagnosis where semantic similarity alone misses logical requirements.
- Alignment-Score Fact Checker
Heuristic Estimator abstract
A cognition component that estimates, for a search node, the remaining cost h(n) to the goal, guiding which candidates a search planner expands first.
- Geometric Distance HeuristicChoose when state coordinates are available and per-node evaluation must be near-free; the sweet spot for grid and road pathfinding.
- Learned Heuristic EstimatorChoose only when inference is cheap relative to node expansion cost, so higher accuracy reduces total search time.
Learned Decision Policy abstract
Learned parameters mapping states to actions, produced by reinforcement learning to maximize expected long-term reward.
- Policy NetworkChoose when states are high-dimensional (images, sensors) or actions are continuous, so tables cannot be stored or visited.
- Tabular Value FunctionChoose when the state-action space is small and discrete enough to enumerate and visit.
Reasoning Verifier abstract
An abstract external verification component that judges the correctness of an agent's reasoning steps or chain, independently of the model that produced them.
- Business Rule Logic VerifierChoose in regulated industries where business rules and domain constraints can be encoded as formal logic and compliance must be verifiably met.
- Circuit Reasoning VerifierChoose when model internals are accessible and verification must reflect what actually happened inside the model rather than its verbal explanation; requires per-domain calibration.
- Fine-Tuned Step VerifierChoose when domain-specific labeled correct/incorrect reasoning steps (human-annotated or synthetic) are available and higher in-domain accuracy is required, e.g., compliance, legal or policy reasoning.
- Formal Proof VerifierChoose for mathematical, symbolic or other formally specified domains where errors have severe consequences (financial calculation, cryptographic protocols, safety-critical systems) and reasoning can be translated into formal statements.
- Symbolic Math VerifierChoose when reasoning contains arithmetic or algebraic manipulations that can be extracted and independently re-computed with a symbolic mathematics library.
- Zero-Shot Step VerifierChoose when labeled reasoning-correctness data is unavailable and moderate agreement with human judgment suffices, using a properly prompted general LLM.
Reflection Critic abstract
An abstract cognition component that critiques a generated output against evaluation criteria, identifies errors or gaps, and triggers a refined generation.
- Dual-Agent CriticChoose when reducing self-reinforcing bias justifies added complexity and cost; requires critic criteria aligned with downstream consumers.
- Self-Reflection CriticChoose when implementation simplicity matters and the generating model has adequate domain knowledge; beware it may reinforce its own misconceptions.
- Stepwise Reasoning VerifierChoose when errors must be caught before they propagate through the chain; substantially improves final reasoning quality compared to single-pass generation with post-hoc evaluation.
Replanner abstract
A component that revises the current plan from execution error observations, e.g., inserting retries, substituting cached sources, or reordering steps.
- Complete ReplannerChoose when failures are global or rare (<10% of executions), unpredictable, memory is constrained, spaces are small, latency can be tolerated, or optimality matters more than continuity.
- Contingency Branch ActivatorChoose when failures are frequent (>30%), predictable from observable conditions (time of day, sensor state), stable, few (2-5 modes), and replanning pauses of 3-5 s are intolerable.
- Incremental Search ReplannerChoose when changes are frequent but localized edge-cost modifications on a stable topology, state spaces are large, memory permits search-tree persistence, and optimal solutions matter.
- Plan RepairerChoose when discrepancies are localized, adaptation must be fast, and plan stability matters (multi-robot committed trajectories, user-communicated ETAs); good-enough solutions suffice.
Risk Attitude Policy abstract
An abstract configuration fixing the curvature of a utility function, and hence whether the agent prefers certainty or gambles of equal expected value.
- Risk-Averse Utility PolicyChoose for safety-critical systems, medical treatment, retirement finance and other domains where worst-case outcomes carry severe consequences.
- Risk-Neutral Utility PolicyChoose only for small-stakes decisions with minimal outcome variance or explicitly risk-neutral domains aggregated across many independent events.
- Risk-Seeking Utility PolicyChoose for venture-style portfolios, innovation strategies with asymmetric upside, or winner-take-all competitive scenarios.
Self-Consistency Aggregator abstract
An abstract cognition component that combines final answers from multiple independently generated reasoning chains or agents into one consensus answer by voting, exposing the agreement distribution.
- Majority Vote AggregatorChoose for easy problems or small sample counts (k=3-5) where vote distributions are already clear and quality-assessment overhead is not justified.
- Quality-Weighted Vote AggregatorChoose for medium-to-hard problems with larger sample counts (k≈10-40) or cost-constrained high-stakes deployments where extracting more signal from fewer samples justifies ~5% quality-scoring overhead.
- Test-Based Vote AggregatorChoose for code generation where implementations vary textually but test cases define correctness.
State Value Estimator abstract
An abstract cognition component that estimates the value of a newly expanded search-tree leaf state, producing the reward signal backpropagated through the tree.
- Rollout SimulatorChoose when no trained value network is available and simulations are cheap, favouring heuristic rollout policies where domain knowledge exists.
- Value Network EvaluatorChoose when training data exists or self-play can generate it and upfront training cost plus per-evaluation inference latency are acceptable.
Structured Reasoning Prompt Template abstract
A prompt template that prescribes the format in which an agent exposes its reasoning, creating consistent step boundaries and evaluation checkpoints in reasoning traces.
- Numbered Step Reasoning TemplateChoose when the simplest structure suffices and clear step boundaries are needed for automated parsing and step-by-step evaluation.
- Tagged Reasoning TemplateChoose when evaluation must extract and assess specific reasoning components, e.g., whether verification actually checks results or conclusions quantify confidence.
Task Planner abstract
A cognition component that decomposes a high-level goal into a concrete ordered set of executable steps before execution begins.
- Contingency Planner
- Flat PlannerChoose for simple sequential workflows (3-10 steps) or optimization-dominant problems where provably optimal solutions matter more than fast feasible plans and the action space is small.
- Graph Search Planner abstract
- Bounded-Suboptimal Search PlannerChoose for real-time games, robots under time pressure and interactive systems where responsiveness outweighs optimality and cost may exceed optimal by up to factor w.
- Memory-Bounded Search PlannerChoose when memory is scarcer than time: embedded systems, mobile devices with strict memory budgets, or search spaces far exceeding available RAM.
- Optimal Heuristic Search PlannerChoose when the state space admits informative admissible heuristics, optimal solutions matter more than computational efficiency, the branching factor is manageable and transitions are deterministic (e.g., route planning on a known warehouse floor plan).
- Uniform-Cost Search PlannerChoose when many destinations share one origin (20+), when the graph is tiny (~50 nodes), when repeated queries on a static graph amortise preprocessing, or when no meaningful goal-distance heuristic exists.
- Bounded-Suboptimal Search Planner
- HTN PlannerChoose when the goal has naturally nested structure, reusable decomposition patterns exist, stakeholders need multiple abstraction views, or staged commitment under uncertainty is needed, and the world is stable during planning.
- LLM Task PlannerChoose for high-level decomposition of open-ended goals where broad knowledge matters and encoding deep domain methods in prompts would be too expensive.
- MCTS PlannerChoose when heuristics are hard to design, the state space is enormous, near-optimal solutions are acceptable for faster computation, transitions are stochastic or partially observable, or asymmetric tree growth pays off (game playing, robotic task planning, chemical retrosynthesis, navigation among dynamic crowds).
- Monte Carlo PlannerChoose for one-shot planning problems with expensive forward-model simulations (e.g., computationally intensive physics), where lower memory use outweighs MCTS's progressive statistical refinement.
- Reactive PlannerChoose for highly dynamic domains where strategic decisions become invalid before tactical decomposition finishes, or where problem structure is undefined.
Thought Exploration Controller abstract
An abstract cognition controller that maintains multiple candidate intermediate thoughts, deciding which to expand, evaluate, prune or combine and when to terminate a deliberate reasoning episode.
- Graph-of-Thought ControllerChoose when a problem decomposes into independent subproblems whose solutions must be synthesized, or intermediate results should inform rather than compete with each other.
- Tree Search Controller abstractChoose when exploration and backtracking matter but synthesis across branches does not.
- Breadth-First Thought Search ControllerChoose when trees are shallow (typically 2-4 steps), evaluation is uncertain so hedging across branches is valuable, or multiple solutions should be compared.
- Depth-First Thought Search ControllerChoose when trees are deep (5+ steps), evaluation is reliable enough to trust greedy choices, or any solution suffices.
- Hybrid Breadth-then-Depth Thought Search ControllerChoose when robustness against early misevaluation at shallow depths and efficient deep exploration are both needed and extra implementation complexity is acceptable.
- Breadth-First Thought Search Controller
Thought Generator abstract
An abstract cognition component that prompts a language model to produce k candidate next thoughts from the current reasoning state, conditioned on existing thoughts.
- Independent Sampling Thought GeneratorChoose for open-ended exploration of vast, unstructured solution spaces where large conceptual leaps are needed (e.g., creative writing plans).
- Sequential Proposal Thought GeneratorChoose for highly constrained, structured solution spaces where candidates should collectively and systematically cover the space (e.g., arithmetic operations).
Thought State Evaluator abstract
An abstract cognition component that uses a language model as a heuristic to assess how promising intermediate reasoning states are, producing signals that guide pruning and selection.
- Value Thought EvaluatorChoose when intermediate states can be checked for objective feasibility (e.g., mathematical reachability, constraint satisfaction) and absolute thresholds are needed for pruning.
- Vote Thought EvaluatorChoose when evaluation criteria are subjective or rating scales poorly defined (e.g., narrative coherence), or when a single best candidate must be selected.
Context Window Manager abstract
A memory component that tracks and manages how much conversation, reasoning, and observation history fits within the model's context window, signalling when older context will be dropped.
- Full Conversation BufferChoose for moderate conversation lengths (roughly 10-20 exchanges) where perfect recall of prior turns matters; switch to sliding-window or summarizing variants when history approaches token limits.
- Hierarchical History CompressorChoose for long-running conversations where historical context determines current interpretation and users reference earlier turns, accepting retrieval latency and the need to decide which history to fetch.
- Importance-Weighted History RetainerChoose for goal-oriented conversations in which certain turns establish critical constraints or preferences that later turns build on, provided an accurate relevance model is available.
- Sliding-Window History TruncatorChoose as the simplest bound when losing early conversation detail is acceptable.
- Summarizing History CompressorChoose when early context (user goals, key facts, binding decisions) remains relevant later and naive truncation would lose it.
- Trajectory PrunerChoose for multi-step agents whose trajectories accumulate dead ends, restatements and stale turns.
Conversation State Store abstract
A store holding per-session conversation history and agent state machine state used to continue a user's conversation across turns.
- External Session State StoreChoose for large-scale deployments (100+ instances), frequent autoscaling that would churn sessions, or when no session may be lost on instance failure.
- Instance-Local Session StateChoose when latency requirements are extreme (sub-100 ms), scale is modest (tens of instances), and session recreation after instance failure is acceptable.
Episode Encoder abstract
An abstract memory component that decides which experiences from an interaction become episodic memories and converts them into structured episode records.
- Event-Based Episode EncoderChoose when every action might be significant later (e.g., debugging complex technical problems where the solution hinges on minor details); avoid for routine standard-pattern interactions where storage overhead outweighs retrieval value.
- Significance-Based Episode EncoderChoose when most interactions follow standard patterns and storage must be reduced while preserving the most informative experiences (failures, unusual outcomes, high-impact states).
Working Memory Buffer abstract
Short-term, in-context storage of the current task's reasoning traces, tool invocations, observations, and reflection insights, bounded by the model's context window.
- Graph Reasoning State StoreChoose when thoughts need multiple parents so independent reasoning chains can merge.
- MCTS Search Tree Store
- Thought Tree StoreChoose when reasoning branches diverge but never need to reconverge.
Adaptive Retrieval Controller abstract
An abstract retrieval-gating component that decides per query whether to retrieve external knowledge or answer from the model's parametric memory, and which retrieval strategy to apply.
- Confidence-Gated Retrieval ControllerChoose when a simple gate suffices and the model's self-assessed certainty is reasonably calibrated; combine with pattern rules and periodic audits to catch miscalibration.
- Query-Type Retrieval RouterChoose when query types differ in their optimal retrieval strategy and more sophisticated routing than a single confidence gate is warranted.
Answer Synthesizer abstract
A software component that prompts an LLM to turn a question plus retrieved or queried results into a grounded natural-language answer.
- Multi-Hop Answer SynthesizerChoose when answers must combine sub-answers from multiple documents and cite the supporting sources for verification.
- Multimodal Answer Synthesizer
Content Deduplicator abstract
An abstract transformation component that detects and removes redundant copies of documents or chunks before they are indexed.
- Cascading DeduplicatorChoose when a corpus contains exact, near and semantic duplicates together (typical in production); orders cheap levels first to minimise cost.
- Exact Hash DeduplicatorChoose when duplicates are identical after cleaning and pairwise similarity comparison would be computationally prohibitive.
- Fuzzy Text DeduplicatorChoose when near-duplicates differ by minor edits, dates or formatting (versioning, format conversion); tune the threshold between conservative and aggressive.
- Near-Duplicate DetectorChoose when content varies slightly (whitespace, minor edits, reformatting) so exact hashing misses duplicates.
- Semantic DeduplicatorChoose when duplicates express the same information in different words (rewrites, summaries, translations); most expensive level.
Context Assembler abstract
A software component that combines retrieved document content and graph relationship context into a grounded prompt context.
Data Source Connector abstract
An abstract extraction component that interfaces with one class of source system to pull raw records or documents in their native format for an ETL pipeline.
- Event Stream ConsumerChoose for real-time or semi-structured event sources when near-instant knowledge updates justify the added complexity.
- File Store ExtractorChoose for unstructured document repositories such as shared drives, Git repositories or cloud object storage.
- Paginated API ExtractorChoose for semi-structured sources exposed through REST or GraphQL APIs (e.g., ticketing, CRM, wiki platforms).
- Source Record ExtractorChoose for structured sources (relational/NoSQL databases, warehouses) whose schemas and audit timestamps enable selective, batch delta queries.
Document Chunker abstract
A software component that splits raw documents into chunks for embedding and entity extraction.
- Fixed-Length ChunkerChoose only when content has no usable structure or timing data; the chapter presents it as the naive baseline that fractures semantic units.
- Hierarchical ChunkerChoose when document structure (sections, subsections) must be preserved so retrieved chunks keep their structural context, accepting more complex retrieval logic.
- Overlapping Window ChunkerChoose when important context spans chunk or section boundaries and continuity must be preserved, accepting multiplied storage.
- Semantic Boundary ChunkerChoose for text documents with structure, splitting at section headers and paragraph breaks rather than arbitrary token counts.
- Time-Indexed Transcript ChunkerChoose for audio transcripts lacking paragraph breaks or headers, when retrieved segments must link back to exact moments in the recording.
- Topic-Shift ChunkerChoose when chunk coherence matters more than uniform chunk size.
Embedding Service abstract
A service that converts text or other media into vector embeddings for similarity search.
- CPU Embedding ServiceChoose only when GPU infrastructure is unavailable and query volume is low (below the few-thousand-queries-per-day GPU break-even); pair with an embedding cache.
- Hosted Embedding API ServiceChoose when building initial development or baseline systems, when no data-sovereignty or air-gap constraint applies, and when per-token API cost is acceptable for the query volume.
- Joint Multimodal Embedding ServiceChoose for rapid deployment on existing text RAG with mostly general imagery (photos, simple diagrams); avoid for information-dense charts needing precise values or OCR.
- Modality-Specific Embedding ServiceChoose when each modality needs its best-suited embedding model (e.g., domain BERT for text, CLIP for images, DePlot for charts) and components must be upgradable independently.
- Self-Hosted GPU Embedding ServiceChoose when data sovereignty, zero-trust or air-gapped operation is required, when long documents must be embedded, or when volume is high enough (beyond a few thousand queries daily) that GPU per-query efficiency beats API pricing; requires operating GPU and serving infrastructure.
- Text Embedding ServiceChoose when all modalities are grounded to text (captions, linearized tables, transcripts) so a single text embedding model and text vector index suffice.
Entity Linker abstract
An abstract software component that resolves entity mentions with varying surface forms to a single canonical entity identifier.
- Fuzzy-Match Entity LinkerChoose for lower-precision applications; matches by edit distance with manual review of uncertain mappings.
- Knowledge-Base Entity LinkerChoose for high-precision requirements; resolves mentions to knowledge-base IDs via an entity linking service.
- Multi-Signal Entity ResolverChoose when the same real-world entity carries different identifiers across several source systems and string similarity alone is insufficient to unify them.
Image Type Classifier abstract
An abstract preprocessing classifier that assigns each image a content class, such as chart/plot versus general image, to drive processing-path selection.
- Heuristic Image Type ClassifierChoose when a simple rule suffices to spot charts, e.g., images containing axis labels, legends or grid patterns, as in the financial-report worked example.
- VLM-based Image Type ClassifierChoose when routing at scale and a vision-language model is already deployed, reusing it for meta-classification instead of adding a separate classification model.
Image-to-Text Grounder abstract
An abstract preprocessing component that converts image content into searchable text (captions or structured data) so it can be embedded and retrieved with standard text retrieval.
- Chart Data ExtractorChoose for information-dense images (charts, plots, graphs, tables) where users need exact figures and computation, e.g., financial or benchmark charts.
- Image CaptionerChoose for natural images, complex scenes, and diagrams with readable text where descriptive detail (objects, spatial relationships, context, visible text) matters more than exact numbers.
Knowledge Graph Store abstract
A store of entities and relationships enabling multi-hop relational and causal reasoning that vector similarity cannot represent.
- Property Graph StoreChoose when rapid development and schema flexibility (properties on nodes and edges, multiple labels, schema evolution) matter more than formal semantics; the industry-standard choice for agent systems.
- RDF Triple StoreChoose when interoperability across systems requires standardized ontologies and formal W3C semantics (e.g., academic knowledge bases like Wikidata).
Knowledge Store Loader abstract
An abstract load-stage component that inserts processed records, embeddings and metadata into a target knowledge store.
Relation Extractor abstract
An abstract software component that identifies typed semantic relationships (triples with properties) between entity pairs mentioned in text.
- Dependency-Parse Relation ExtractorChoose as a reliable baseline for explicit subject-verb-object relationships; it misses implicit relationships stated without verbs.
- Neural Relation ExtractorChoose when implicit relationships (e.g., appositives like 'Google, a major Anthropic investor') must be captured for higher recall.
Reranker abstract
A software component that deduplicates and re-scores candidate results from one or more retrieval sources into a single ranking.
- Cross-Modal Reranker
- Lexical RerankerChoose instead of full hybrid fusion for extremely latency-sensitive applications (sub-50 ms p99) that still need keyword precision.
- Multi-Signal Relevance Ranker
Retriever abstract
An abstract software component that fetches the information needed to answer a query from an indexed knowledge source.
- Dense-Sparse Hybrid RetrieverChoose in latency-sensitive RAG agents where retrieval dominates time-to-first-token.
- Graph RetrieverChoose when queries require following typed relationships across multiple hops (e.g., conflict-of-interest chains, investor overlap).
- Hybrid Retriever abstractChoose when queries need both semantic similarity and relationship traversal.
- Graph-Constrained Vector Retriever
- Graph-Enhanced RetrieverChoose when retrieved chunks need relationship context invisible to vector search (authorship, employment history, contradictions between documents).
- Parallel Fusion RetrieverChoose when you cannot predict whether vector search or the graph will surface the needed information, trading complexity for completeness.
- Retrieval-Augmented Graph RetrieverChoose when the graph is too large for exhaustive traversal (millions of nodes, billions of edges) and vector search can identify the entities around which to scope traversal.
- Keyword RetrieverChoose alongside dense vector search when sparse keyword matching can return results faster, run in parallel to cut retrieval latency.
- Multimodal Retriever abstract
- Grounded Text RetrieverChoose for information-dense visuals (financial reports, scientific charts, technical diagrams) where precise chart content must be retrievable and existing text RAG infrastructure should stay unchanged.
- Per-Modality Fan-Out RetrieverChoose for research or experimentation with best-in-class per-modality embedding models, or mature-MLOps production systems able to absorb the extra complexity and cost.
- Unified Embedding RetrieverChoose for general imagery, rapid prototyping, or minimal change to an existing text RAG stack with a single vector store.
- Grounded Text Retriever
- Vector Retriever abstractChoose when queries need conceptual understanding or semantic similarity across varied terminology (simple Q&A, documentation search, customer support).
- Approximate Vector Search RetrieverChoose when sub-linear, millisecond-latency search over millions of vectors matters more than exact, reproducible top-K results.
- Coarse-to-Fine Vector RetrieverChoose when embeddings are MRL-trained and corpus scale makes full-dimensional search latency or memory prohibitive.
- Exact Vector Search RetrieverChoose when deterministic, reproducible retrieval is required (scientific experiments, compliance audits, A/B testing), accepting slower queries.
- Metadata-Filtered Retriever
- Approximate Vector Search Retriever
Text Embedding Model abstract
A trained encoder model that maps text (queries, knowledge chunks, episode summaries) to fixed-dimension dense vectors whose cosine similarity approximates semantic relatedness.
- Cost-Optimized Embedding ModelChoose for cost-sensitive applications with moderate accuracy needs and as the default starting point to establish baseline retrieval performance; open-source variants when self-hosting without API dependencies is required.
- Domain-Specific Embedding ModelChoose for specialized applications with technical or industry-specific terminology, where domain models outperform general-purpose ones.
- General-Purpose Embedding ModelChoose for broad coverage across diverse content types.
- High-Accuracy General Embedding ModelChoose when retrieval accuracy directly drives user experience and passages are short to medium length, justifying higher per-token cost.
- Long-Context Embedding ModelChoose for enterprise deployments processing long technical documents, contracts or papers that routinely exceed 8,000 tokens, typically self-hosted for data sovereignty.
- Retrieval-Optimized Embedding ModelChoose for search-specific applications over short passages where ranking precision and near-duplicate discrimination matter more than long context.
Vector Index Build Configuration abstract
A build-time configuration fixing a vector collection's index type, graph connectivity (M), construction candidate-list size (efConstruction) and distance metric; changing it requires full re-indexing.
- Flat Index ConfigurationChoose only for small collections (under about 100K vectors) where exact results are needed.
- HNSW Index ConfigurationChoose for large collections (beyond ~10 million vectors) where search latency outweighs higher memory usage.
- IVF Index ConfigurationChoose as a balanced default for mid-size collections needing good recall with manageable memory.
Vector Index Store abstract
A database that stores vector embeddings and answers similarity queries.
- CPU Vector Index StoreChoose when collections and query rates are modest enough that CPU-based index building and search latency are acceptable.
- Distributed Vector Index StoreChoose when query throughput, storage efficiency, multi-tenancy or on-premises constraints demand capabilities simpler vector databases lack and an infrastructure team can absorb the operational complexity.
- Embedded Vector Index StoreChoose for local development, notebooks, proofs of concept, education and embedded applications with modest scale (well under one million vectors).
- Filter-Optimized Vector Index StoreChoose when filtered vector search is the core workload (recommendations under business-rule constraints, document search with access control, analysis within temporal windows); available self-hosted or managed.
- GPU-Accelerated Vector Index StoreChoose for large (up to billion-scale) collections, continuous ingestion with frequent index rebuilds, or multi-hop agent workflows querying many times per request under tight latency.
- Managed Vector Index StoreChoose when rapid deployment and operational simplicity outweigh infrastructure control and cloud-hosted storage is acceptable.
- Modality-Specific Vector Store
- Relational Vector Extension StoreChoose when an organization already operates the relational database at scale, datasets are moderate (<1M vectors), and transactional consistency across relational and vector operations is required.
- Self-Managed Vector Index StoreChoose when a team comfortable operating infrastructure needs to deploy anywhere (on-premises for data sovereignty, own cloud for cost control) without sacrificing developer experience.
Vector Store Query API abstract
The network interface through which clients authenticate to a vector store and submit schema, ingestion and search requests.
- Vector Store REST APIChoose for broad client compatibility, debugging and moderate query volumes (below ~100 requests per second).
- Vector Store gRPC APIChoose when query volume exceeds ~100 requests per second or real-time applications need the lowest latency.
Attention Kernel abstract
An abstract accelerator kernel implementation that computes transformer self-attention over the current sequence and cached key-value projections.
- Fused Block-wise Attention KernelChoose by default for any transformer deployment, especially long contexts on memory-bandwidth-limited GPUs; benchmark against standard attention for very short sequences.
- Sparse Attention KernelChoose when much larger effective contexts are needed and occasionally missing long-range dependencies that full attention would capture is acceptable.
- Standard Attention KernelChoose only when benchmarks show fused kernels are slower, e.g., very short sequences (under 256 tokens) with small batches.
Draft Token Proposer abstract
Model weights that propose K speculative future tokens for a target model to verify in one parallel forward pass, selected for distributional alignment with the target rather than task accuracy.
- Self-Speculative Decoding HeadChoose when memory is constrained or batch size must be preserved (e.g., edge devices, high-batch serving).
- Speculative Draft Model abstract
- Distilled Draft ModelChoose for production where improved acceptance justifies 50-100 GPU-hours of training; use online distillation at scale to track target shifts.
- Independent Draft ModelChoose for rapid prototyping where zero training cost and immediate deployment outweigh modest acceptance.
- Distilled Draft Model
Foundation LLM abstract
General-purpose large language model weights, ranging from frontier models used for planning to smaller models used for execution.
- Domain-Adapted Base Model
- Fine-Tuned Agent Model
- Large Language Model TierChoose for complex queries or premium users where the quality improvement justifies cost and sharding complexity.
- Reasoning Language ModelChoose for complex queries where accuracy improvements justify the substantially higher token cost.
- Reference Policy Model
- Small Language Model TierChoose for simple or domain-specific queries where benchmarking shows it meets quality thresholds.
- Standard Language Model TierChoose for queries classified as moderately complex.
Inference Backend abstract
A pluggable execution module, loaded dynamically by an inference server according to model configuration, that runs a model in one specific framework runtime.
- LLM Generation BackendChoose for large language model generation, where the engine's continuous or in-flight batching outperforms request-level batching.
- Portable LLM Runtime BackendChoose for uncommon or heterogeneous GPUs (RTX 4090, A10G, V100, T4), edge and mixed research clusters, or development prioritizing iteration speed.
- Pre-compiled Engine BackendChoose when pre-compiled engines exist for the deployed GPU (e.g., A100 40/80GB, H100 80GB, L40S 48GB).
- Tensor Framework BackendChoose for non-autoregressive models (classifiers, encoders, vision, tree-based models, custom Python logic) that benefit from server-side dynamic batching; OpenVINO for Intel CPU or edge CPU-only infrastructure.
Inference Batch Scheduler abstract
A server-side scheduler that groups concurrent inference requests into shared GPU executions to raise hardware utilisation, transparently to clients.
- Dynamic Batch SchedulerChoose for models on backends that accept batched tensors (TensorFlow, PyTorch, ONNX Runtime, TensorRT); do not use for engines with internal continuous batching.
- In-Flight Batch SchedulerChoose for autoregressive LLM generation where sequences finish at different times and latency and throughput must be balanced without artificial batch-assembly delays.
- Sequence Batch SchedulerChoose when the served model is stateful (recurrent networks, language models with hidden state, streaming speech recognition, conversational models carrying context) so every request of a sequence must reach the same model instance.
- Static Batch SchedulerChoose for offline batch workloads (overnight reports, bulk document processing, scheduled evaluation) where no user waits.
Inference Service Image abstract
A container image packaging an inference runtime, engine-selection logic and a standard API so the same image deploys identically across cloud, data center and workstation GPUs.
- Model-Specific Inference ImageChoose for production with SLA or compliance requirements (SOC 2, HIPAA, FedRAMP) and mission-critical services where 10-15% latency gains justify reduced flexibility.
- Multi-Model Inference ImageChoose for research, experimentation, custom fine-tuned model pipelines, multi-model task switching and prototyping where experimentation velocity outweighs production stability.
Inference Serving Configuration abstract
A configuration artifact fixing the served model, sampling temperature, maximum output tokens, engine optimisation, and batching and streaming modes for an inference endpoint.
- Balanced Batching ConfigChoose for mixed or semi-interactive workloads, such as customer service chatbots where users tolerate 300-500 ms latency and traffic allows some batching.
- Inference Queue PolicyChoose when serving heterogeneous users with tiered SLAs (premium <50ms, standard <200ms, best-effort batch) on shared infrastructure.
- Latency-Oriented Batching ConfigChoose for interactive query-time workloads such as visual question answering where sub-second responses are required.
- Latency-Tuned Decoding ConfigurationChoose for interactive applications (chat, code completion, real-time moderation) with short, factual responses where deterministic, concise output is acceptable.
- Model Deployment Profile abstract
- Latency-Optimized Deployment ProfileChoose for latency-critical interactive applications where time-to-first-token SLOs dominate.
- Memory-Optimized Deployment ProfileChoose when available GPU memory is insufficient for full-precision weights plus activations, such as smaller GPUs or edge devices.
- Throughput-Optimized Deployment ProfileChoose when maximising queries per second matters more than minimum time-to-first-token and multiple GPUs are available.
- Latency-Optimized Deployment Profile
- Throughput-Oriented Batching ConfigChoose for offline RAG preprocessing (bulk captioning or extraction) where images processed per second matters more than per-image latency.
KV Cache Allocator abstract
An abstract inference-engine component that allocates and releases accelerator memory for each request's key-value cache.
- Contiguous KV Cache AllocatorChoose (with buffers pre-allocated for maximum length) only when sequence lengths are bounded and predictable, accepting higher initial memory use.
- Paged KV Cache AllocatorChoose by default for production serving, especially workloads with high variance in generation length; benefit is minimal for fixed-length outputs.
KV Cache Eviction Policy abstract
An abstract policy deciding which requests' KV caches to evict or defer when cache demand exceeds available memory.
- Cost-Aware Hybrid KV Cache Eviction PolicyChoose in production when recency, priority and recomputation cost must be balanced together.
- FIFO KV Cache Eviction PolicyChoose when all requests have similar priority.
- LRU KV Cache Eviction PolicyChoose for interactive sessions where recent activity signals continued engagement and idle sessions can be suspended.
- Priority KV Cache Eviction PolicyChoose when requests carry differentiated SLAs (e.g., premium versus free tier).
Native Function-Calling API abstract
A model-serving interface that accepts tool definitions as JSON schemas and returns selected tool calls as guaranteed-valid structured JSON objects instead of free text requiring parsing.
- Hosted Provider Inference APIChoose when data residency, security policy, and cost considerations permit sending prompts to a cloud-hosted LLM provider.
- Self-Hosted Inference EndpointChoose when data residency requirements, security policies, or cost considerations prevent using cloud-hosted LLM APIs, or for high-throughput tool-heavy agents.
Optimized Inference Engine abstract
A compiled, precision-reduced model engine produced for low-latency, high-throughput serving.
- FP16 Inference EngineChoose as the default when accuracy cannot be compromised (e.g., precise numerical reasoning where even 1% degradation is unacceptable).
- FP4 Inference EngineChoose for maximum compression and throughput in use cases that accept ~3-7% accuracy loss.
- FP8 Quantized EngineChoose on GPUs with FP8 support when roughly doubled throughput with minimal accuracy loss is acceptable for the model.
- INT4 Quantized EngineChoose only when 2-5% accuracy loss is acceptable and maximum memory and cost reduction is needed.
- INT8 Quantized EngineChoose when 1-2% accuracy reduction is tolerable (e.g., conversational or human-reviewed moderation workloads) in exchange for roughly halved infrastructure cost; requires representative calibration.
- TF32 Inference EngineChoose when targeting Ampere-or-newer hardware exclusively and near-FP32 accuracy is required without a quantization workflow.
Query Complexity Assessor abstract
An abstract component that classifies an incoming query's complexity (e.g., simple, moderate, complex) before generation so that a router can select an appropriately sized model.
- Query Complexity ClassifierChoose when query phrasing varies widely and higher classification accuracy justifies the added latency and cost of a classifier model.
- Rule-Based Complexity ClassifierChoose when query classes follow clear lexical patterns and zero added latency and cost matter more than robustness to phrasing variation.
Response Cache abstract
A cache of complete agent/LLM responses keyed by (normalized) query so repeated requests are answered without inference.
- Database Response CacheChoose when audit requires persisting all outputs, capacity exceeds practical RAM, cache updates must be atomic with other data, or SQL querying of cached data is needed.
- Distributed Response CacheChoose for horizontally scaled systems with >20% query repetition or where cache must survive instance restarts.
- In-Process Response CacheChoose for single-instance deployments, session-scoped data, or sub-millisecond latency needs where even Redis overhead matters.
- Semantic CacheChoose when users ask the same questions in different wording and incremental savings exceed embedding costs.
Vision-Language Model abstract
A generative multimodal model combining a vision encoder with a language model via cross-attention, producing captions and answers about images, including reading visible text.
- Balanced Vision-Language ModelChoose for cloud deployment needing a balance of accuracy and compute.
- Edge Vision-Language ModelChoose for edge deployment on embedded GPU devices where memory and power are constrained.
- High-Accuracy Vision-Language ModelChoose for demanding applications requiring the highest visual reasoning accuracy.
Fine-Tuning Pipeline abstract
A software component that adapts model weights to a domain or task using curated training data.
- Full-Parameter Fine-TunerChoose when substantial multi-GPU infrastructure is available and all model weights are to be modified.
- LoRA Fine-Tuner abstractChoose when GPU memory is constrained and near-full-fine-tuning quality is needed by training only a small adapter.
- QLoRA Fine-TunerChoose when even LoRA exceeds available hardware memory and a small quality loss is acceptable.
- QLoRA Fine-Tuner
Policy Learner abstract
An abstract model-adaptation component that updates a learned policy model from interaction experience or expert demonstrations.
- Imitation Learner abstractChoose when expert demonstrations are available and autonomous exploration is too slow, costly, or unsafe.
- Behavior Cloning TrainerChoose when demonstrations comprehensively cover deployment situations, the expert is consistent, and decision sequences are short enough that errors do not compound.
- DAgger TrainerChoose when experts can provide repeated labels and the learned policy can be safely executed during training to expose its failure modes.
- Inverse RL Reward LearnerChoose when demonstrations come from multiple experts with different strategies, deployment differs qualitatively from demonstrations, or the reward is needed for explanation or policy evaluation.
- Behavior Cloning Trainer
- Multi-Agent Policy Learner abstract
- Centralized-Training Decentralized-Execution LearnerChoose when training can occur with full observability (e.g., simulation) but deployment requires distributed execution with limited communication.
- Independent Multi-Agent LearnerChoose when agents interact loosely so the environment appears approximately stationary to each agent.
- Centralized-Training Decentralized-Execution Learner
- Reinforcement Learning Policy LearnerChoose when a reward signal is available and the agent can safely explore (e.g., in simulation), with ample data and compute.
Preference Elicitor abstract
An abstract component that derives utility or reward function parameters representing a principal's preferences from observed behaviour instead of explicit engineering.
- Inverse Reward LearnerChoose when expert demonstrations (e.g., thousands of hours of human driving) are available and explicit preference specification is impractical.
- Revealed Preference LearnerChoose when users' selections among presented options and feedback are continuously observable, enabling personalization without explicit preference configuration.
Preference Optimizer abstract
An abstract model-adaptation component that updates a supervised-fine-tuned policy model so its outputs better match human preferences expressed as comparative judgments.
- Direct Preference OptimizerChoose first for simpler preference-learning tasks where direct optimization suffices and lower compute is desired.
- RLHF Policy OptimizerChoose for complex multi-dimensional preferences where explicit reward modeling provides interpretability benefits.
Reward Scorer abstract
An abstract component that computes the scalar reward signal for a prompt-response pair used to guide reinforcement-learning policy optimization.
- Composite Reward ScorerChoose when human preference must be grounded by verifiable correctness, e.g., to prevent optimizing toward fluent hallucinations or annotator bias.
- Constitutional Reward ScorerChoose when harmlessness rewards must scale without human labels and reflect explicit, inspectable principles rather than implicit annotator preferences.
- Ensemble Reward ScorerChoose when reward hacking via single-pattern exploitation is a concern and the cost of training multiple reward models is acceptable.
- Fairness-Constrained Reward ScorerChoose for high-stakes decisions such as lending where the policy must satisfy a chosen fairness metric (e.g., equalized odds) while keeping legitimate business criteria.
- Preference Reward ScorerChoose when desired behaviour is a subjective judgment (e.g., tone, trade-offs) that rigid rules cannot capture and annotator bias is controlled.
- Segment-Routed Reward ScorerChoose when user segments legitimately hold divergent preferences and sufficient preference data can be collected for each segment.
Rule Learner abstract
An abstract adaptation component that proposes new or refined if-then rules from labelled decision cases while preserving interpretable rule structure.
- Case-Based Rule RefinerChoose when an expert-provided initial rule set exists and should be refined from accumulated cases where its decisions proved wrong.
- Inductive Rule LearnerChoose when rules must be induced from historical labelled examples (e.g., past approved/denied applications) rather than from an existing expert rule set.
API Gateway Proxy abstract
An abstract reverse proxy at the system entry point that applies cross-cutting traffic policies and forwards requests to backend agent services.
- High-Performance Reverse ProxyChoose when maximum performance and operational simplicity outweigh API-management features, authentication/rate limiting live elsewhere, and routing plus load balancing suffice.
- Policy Plugin GatewayChoose when centralised policy management, varied authentication schemes (OAuth2, JWT, API keys), traffic transformation or complex routing justify added latency and configuration complexity.
Agent Hosting Platform abstract
An abstract compute substrate on which agent services are packaged, executed and scaled, trading operational control and warm capacity against management overhead and idle cost.
- Container OrchestratorChoose when consistent latency matters more than cost (24/7 warm instances, no cold starts), Kubernetes already runs other services to amortise overhead, the team has strong Kubernetes expertise, continuous background processing is needed, or debugging requires persistent logs and long-term metric retention.
- Serverless Function RuntimeChoose when traffic shows high variance or long idle periods, deployment velocity is critical, the team lacks infrastructure expertise, cost must scale with actual usage, or the system processes discrete events rather than continuous streams.
Autoscaler abstract
A control component that adjusts the number of replicas of a workload to match demand, turning fixed infrastructure cost into variable cost.
- Latency-Target AutoscalerChoose for latency-sensitive deployments with explicit response-time SLOs.
- Metric-Driven AutoscalerChoose for stateless agents with unpredictable, high-variance traffic (e.g., multi-tenant SaaS) that tolerate 60-90 s scale-up lag.
- Queue-Depth AutoscalerChoose for agents that primarily orchestrate external API or tool calls, where CPU does not reflect capacity.
- Resource-Utilization AutoscalerChoose CPU for compute-bound agents dominated by inference or reasoning loops, and memory for agents holding large context windows or document embeddings.
- Scheduled ScalerChoose when traffic patterns are extremely predictable, making schedule-based capacity simpler and equally effective.
Dataframe Compute Engine abstract
An abstract data-processing engine that executes dataframe operations such as filtering, deduplication and text normalization for ETL transformation stages.
- CPU Dataframe EngineChoose for datasets under ~10 million documents, infrequent (monthly or rarer) processing, simple quality filters without deduplication, or limited capacity to manage GPU infrastructure.
- GPU-Accelerated Dataframe EngineChoose for billions of documents where CPU pipelines take multiple days, regular full reprocessing of large corpora, fuzzy deduplication at scale, or when GPU infrastructure already exists.
Distributed Model Executor abstract
An execution component that partitions a model's training or inference computation across multiple GPUs according to a parallelism strategy and synchronizes partial results through collective communication.
- Data Parallel ExecutorChoose for models under ~13B that fit on one GPU; scales nearly linearly.
- Fully Sharded Data Parallel ExecutorChoose for training models whose weights plus optimizer state exceed GPU memory (e.g., 70B with 280 GB state on 4x80GB).
- Pipeline Parallel ExecutorChoose for cross-node scaling over slower interconnects where memory efficiency matters more than per-sample latency.
- Tensor Parallel ExecutorChoose for memory-constrained models (30-40B+) and latency-sensitive inference, confined to a single-node high-bandwidth GPU interconnect domain.
Edge Device abstract
An abstract resource-constrained compute device at or near the point of data generation that runs inference locally, possibly without network connectivity.
- Edge FPGA DeviceChoose for specialised applications needing ultra-low latency when hardware programming expertise is available.
- Edge GPU DeviceChoose for robotics and autonomous systems needing programmable acceleration across many model types, accepting higher power draw.
- Edge NPU DeviceChoose for mobile and embedded products where dedicated AI acceleration must consume minimal power.
- Edge TPU DeviceChoose for battery-powered devices running standard neural network architectures.
- Microcontroller DeviceChoose for sensor-scale, battery-powered deployments where memory and power budgets are minimal.
GPU Allocation Unit abstract
An abstract schedulable unit of accelerator capacity (whole GPU, hardware partition, or time-shared slot) onto which the orchestrator places a single workload.
- Dedicated GPU DeviceChoose for large models (70B+), high-concurrency single-tenant deployments, or latency-critical workloads needing full memory bandwidth and large continuous batches.
- GPU PartitionChoose for multi-tenant SaaS with SLA guarantees, mission-critical agents, regulatory isolation needs, or production agents that must not be slowed by co-located batch jobs.
- Time-Sliced GPU ShareChoose for internal development, trusted co-located workloads, long-running batch jobs, or cost-constrained deployments that tolerate latency variance and cannot afford partitioning overhead.
GPU Interconnect abstract
The intra-node data path linking GPUs for peer memory access and collective communication, whose bandwidth and latency bound multi-GPU parallel efficiency.
- GPU Switch FabricChoose for 8+ GPU nodes and high-batch tensor-parallel inference where all-reduce bandwidth dominates.
- PCIe GPU BusChoose for smaller or lower-utilization deployments and data-parallel workloads where interconnect bandwidth is not the bottleneck.
- Peer-to-Peer GPU LinkChoose for tensor-parallel inference or training of >40B models with sustained utilization (100+ GPU-hours/month).
GPU Node abstract
An accelerated compute host (single or multi-GPU, linked by high-bandwidth interconnect) on which inference and agent workloads run.
- Edge GPU DeviceChoose for robotics and autonomous systems needing programmable acceleration across many model types, accepting higher power draw.
- On-Demand GPU NodeChoose for critical real-time production workloads requiring guaranteed availability.
- Reserved GPU NodeChoose for the predictable, continuous baseline load present even in low-traffic periods.
- Spot GPU NodeChoose for batch, development, and embarrassingly parallel background tasks that tolerate interruption.
GPU Partition Layout abstract
A declarative allocation of a physical GPU into isolated instances with fixed memory and compute fractions assigned per workload.
- Mixed GPU Partition LayoutChoose for multi-tier offerings and diverse model sizes (7B-70B) when the team accepts extra scheduling and configuration complexity for better utilisation.
- Uniform GPU Partition LayoutChoose for homogeneous workloads and SaaS platforms with many similar-sized tenants where operational simplicity is the priority.
- Unpartitioned GPU LayoutChoose for 70B+ models, high-concurrency single-tenant or latency-critical deployments needing a full GPU.
Load Balancer abstract
A distribution component that spreads requests across multiple instances of an agent or service using health checks.
- Cache-Aware Inference RouterChoose when inference replicas hold non-equivalent state (KV caches, loaded adapters) and multi-agent routing must track which pods cache which contexts.
- DNS Load Balancer
- Layer-4 Load BalancerChoose when all instances are identical and all requests equivalent, so raw throughput and simplicity outweigh request-aware routing.
- Layer-7 Load BalancerChoose when request-based routing, application-aware health checking, TLS termination, or cookie-based session affinity is required.
- Least-Connections Load BalancerChoose as the default when request durations vary widely (e.g., 90% under 2 s, 10% over 10 s).
- Metric-Aware Load BalancerChoose only when simpler heuristics (least connections, weighting) fail to maintain SLOs and observability infrastructure can supply fresh metrics.
- Round-Robin Load BalancerChoose only when all requests have similar processing cost and all replicas identical capacity.
- Session Affinity Load BalancerChoose as the simplest cache-locality option for conversational agents whose follow-up messages come from the same client.
- Weighted Round-Robin Load BalancerChoose when replicas run on heterogeneous infrastructure with measurable capacity differences, or for weighted gradual rollouts.
Rollout Manager abstract
An abstract release component that replaces a running agent or model version with a new one according to a rollout strategy while bounding user impact and enabling rollback.
- Blue-Green Deployment SwitcherChoose when validation completes before cutover and instant switch-over with instant failback to a fully running previous version is required.
- Canary Rollout ControllerChoose when a new version should be validated on real user traffic with limited exposure before full rollout.
- Rolling Update ControllerChoose gradual one-at-a-time replacement when safety matters more than update speed; replacing all instances at once is faster but risky.
Serverless Hosting Plan abstract
An abstract configuration selecting how a serverless runtime provisions capacity for a function, trading cold-start latency and timeout limits against idle cost.
- On-Demand Consumption PlanChoose for variable, unpredictable or infrequently triggered agent workloads.
- Pre-warmed Capacity PlanChoose for latency-sensitive agents requiring consistent performance; avoid enabling it for all functions, since it pays for idle capacity and removes serverless's main cost benefit.
- Reserved Capacity PlanChoose for predictable, constant high-volume agent workloads where reserved capacity costs less than pay-per-execution billing.
Version Rollback Controller abstract
An abstract release component that returns production traffic from a newly deployed agent or model version to the previous known-good version.
- Gradual Rollback ControllerChoose for moderate quality degradation, performance issues or non-critical bugs.
- Immediate Rollback ControllerChoose for critical issues: error rate > 5%, latency > 200% of baseline, system crash/unavailability or detected data corruption.
Workload Controller abstract
An abstract orchestrator control loop that continuously maintains the declared replica set of one containerised workload, creating, replacing and updating instances to match desired state.
- Stateful Workload ControllerChoose when agents or stores need persistent identity, ordered initialization (e.g., coordinator before workers) or reconnection to specific persistent volumes.
- Stateless Workload ControllerChoose for stateless agent replicas that load the same model, keep no local state and execute tasks independently (most worker agents).
Agent Version Experimenter abstract
An abstract production experimentation component that compares a candidate agent version with the current production version on real traffic before full deployment.
- A/B Test Traffic SplitterChoose when comparing versions on live outcome metrics with users actually served by the new version, controlling for temporal and population effects.
- Shadow Test RunnerChoose when a new version must be assessed on production inputs before committing to deployment, without executing its actions.
Benchmark Environment abstract
An interactive, reproducible task environment (e.g., operating system shell, database, knowledge graph, web application, game) that exposes actions and state to an agent under evaluation over multi-turn episodes.
- Decision Scenario SimulatorChoose to test utility-driven decisions across diverse and edge-case scenarios before deployment.
- Simulated Web Environment
Evaluation Sampling Policy abstract
A configuration setting what fraction of production interactions are evaluated at each deployment stage and when sampling rates increase in response to anomalies.
- Exhaustive Evaluation Sampling PolicyChoose when comprehensive per-trace ground truth justifies ~2-5% of inference cost and ~500 ms added latency per request.
- Random Evaluation Sampling PolicyChoose when a representative baseline-quality estimate is needed with no production latency impact.
- Strategic Evaluation Sampling PolicyChoose for production systems where evaluation should focus on problem areas; recommended default (5-10% with error oversampling).
Metrics Dashboard abstract
A visualization component that queries stored metric time series and renders graphs, heatmaps and gauges organised around operational questions.
- Adoption Dashboard
- Agent Performance DashboardChoose when managers must review multi-agent collaboration effectiveness and escalation patterns across thousands of conversations.
- Compliance Dashboard
- Cost Reporting DashboardChoose when the audience is budget owners or executives who need trend, model-mix, unit-economics and budget-versus-actual views rather than operational health panels.
- Risk Dashboard
Metrics Endpoint abstract
An HTTP endpoint (conventionally /metrics) on an instrumented agent or service that exposes its counters, gauges, histograms and summaries in a scrapeable text format.
Performance Profiler abstract
An abstract observability component that captures execution activity of an agent or inference workload at a chosen granularity so elapsed time and resource use can be attributed to stages.
- Execution ProfilerChoose to decide whether multi-step agent latency is dominated by LLM inference or external tool latency, and whether multi-agent parallelism is effective or serialized.
- GPU System ProfilerChoose for routine and production diagnosis of where time is spent across CPU-GPU boundaries, since event-driven tracing with CPU sampling keeps overhead at about 1-3%.
- Inference Engine ProfilerChoose when system profiling shows inference dominates and the time split across attention kernels, KV-cache access, quantized operations and batching must be understood.
- Kernel ProfilerChoose only for controlled deep-dive optimization of individual kernels (warp divergence, memory access, register use); its 10-100x slowdown precludes continuous production use.
Production Quality Monitor abstract
An abstract monitoring component that computes model-quality, drift and data-quality metrics over production predictions against a reference baseline and raises alerts on anomalies, compensating for silent failures and delayed ground truth.
- Batch Quality MonitorChoose for non-real-time applications, cost-sensitive monitoring and historical analysis.
- Streaming Quality MonitorChoose for critical or high-volume applications that need real-time alerts.
Reasoning Quality Scorer abstract
An evaluation component that combines per-dimension reasoning scores (intra-step correctness, inter-step consistency, informativeness, relevancy) into an overall reasoning quality result.
- Minimum Aggregation Quality ScorerChoose when agents must not receive full credit for correct answers reached through flawed reasoning and the bottleneck dimension should be surfaced.
- Reasoning Path Quality Classifier
- Weighted Aggregation Quality ScorerChoose when dimensions carry different importance, e.g., high-stakes medical reasoning weighting correctness over efficiency or consumer chatbots weighting relevancy and informativeness.
Response Scorer abstract
An abstract evaluation function that scores one agent response against ground truth, policy, or quality criteria and returns a normalized score or pass/fail indicator.
- Exact Match ScorerChoose when valid answers have a single canonical form; it fails on correct paraphrases.
- Fuzzy Match ScorerChoose when correct answers may be paraphrased (e.g., differently worded delivery dates) but wrong answers must still fail.
- Keyword Match ScorerChoose for cheap detection of required/forbidden language (e.g., acknowledgment phrases, unauthorized guarantees).
- LLM Judge abstractChoose when the quality dimension is qualitative and resists simple keyword or rule checks.
- AI-Feedback Preference LabelerChoose when preference labels must scale without human annotation time and be traceable to explicit principles; complement with human labels where contextual nuance matters.
- Chain-of-Thought JudgeChoose when evaluation accuracy, hallucination detection and user trust matter and secondary verification of the evaluator's logic is needed.
- Score-Only JudgeChoose only where lower evaluation accuracy and unverifiable judgments are acceptable; the chapter reports it underperforms explained judges.
- AI-Feedback Preference Labeler
- Rule-Based Compliance ScorerChoose when success depends on structured constraints such as refund limits, dates or eligibility.
- Semantic Similarity ScorerChoose when correct answers may be phrased differently from the ground truth and accuracy must be judged by meaning.
- Task Success Evaluator abstract
- Milestone EvaluatorChoose for complex tasks admitting partial credit, where locating the failing intermediate step guides optimization.
- State Outcome ScorerChoose for transactional domains where success has clear state-verifiable criteria (e.g., refund issued, inventory updated).
- Trajectory Matching EvaluatorChoose only when a task has a single valid execution path; the text warns it penalizes valid alternative routes of stochastic agents.
- Milestone Evaluator
- Tool Efficiency ScorerChoose when the agent exposes a tool-call log and operational efficiency matters.
- Trajectory ScorerChoose when deployment requires auditability, interpretability or trust calibration, not just correct outcomes.
Content Safety Filter abstract
An abstract filter that inspects model input or output text for harmful, toxic or policy-violating content and flags, blocks or passes it before delivery.
- Allow-List Output FilterChoose for highly constrained applications with well-defined, limited acceptable outputs (template-based bots, approved code sets, near-zero risk tolerance); avoid for general-purpose applications needing natural, flexible responses.
- Deny-List Content FilterChoose when microsecond latency, implementation simplicity and deterministic, auditable behaviour matter, e.g., as a fast first-pass check for blatant violations before more expensive ML classifiers; avoid as sole defense where context matters or adversaries use misspellings and character substitutions.
- Toxicity ClassifierChoose when adversaries consistently evade keyword filters or when context and semantics matter (coded language, microaggressions), accepting labeled-data needs, millisecond latency and residual adversarial susceptibility.
Domain Compliance Rail abstract
An abstract output guardrail that checks a proposed response against domain-specific regulatory rules and, on violation, halts delivery and returns a standardized refusal with a logged reason.
- Pattern Compliance CheckerChoose as a fast, deterministic check for explicit regulated phrasing (e.g., 'you should buy', price predictions); insufficient alone because paraphrased advice evades surface patterns.
- Semantic Compliance ClassifierChoose when adversaries or the model paraphrase regulated content, so advice-giving intent must be detected regardless of phrasing, accepting higher false positives and compute cost.
Execution Sandbox abstract
An isolated execution environment for running untrusted, agent-generated code without exposing host systems.
- Dedicated Hardware SandboxChoose when threat modelling demands the strongest possible isolation per customer and its cost and operational complexity are acceptable.
- MicroVM SandboxChoose for adversarial multi-tenant, batch or high-value workloads (e.g., financial services handling customer funds) where security outweighs VM boot latency and memory overhead; also the fallback for gVisor-incompatible workloads.
- Shared-Kernel Container SandboxChoose for low-risk development environments running trusted code; insufficient for multi-tenant, sensitive-data or autonomous production workloads because of shared-kernel escape risk.
- Syscall-Interception SandboxChoose when stronger isolation than standard containers is needed with modest overhead, e.g., interactive real-time services that cannot tolerate VM boot latency, provided the workload's syscalls are supported.
Jailbreak Detector abstract
A detection component that identifies jailbreak and prompt-injection attempts in user input, such as role-play manipulations or known attack preambles.
- Classifier Jailbreak DetectorChoose for deeper analysis in high-stakes applications (financial services, healthcare, government).
- Heuristic Jailbreak DetectorChoose for fast initial screening of all traffic; effective against documented attacks but not novel ones.
Moderation Inference Service abstract
An abstract service that executes content-classification inference for moderation filters and returns category scores.
- Self-Hosted Moderation ServiceChoose when lower latency is needed and content must not be exposed to third parties (critical for HIPAA or GDPR), and ML infrastructure expertise is available.
- Third-Party Moderation ServiceChoose when ease of integration outweighs dependency on an external service and per-request pricing.
Output Bias Detector abstract
A runtime post-filter that scores each response for biased or discriminatory language, including dog whistles, stereotyping and microaggressions, and blocks it or flags uncertain cases for human review.
- Classifier Bias DetectorChoose for robust, scalable detection on user-facing content, including implicit bias without explicit markers; requires labeled training examples and periodic retraining.
- LLM-Judge Bias DetectorChoose for high-stakes decisions needing maximum detection accuracy, accepting a full LLM inference of latency and per-evaluation cost.
- Rule-Based Bias DetectorChoose for fast, interpretable, immediate filtering of explicit bias indicators; insufficient alone because paraphrasing evades it and it misses contextual bias.
PII Detector abstract
An abstract privacy component that locates personally identifiable information spans in document text and reports each with a PII type, detection method and confidence score.
- Context-Aware PII DetectorChoose for identifiers that are PII only in context (patient IDs, account numbers, internal IDs next to field labels).
- Multi-Strategy PII DetectorChoose when documents mix structured, named and contextual PII (e.g., regulated healthcare content); avoids single-strategy blind spots.
- NER PII DetectorChoose for unstructured PII such as person names, organizations and locations that regex cannot distinguish from normal text; costlier than patterns.
- Pattern PII DetectorChoose for structured PII with predictable formats (SSNs, credit cards, phone numbers, emails); cheapest strategy but misses names and contextual identifiers.
Policy Enforcement Mode Configuration abstract
An abstract configuration setting, per policy, whether the policy engine only logs evaluation outcomes, blocks violations, or blocks and launches automated remediation.
- Full Enforcement ModeChoose once policies have been calibrated in monitor and soft modes and their thresholds reflect organizational reality.
- Monitor-Only Enforcement ModeChoose when first deploying policies, to observe agent behaviour, detect over-triggering or never-triggering policies, and calibrate thresholds before enforcement.
- Soft Enforcement ModeChoose after baseline calibration when severe-consequence, low-error-tolerance policies (production database deletion, SSN exfiltration, exceeding authority limits) must be enforced while other policies still need tuning.
Tool Hallucination Detector abstract
An abstract detector that estimates, from model confidence signals, whether a proposed tool call or parameter is likely hallucinated, flagging low-confidence calls for verification.
- Entropy-Based Hallucination DetectorChoose when the generating model's token probability distributions are observable during parameter generation, so uncertainty can be measured without a separate model.
- Verifier-Model Hallucination DetectorChoose when labeled examples of correct and hallucinated tool calls are available to train a well-calibrated verifier, and a second model check per call is acceptable.
Violation Response Policy abstract
An abstract configuration stating what the system does when an output is judged to violate, or plausibly violate, a principle.
- Blocking Violation Response PolicyChoose for clearly prohibited categories (e.g., illegal activity) where erring on the side of safety outweighs refusing some acceptable requests.
- Content-Modification Violation Response PolicyChoose when the violating portion can be removed or rewritten so the remaining response stays useful.
- Human-Escalation Violation Response PolicyChoose for ambiguous cases with possible legitimate justification (fiction, security research) or high-stakes dilemmas, accepting slower responses for greater accuracy.
- Monitor-Only Violation Response PolicyChoose for borderline outputs approaching policy boundaries where immediate blocking is not required; use blocking mode for inviolable legal, ethical or safety constraints.
- Warning-Label Violation Response PolicyChoose when preserving user autonomy matters and the concern (e.g., possible misinformation, low confidence) warrants flagging rather than blocking.
AI Auditor abstract
An abstract independent reviewer role that examines AI governance, technical, operational, safety and regulatory evidence to determine conformance and reports graded findings.
- External AuditorChoose when external validation is needed for customers, regulators or certification (e.g., ISO/IEC 42001), typically annually with a formal report.
- Internal AuditorChoose for continuous assurance: monthly or quarterly comprehensive audits with immediate finding reporting and verification that controls operate as designed.
Approval Timeout Fallback Policy abstract
An abstract policy setting the automatic outcome when no approver in the escalation chain responds before the final deadline.
- Auto-Approve Timeout FallbackChoose when actions are reversible and organizational risk tolerance favors velocity over oversight (aggressive).
- Auto-Reject Timeout FallbackChoose when actions are irreversible or risk tolerance is low (conservative; prevents unauthorized actions).
- Escalate-Up Timeout FallbackChoose when the decision must still be made promptly and higher authority is available to make it.
- Post-Hoc Review Timeout FallbackChoose for time-critical decisions whose delay harms users (e.g., emergency claims or treatments) when approval queues exceed their SLA.
- Retry-Queue Timeout FallbackChoose when the request remains valid later and neither declining nor escalating further is appropriate.
Bias Mitigator abstract
An abstract fairness component that corrects detected bias in a decision model by intervening on its training data, its training objective, or its output decisions.
- Decision Threshold AdjusterChoose (post-processing) to correct a trained model's decisions without retraining, e.g., as an adaptive intervention on production drift; post-hoc correction usually gives weaker fairness-accuracy trade-offs than in-training constraints.
- Fairness-Constrained TrainerChoose (in-processing) when fairness should be a first-class optimization goal; often yields better fairness-accuracy trade-offs than post-hoc correction because the model learns representations that satisfy the constraint.
- Representation RebalancerChoose (pre-processing) when training data underrepresents groups or encodes historical imbalance; reweighting preserves the original distribution, while resampling/augmentation change the dataset.
Constitution abstract
A versioned, human-readable set of explicit natural-language principles (e.g., helpful, harmless, honest; domain rules) that guides model training and runtime behaviour and can be publicly inspected and debated.
- Regional ConstitutionChoose for globally deployed systems whose jurisdictions differ in law and cultural norms (e.g., political speech, religious expression), accepting higher design, evaluation and maintenance cost.
- Unified ConstitutionChoose when one consistent value set can serve all users, explicitly acknowledging limited scope (particular value commitments) rather than claiming universal validity.
- User-Configurable ConstitutionChoose when users legitimately differ in how they weight values such as privacy versus convenience, and customization can be bounded by non-negotiable principles.
Lawful Basis Record abstract
An abstract documented justification establishing which of GDPR's six lawful bases authorizes a specific personal-data processing purpose.
- Consent BasisChoose when processing is genuinely optional and not necessary for the service (e.g., marketing, optional research participation, analytics cookies); avoid when withdrawal would disrupt ongoing operations such as longitudinal research.
- Contract BasisChoose when processing is genuinely necessary, not merely convenient, for contract performance (e.g., shipping address and payment for a purchase, employee bank details, treatment as healthcare service delivery).
- Legal Obligation BasisChoose when a specific statute mandates the processing (e.g., tax record retention, workplace safety documentation, medical record retention); vague references to 'legal compliance' are insufficient.
- Legitimate Interests BasisChoose when an organizational or third-party interest (e.g., fraud prevention, medical research with safeguards, behavioural analytics) does not override individual rights and consent would be operationally fragile.
- Public Task BasisChoose for governmental functions, public health initiatives or publicly funded research exercising official authority, identifying the specific public interest or legal provision.
- Vital Interests BasisChoose only narrowly, where failure to process could result in serious harm or death (e.g., treating an unconscious patient, child protection).
Privacy-Preserving Fairness Auditor abstract
An abstract fairness-audit component that computes group fairness metrics while preventing identification of individuals' demographic data or outcomes.
- Differentially Private Fairness AuditorChoose when fairness metrics must be released with provable guarantees that no individual's participation or outcome can be inferred, and sample sizes are large enough to tolerate added noise.
- Federated Fairness AuditorChoose when several institutions (hospital networks, bank consortia) individually lack the data volume or diversity for statistically powerful bias detection but cannot centralize raw data.
- Secure Multiparty Fairness AuditorChoose when parties must jointly compute aggregate fairness metrics while no party may learn any other party's demographic data or predictions, accepting cryptographic overhead.
Risk Scorer abstract
An abstract governance component that assigns each identified AI risk a level by combining its estimated likelihood and potential impact, enabling prioritization.
- Qualitative Risk Matrix ScorerChoose when likelihood and impact can only be judged in bands and qualitative judgment about context and acceptability matters more than numerical precision.
- Quantitative Risk ScorerChoose when numerical evidence is needed to track trends, compare mitigation alternatives and trigger action at defined thresholds, e.g., continuous re-scoring.
Agent User Interface abstract
An abstract user-facing interaction surface through which humans direct agents and perceive agent status, reasoning, uncertainty, evidence, and results.
- Command Palette with Agent SuggestionsChoose for IDEs, productivity and design tools where power users value keyboard speed and the agent can predict intent from context; avoid when discoverability for new users, multi-turn dialogue, or extensive explanation is required.
- Conversational (Chat) InterfaceChoose for general Q&A, customer support, simple task automation, and exploratory conversations; avoid for workflows needing visual process representation, dense structured data, or parallel information streams.
Automatic Speech Recognizer abstract
An abstract speech-to-text service that converts spoken audio into text transcripts for consumption by an agent or downstream processing.
- Speech TranscriberChoose for post-hoc analysis (call-center QA, meeting summaries, voicemail), maximum accuracy (legal, medical), speaker diarization or overlapping multi-speaker audio, where higher latency is acceptable.
- Streaming Speech RecognizerChoose when real-time interaction is required (live customer-service calls, voice assistants, live accessibility captioning) and latency matters more than perfect accuracy.
Explanation Expansion Controller abstract
An abstract presentation-control component that decides when collapsed explanation layers are expanded to reveal deeper reasoning, uncertainty or data detail.
- On-Demand Explanation ExpanderChoose when users can be expected to seek detail themselves (low time pressure, routine or low-stakes decisions); insufficient alone because users rarely seek additional information spontaneously.
- Proactive Explanation ExpanderChoose when confidence may fall below thresholds, decisions may deviate from historical patterns, or high-stakes scenarios arise, so critical detail must surface without user initiative.
Layered Explanation View abstract
An abstract presentation tier that renders one agent decision at a disclosure depth matched to a stakeholder persona's expertise, information needs and decision stakes.
- Exhaustive Audit Explanation ViewChoose for expert analysts and compliance or regulatory auditors conducting deep audits or fairness reviews of high-stakes decisions.
- Intermediate Explanation ViewChoose for regular users with moderate expertise (loan officers, credit analysts, moderators) and medium-stakes decisions: key factors, feature weights, confidence, comparisons and governing thresholds.
- Summary Explanation ViewChoose for novice users, customers and affected parties, and for routine low-stakes decisions: plain-language summary of primary factors with actionable improvement or recourse guidance.
Streaming Transport abstract
An abstract client-facing transport through which incremental agent output is delivered from server to client while it is still being generated.
- Server-Sent Events StreamChoose for most agent streaming workflows: the user submits a query, the agent streams a response, and no mid-stream client-to-server interaction is needed.
- WebSocket ChannelChoose for genuinely interactive scenarios: agents expecting mid-stream user feedback or interruption, collaborative multi-user agent sessions, or real-time dashboards whose pushed updates need client acknowledgment.
Annotation Task Format abstract
An abstract configuration specifying how human judgments are elicited for each prompt, such as absolute ratings, pairwise comparisons or multi-way rankings of candidate responses.
- Absolute Rating FormatUsed in early RLHF work; generally avoid for preference learning because annotators interpret scales inconsistently across people and over time.
- Multi-Way Comparison FormatChoose when more information per annotation is worth higher annotator cognitive load and potentially lower judgment consistency.
- Pairwise Comparison FormatChoose as the default (gold standard) elicitation format: comparative judgments are more reliable than absolute ratings and scale well.
Escalation Protocol abstract
An abstract specification of what the system does when an agent approaches or crosses a decision boundary, naming notified roles, decision owner, response-time objective and fallback behaviour.
- Approval Escalation ProtocolChoose when the decision carries significant consequences but does not require immediate action.
- Emergency Escalation ProtocolChoose when agents detect high-severity situations (security incidents, system failures affecting critical operations) requiring urgent human attention.
- Notification Escalation ProtocolChoose when stakeholders need awareness of a boundary event but the action may proceed (e.g., internal transfers above $10,000 with enhanced logging).
Feedback Collector abstract
A component that captures per-response user ratings and structured reasons, such as thumbs up/down with follow-up questions, for immediate conversation repair and longer-term improvement.
- Binary Rating Feedback PromptChoose when high feedback volume matters more than information density; single-click ratings generate high volume but low information per response.
- Deferred Feedback SurveyChoose when periodic overall satisfaction, frequency and feature-request data are needed (Ref10.03), accepting lower response rates as context fades.
- Failure-Triggered Feedback PromptChoose when training data should focus precisely on current weaknesses rather than random optional feedback.
- Structured Correction Feedback FormChoose when rich training signals are needed and lower participation due to higher user effort is acceptable.
Human Validation Checkpoint abstract
An abstract oversight checkpoint at which humans validate agent actions, either before execution (approval-before) or after execution on completed actions (approval-after).
- Approval GatewayChoose for high-consequence irreversible decisions (financial transactions, contractual commitments, changes affecting many users) where errors must be prevented before impact, accepting reduced throughput.
- Post-Action Review SamplerChoose for lower-stakes reversible actions where operational velocity matters; requires the ability to reverse completed actions when reviews reveal errors.
Oversight Gate abstract
An abstract decision gate that evaluates each proposed agent action and selects the human-control pattern it requires: automatic execution, notification, approval, or monitoring.
- Confidence GateChoose when decisions carry a usable confidence score and escalation should depend on agent certainty; the text presents risk gating as complementary for high-impact actions.
- Escalation Agent
- Exception GateChoose when certain case characteristics (unusual amount, rare scenario, conflicting signals, policy exception) warrant specialist judgement regardless of agent confidence or aggregate risk score.
- Risk GateChoose when escalation must reflect action impact (financial amount, deletion, external calls, customer reach) regardless of agent confidence.
Preference Annotator abstract
A human evaluator who compares pairs of alternative agent responses and indicates which better satisfies specified criteria such as helpfulness, tone, accuracy or policy compliance.
- Domain Expert AnnotatorChoose for specialized domains (healthcare, finance, autonomous vehicles, legal) where judging clinical accuracy, regulatory compliance or safety requires expertise; costs more than general annotators.
- General Preference AnnotatorChoose for general-purpose content where volume and cost matter; unsuitable for specialized medical, legal, financial or safety-critical content.