Observability & Evaluation · Software component

Feature Activation Monitor

Software componentObservability & EvaluationObservability & Evaluationarc:FeatureActivationMonitor

An interpretability component that decomposes model activations into sparse interpretable features during inference and detects activation of features associated with reasoning errors.

Responsibility. Detects error-associated interpretable features in model activations.

Also known as: Feature activation monitoring

monitorstriggersis configured byhostsFoundation LLM: monitorsFoundation LLMFeature Steering Controller: triggersFeature Steering Control…Interpreted Feature Library: is configured byInterpreted Feature Libr…Sparse Autoencoder: hostsSparse Autoencoder
Direct neighbourhood (hover for relationship types)

Relationships

hosts structural

is configured by structural

triggers dynamic

monitors assurance

Classification

Patterns
Sparse autoencodersDictionary learning
Quality attributes
Maintainability (ISO/IEC 25010)Explainability (NIST AI RMF: explainable and interpretable)
Risks mitigated
Confirmation biasPattern matching without causal reasoning

Sources

  1. Ch3.6: T. Nguyen, "Trace Analysis and Execution Debugging," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.6. ISBN: 9798244538229.