Observability & Evaluation · Model asset

Sparse Autoencoder

Model assetObservability & EvaluationObservability & Evaluationarc:SparseAutoencoder

A model trained to decompose a language model's activations into sparse combinations of interpretable features.

Responsibility. Maps raw activations to sparse interpretable features.

Also known as: SAE

deployed onFeature Activation Monitor: deployed onFeature Activation Monitor
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

Classification

Patterns
Dictionary learning

Sources

  1. Ch3.6: T. Nguyen, "Trace Analysis and Execution Debugging," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.6. ISBN: 9798244538229.