Observability & Evaluation · Software component
Metrics Collector
Software componentObservability & EvaluationObservability & Evaluationarc:MetricsCollector
A telemetry component aggregating per-agent latency, error, fallback, token-consumption and confidence metrics.
Responsibility. Aggregates per-agent performance metrics.
Also known as: Per-agent metrics, Prometheus, Action Metrics Aggregator, Reasoning quality metric tracking, Online efficiency monitoring, Production instrumentation, Metrics scraper, Time-series metrics server, Continuous lightweight performance monitoring, Fleet-wide metrics, GPU metrics scraper, Search diagnostics logging, RAG metrics (performance, quality, cost, reliability), Prometheus server, Latency percentile tracker, Prometheus scraper, Event Monitoring Layer, Agent-specific, interaction and system-level metrics
Relationships
deployed on structural
is configured by structural
invokes dependency
is invoked by dependency
reads dependency
writes dependency
receives data from dynamic
receives telemetry from dynamic
- A/B Test Traffic Splitter Ref8.03
- Agent Controller abstract Ch1.8 Ch4.2 +2
- Agent Version Experimenter abstract Ch3.3
- Autoscaler abstract Ch4.7
- Circuit Breaker Ch2.8
- Coherence Continuity Scorer Ch3.9
- Dialogue Flow Manager Ch10.1
- Document Quality Filter Ch6.3A
- Edge Device Agent Ch4.6
- Escalation Handler Ch8.4
- Execution Profiler Ref3.05
- Failure Category Classifier Ch8.2B
- GPU Telemetry Exporter Ch4.4 Ch4.5 +4
- Graceful Degradation Manager Ch2.8
- Guardrail Orchestrator Ref7.03 Ref9.09
- Inference Server Ch2.7 Ch4.5
- Ingestion Pipeline Orchestrator Ch6.3A Ch6.3B
- Knowledge Graph Store abstract Ch1.7B
- LLM Inference Service Ch1.3 Ch1.8
- LLM Judge abstract Ch3.9
- MCTS Planner Ch5.5
- Model Router Ch2.8 Ch4.2
- Parallel Sub-Query Retrieval Controller Ch6.6
- Plugin Kernel Orchestrator Ch2.1
- RAG Query Orchestrator Ch6.5
- Reasoning Quality Scorer abstract Ch3.9
- Response Streamer Ch2.9
- Self-Managed Vector Index Store Ch6.2B
- Service Container Ch2.5
- Speculative Decoder Ch7.1A
- Supervisor Agent Ch1.3
- Task Success Evaluator abstract Ch8.4
- Text Embedding Service Ch6.3B
- Token Cost Meter Ch3.10 Ch4.1 +1
- Tool Call Accuracy Evaluator Ch3.8
- Tool Executor Ch3.10
- Trace Collector Ch3.6
- Trajectory Matching Evaluator Ch3.8
- Vector Batch Ingestor Ch6.3B
- Worker Agent abstract Ch1.3
sends data to dynamic
- Alert Manager Ch1.8 Ch3.8 +5
- Continuous Compliance Monitor Ref9.09
- Custom Metrics Adapter Ch4.5 Ref4.03
- Evaluation Result Analyzer Ch3.9
- Experiment Guardrail Monitor Ch3.1A Ch4.2
- Metrics Dashboard abstract Ch4.5
- Quality Drift Detector Ch3.9 Ch3.10
- SLO Monitor Ch1.7B Ch4.2
- Service Degradation Predictor Ch10.2
- Agent Behavior Anomaly Detector Ch3.3
monitors assurance
produces lifecycle
Design guidance
- SHOULD report p95/p99 latency per agent, error and fallback rates, token consumption, and confidence scores over time.
- SHOULD track latency percentiles (p50/p95/p99), failure rates, graph growth, index cache hit rates, and memory usage.
- SHOULD capture per-request token consumption, API call frequency and latency, tool execution times, and error and retry patterns from live interactions without overwhelming log volume.
- SHOULD pull metrics from service endpoints so services need not know where metrics go and monitoring outages do not back up producers.
- SHOULD capture LLM-specific metrics (time to first token, token usage, cost per interaction, active sessions, conversation turns, end-to-end STT-LLM-TTS latency).
- SHOULD aggregate metrics by service name rather than by container instance.
- SHOULD collect GPU utilisation, throughput and error rates, latency percentiles with queue-wait vs compute split, batch sizes achieved, cache hit rates and concurrent request counts for every served model.
- SHOULD surface both instance-level and fleet-level metrics (load balance across replicas, per-replica error rates, aggregate capacity vs SLO).
- SHOULD retain metrics for trend analysis and capacity planning.
- SHOULD scrape vector store metrics every 15-30 s and track capacity, performance (p50/p95/p99, ingestion rate, cache hit rate) and health metrics every 1-5 min.
- SHOULD track guardrail input rejection rate, output filtering rate, false positive rate and response latency (Ref7.03).
- SHOULD pull metrics from inference /metrics endpoints on a fixed interval and store them as time series for querying and alerting.
- SHOULD track average batch size, queue depth, request-latency distribution, GPU utilization and model versus queue latency for batched inference.
- SHOULD track P50/P95/P99 latency, tokens/s, time to first token, GPU/memory/bandwidth utilization and cost per token/request.
- MUST compute latency percentiles (P50, P95, P99) from the full distribution rather than reporting averages alone.
- SHOULD track end-to-end latency, time to first token (TTFT) and token generation rate together, prioritising by interface type (streaming chat vs batch processing).
- SHOULD collect operational (availability, latency, error rate, queue depth), financial (cost per request, per successful request, by agent, by model) and quality (accuracy, satisfaction, goal achievement, tool call success) metrics (Ref8.01).
- SHOULD maintain separate counters for safety violations and infrastructure errors, labelled by violation or error type.
- SHOULD collect agent-specific metrics (decision and confidence distributions, error rate per decision type), interaction metrics (communication patterns, resource use, queuing) and system-level metrics (throughput, cascading failures, end-to-end latency).
Quantitative guidance
As stated by the sources; verify before use.
- API call frequency = total external API calls / tasks; loan agent reduced from 12 to 4.8 calls per application (Ch3.10).
- Scrape intervals of 15-60 s trade storage against metric granularity (Ch4.1).
- GPU scrape job example interval 1 s; minimum Prometheus sizing 2 CPU / 4 GB RAM, 50 GB+ storage for 1-week retention; example 30-day retention on 100 Gi (Ref4.05).
- Scrapes NIM endpoints every 15-30 seconds (Ch7.2).
- Two agents with equal 600 ms averages: A P50 550 / P95 750 / P99 900 ms vs B P50 400 / P95 2,000 / P99 8,000 ms (Ch8.1).
- A 500 ms average can coexist with P99 above 10-12 s; at 100,000 queries/day the 1% tail is 1,000 users/day, estimated ~$21,000/week (~$1.09 M/yr) in follow-up support (Ch8.1).
- Scrapers typically collect at 15-30 s intervals (Ch8.1).
- Example: an agent with 98% overall accuracy may show 40% error rates on rare decision categories; approval rate shifting from 70% to 90% signals behavioural drift (Ch10.4).
Classification
- Patterns
- Pull-based scrapingCounters, gauges, histograms, summariesService discovery of exportersTime-series retention for capacity planning
- Technologies
- Neo4j monitoring APIPrometheusGrafanaPhoenixWeaveLangfuseOpenTelemetryprometheus_clientPrometheus Operator (ServiceMonitor, PrometheusRule)kube-prometheus-stackDatadogStatsDNew Relic
- Quality attributes
- Maintainability (ISO/IEC 25010)Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Undetected bottlenecksSilent quality degradation
Sources
- Ch1.3: T. Nguyen, "Multi-Agent Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.3. ISBN: 9798244538229.
- Ch1.7B: T. Nguyen, "Relational Reasoning with Knowledge Graphs - Hybrid RAG+KG Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.7B. ISBN: 9798244538229.
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch2.5: T. Nguyen, "Semantic Kernel," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.5. ISBN: 9798244538229.
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch2.9: T. Nguyen, "Streaming and Real-Time Responses," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.9. ISBN: 9798244538229.
- Ch3.1A: T. Nguyen, "Implement Evaluation Pipelines and Task Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.1A. ISBN: 9798244538229.
- Ch3.6: T. Nguyen, "Trace Analysis and Execution Debugging," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.6. ISBN: 9798244538229.
- Ch3.8: T. Nguyen, "Action Accuracy Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.8. ISBN: 9798244538229.
- Ch3.9: T. Nguyen, "Reasoning Quality," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.9. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch4.1: T. Nguyen, "Introduction to AI Agent Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.1. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch5.5: T. Nguyen, "Monte Carlo Tree Search Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.5. ISBN: 9798244538229.
- Ch6.2B: T. Nguyen, "Production Vector Database Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2B. ISBN: 9798244538229.
- Ch6.3A: T. Nguyen, "ETL Pipeline Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.3A. ISBN: 9798244538229.
- Ch6.3B: T. Nguyen, "ETL Worked Example - Load Phase & Pipeline Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.3B. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ch6.6: T. Nguyen, "Query Decomposition and Adaptive Retrieval," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.6. ISBN: 9798244538229.
- Ch7.1A: T. Nguyen, "Advanced Implementation with Nvidia NEMO Framework and Nvlink," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1A. ISBN: 9798244538229.
- Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
- Ch8.1: T. Nguyen, "Latency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.1. ISBN: 9798244538229.
- Ch8.2B: T. Nguyen, "NeMo Guardrails Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.2B. ISBN: 9798244538229.
- Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.
- Ch10.2: T. Nguyen, "Proactive Agents," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.2. ISBN: 9798244538229.
- Ch10.4: T. Nguyen, "Human-in-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.4. ISBN: 9798244538229.
- Ref3.05: NVIDIA, "NVIDIA NeMo Agent Toolkit FAQs," NVIDIA NeMo Agent Toolkit Documentation, v1.8. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/agent-toolkit/latest/resources/faq.html
- Ref4.03: M. Zhang, J. Wyman, I. M. Bhosale, and W. Tan, "Scaling LLMs with NVIDIA Triton and NVIDIA TensorRT-LLM using Kubernetes," NVIDIA Technical Blog, Oct. 22, 2024. [Online]. Available: https://developer.nvidia.com/blog/scaling-llms-with-nvidia-triton-and-nvidia-tensorrt-llm-using-kubernetes/
- Ref4.05: NVIDIA, "Setting up Prometheus," NVIDIA GPU Telemetry Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-telemetry/latest/kube-prometheus.html
- Ref4.08: NVIDIA, "Agentic AI in the Factory," NVIDIA Enterprise AI Factory Design Guide White Paper. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/ai-enterprise/planning-resource/ai-factory-white-paper/latest/agentic-ai-in-the-factory.html
- Ref7.02: NVIDIA, "Batchers," NVIDIA Triton Inference Server User Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/batcher.html
- Ref7.03: NVIDIA, "Overview," NVIDIA NeMo Guardrails Library Developer Guide. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/guardrails/about-nemo-guardrails-library/overview
- Ref7.04: NVIDIA, "NVIDIA NIM," NVIDIA Docs. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nim/
- Ref7.15: "Advanced Agentic AI Optimization Techniques," unpublished reference note (15-Advanced-Agentic-Optimization.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.16: "Production Monitoring and Operations for Agentic AI," unpublished reference note (16-Production-Monitoring-Operations.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref7.18: "Chapter 7 Summary: NVIDIA Platform Implementation," unpublished reference note (18-Chapter-7-Summary.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.01: LangChain, "LangSmith observability: AI agent observability platform," LangChain. Accessed: Sep. 27, 2026. [Online]. Available: https://www.langchain.com/langsmith/observability
- Ref9.09: "Compliance Automation and Tools," unpublished reference note (09-Compliance-Automation-Tools.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note