Observability & Evaluation · Software component
Evaluation Failure Analyzer
Software componentObservability & EvaluationObservability & Evaluationarc:EvaluationFailureAnalyzer
An analysis component that categorizes failed evaluation or production interactions into failure modes and traces them to responsible agent components.
Responsibility. Attributes observed failures to failure categories and root-cause components.
Also known as: Failure pattern analysis, Root cause analysis, Error clustering analysis, Stratified failure analysis, Error propagation analysis, Error Categorization, Error Analysis, Impact Prioritizer, Failure attribution, Detailed failure analysis, Hallucination feedback loop, Trace pattern mining
Relationships
is configured by structural
reads dependency
receives data from dynamic
- Evaluation Harness Ch3.3 Ch3.4 +1
- LLM Judge abstract Ch3.10
sends data to dynamic
triggers dynamic
evaluates assurance
Design guidance
- SHOULD distinguish surface issues fixable by instruction changes from capability gaps requiring architectural change.
- SHOULD stratify performance by task type and complexity rather than rely on aggregate success rate.
- SHOULD compare runs with full retrieval provided versus self-retrieval to attribute failures to reasoning or retrieval.
- SHOULD group failures by root cause rather than surface symptom.
- SHOULD distinguish retrieval errors from genuine knowledge gaps, as they need different remedies.
- SHOULD recompute impact scores after every validated fix.
- SHOULD compare failure logs of baseline and ablated conditions to estimate a component's genuine contribution net of cascading effects.
- SHOULD analyse which specific cases fail under ablation, not only aggregate metrics.
- SHOULD analyse at least 100-500 outputs before implementing mitigations, distinguishing systematic patterns from isolated incidents.
- SHOULD route findings to knowledge base updates, retrieval refinements or prompt iterations according to root cause.
- SHOULD mine exported traces of both successful and failed workflows to cluster failures by signature before fixing.
Quantitative guidance
As stated by the sources; verify before use.
- Graceful degradation example 90/70/50% on simple/medium/hard versus cliff-edge 90/90/20% (Ch3.3).
- Example distribution of 130 failures: parameter 32%, retrieval 25%, tool selection 15%, timeout 12%, reasoning 8%, hallucination 8% (Ch3.4).
- Impact scores: parameter 1,280; tool selection 630; retrieval 500; timeout 216; reasoning 168; hallucination 144 (Ch3.4).
- Illustrative: a no-tools ablation with 30% failures (20% reasoning, 10% format) versus a 5% baseline implies the tool's genuine contribution is ~10% rather than the observed 25% (Ch3.7).
- Root cause analysis yields 2-3x higher hallucination reduction than generic mitigations (Ch3.10).
- 1,000 exported traces (920 successful, 80 failed) yielded deadlock 35%, out-of-order execution 40%, state consistency 25% (Ch8.2A).
Classification
- Patterns
- Failure taxonomy: reasoning, decision-making, instruction-following, planningEnvironment-specific error analysisOracle-retrieval ablation to isolate reasoning failuresSuccess rate by complexity bucketCross-website consistency analysisFailure taxonomy classificationRoot cause analysisImpact scoring (Frequency x Severity x Effort)Iterative re-prioritisation after each fixPer-case failure comparison between baseline and ablated conditionsDetection-to-improvement feedback loop
- Quality attributes
- Maintainability (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Aggregate metrics masking task-specific capability gapsOptimising by intuition instead of dataOver-investing in hard, low-impact failure classesCascading failures inflating apparent component importanceAggregate metrics hiding rare edge-case failuresOverreaction to salient isolated incidentsGeneric untargeted mitigations
Sources
- Ch3.2: T. Nguyen, "Compare Agent Performance Across Tasks and Datasets," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.2. ISBN: 9798244538229.
- Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch3.7: T. Nguyen, "Tool Usage Auditing," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.7. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.
- Ch8.2A: T. Nguyen, "Error Rates and Reliability," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.2A. ISBN: 9798244538229.
- Ref3.10: "Powering the Next Generation of AI Agents," unpublished reference note (10-Powering-Next-Generation-AI-Agents.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.07: "Agent Health Checks and Diagnostics," unpublished reference note (07-Agent-Health-Checks-Diagnostics.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note