Observability & Evaluation · Data artifact
Regression Test Suite
Data artifactObservability & EvaluationObservability & Evaluationarc:RegressionTestSuite
A versioned set of known-good and known-bad agent tasks, including core, edge-case and historically failed tasks, re-run after every model or agent update to detect performance regressions.
Responsibility. Holds the tasks whose results must not regress across agent versions.
Also known as: Known good/bad trace dataset, Regression test dataset
Relationships
is read by dependency
is written by dependency
Design guidance
- SHOULD separate core-functionality, edge-case and historical-failure categories, each with its own pass threshold.
Quantitative guidance
As stated by the sources; verify before use.
- Core tasks success >99%; edge cases >90%; all previously fixed issues must remain fixed (Ref8.03).
Classification
- Patterns
- Regression testingEvaluation flywheel
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Reintroduction of previously fixed failures after model updates
Sources
- Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note