Observability & Evaluation · Data artifact
Evaluation Dataset
Data artifactObservability & EvaluationObservability & Evaluationarc:EvaluationDataset
A curated, versioned collection of test queries paired with ground-truth answers and metadata (task type, difficulty, expected reasoning) used as a reproducible benchmark for agent quality.
Responsibility. Provides a fixed, labeled benchmark against which every agent version is measured.
Also known as: Test dataset, Ground-truth dataset, Golden dataset, Regression test suite, Benchmark dataset, Benchmark suite, Test set, Smoke test set, Test Set, Tuning Set, Benchmark Suite, Action Accuracy Test Set, Offline test set, Held-out test set, Curated benchmark dataset, Fixed evaluation dataset, Stratified evaluation set, Hallucination benchmark, Adversarial test set, GSM8K, SVAMP, ARC, TriviaQA, MMLU, HumanEval, Development dataset with known solutions, Relevance judgment set, Labeled relevance query set, Golden test set, RAG test set, Ground truth test cases, Query test set
Relationships
is read by dependency
- Confidence Calibration Analyzer Ch3.10 Ch5.2
- Evaluation Harness Ch3.1A Ch3.1B +15
- Helpfulness Preservation Evaluator Ch9.5
- Keyword Match Scorer Ref7.07
- RAG Evaluator Ch6.4 Ch6.5
- Reasoning Chain Validator Ch3.3
- Reasoning Faithfulness Tester Ch5.1
- Response Scorer abstract Ch3.1A
- Retrieval Quality Evaluator Ch5.8 Ch6.2A +1
- Task Success Evaluator abstract Ch3.3
- Test Case Mutator Ch3.8
is written by dependency
receives data from dynamic
is produced by lifecycle
Design guidance
- SHOULD represent the production query distribution, including common cases, rare edge cases and adversarial examples.
- SHOULD add a test case for every fixed bug so the regression cannot silently return.
- SHOULD be extended with production queries where offline and online results diverged, to close the prediction gap.
- SHOULD mirror production proportions of routine, moderately complex and edge-case queries rather than academic difficulty.
- SHOULD capture authentic failure modes reported by users (realistic typos, rambling multi-part questions) rather than sanitized approximations.
- SHOULD include edge cases, adversarial scenarios and anti-automation defences, not only common cases.
- SHOULD be stratified by task complexity so degradation patterns are visible.
- MUST be versioned, recording when cases were added, modified or retired and the production phenomena that motivated each change.
- SHOULD include multi-turn scenarios measuring state management and context retention.
- SHOULD cover simple retrieval, multi-step resolution, ambiguous edge cases and backend-integration scenarios.
- SHOULD include out-of-distribution slices (later time periods, adversarial phrasing) to probe robustness.
- MUST include failure, empty-result, error, timeout and partial-success tool scenarios, not only happy paths.
- SHOULD systematically cover boundary, invalid, Unicode, locale and temporal edge values for parameters.
- MUST maintain strict train-test separation to prevent overfitting.
- SHOULD mirror production query distributions with a balanced mix of common cases, edge cases and adversarial examples.
- SHOULD capture realistic user journeys, including complex multi-turn questions, rather than artificial test scenarios.
- SHOULD be extended with insights from production failures to prevent recurrence.
- SHOULD contain 100+ samples per experimental condition for adequate statistical power.
- MUST be held constant across baseline and all ablated conditions.
- SHOULD include easy and difficult cases and be stratified by task type so category-specific component importance is visible.
- MUST include edge cases, adversarial examples (impossible questions, contradictory constraints) and production-sampled queries.
- SHOULD include error-condition scenarios (retrieval failures, API timeouts, context truncation, partial outages).
- SHOULD be built from sampled production traces and representative tasks, with domain-expert ground truth, stratified by difficulty and including negative cases.
Quantitative guidance
As stated by the sources; verify before use.
- Offline predicted 92% accuracy vs 87.8% online TSR, a 4.2-point gap (Ch3.1A).
- A 1,000-query set may still miss a query category that is 5% of production traffic (Ch3.1A).
- Example production mix to mirror: 70% routine, 25% moderately complex, 5% edge cases (Ch3.2).
- In-distribution vs out-of-distribution gaps often exceed 30 points (e.g., 85% -> 55%; HotpotQA 95% vs 2WikiMultiHopQA 60%) (Ch3.2).
- Healthcare example added 150 feedback-derived regression cases (50 paraphrase-consistency, 50 lay-language, 50 brand-name) (Ch3.2).
- 47% of Mind2Web tasks became invalid or had outdated ground truth within 18 months (Ch3.2).
- Smoke sets 50-100 core scenarios; medium sets 500-1,000 examples; comprehensive release sets 5,000+ examples (Ch3.3).
- 47% of Mind2Web tasks became invalid or had outdated ground-truth trajectories within months (Ch3.3).
- HotpotQA: 112,779 QA pairs; train-easy 18,089, train-medium 56,814, train-hard 15,661 (Ch3.3).
- Challenge baseline requires at least 100 test cases; worked examples use 500-task and 1,000-case suites (Ch3.4).
- Happy-path-only test set showed 92% action accuracy yet failed on production edge cases (Ch3.8).
- A customer-service agent scoring ~95% offline on simple FAQ questions still failed on complex multi-turn production queries (Ch3.5).
- Illustrative: a tool ablation showing 3% aggregate degradation hid 40% degradation on data-analysis tasks in an 80/20 creative-writing/data-analysis mix (Ch3.7).
- If adversarial hallucination rates exceed 20%, production requires additional safeguards (Ch3.10).
- CI evaluation set of 20-50 cases spanning difficulty levels, query types (factual, analytical, synthesis) and failure modes (Ch4.2).
- Retrieval tuning set of 200-500 queries with known relevant documents (Ch6.2A).
- Accuracy SLA requires a golden set of 1,000+ verified query-answer pairs tested weekly against production (Ch6.4).
- Custom enterprise benchmark: 50-100 test cases minimum (Ref8.03).
Classification
- Patterns
- Stratification by task complexity, domain, modality and timeIn-domain / out-of-domain / temporally shifted / multimodal splitsFeedback-driven test case augmentationTiered evaluation depths (smoke, medium, comprehensive)Stratification by complexity, hop count and reasoning typeBenchmark versioningK-fold partitioningComplexity-tiered test casesAdversarial and edge-case inclusionRoutine, edge-case and adversarial strataProduction-trace backfillOffline evaluationProduction-failure feedback into test setsStratified sampling across task types
- Technologies
- JSONCSVHotpotQA2WikiMultiHopQAMuSiQueFinGAIAGAIAAstaBenchMind2WebWebArenaDCA-BenchJSONLXLSParquet
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Offline/online prediction gap from unrepresentative test dataReintroduction of previously fixed bugsMasked degradation from mixed-difficulty aggregatesStale ground truth as sources evolveTest sets under-representing edge cases and production distributionsHappy-path-only coverage gapsSelection bias in test setsOverfitting to narrow test casesTrain-test leakageDataset-dependency confounds (unrepresentative task mix)Differences caused by dataset difficulty rather than componentsTest rates that fail to predict productionMissing adversarial coverage
Sources
- Ch3.1A: T. Nguyen, "Implement Evaluation Pipelines and Task Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.1A. ISBN: 9798244538229.
- Ch3.1B: T. Nguyen, "Implement Evaluation Pipelines and Task Benchmarks - Guided Practice," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.1B. ISBN: 9798244538229.
- Ch3.2: T. Nguyen, "Compare Agent Performance Across Tasks and Datasets," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.2. ISBN: 9798244538229.
- Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch3.7: T. Nguyen, "Tool Usage Auditing," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.7. ISBN: 9798244538229.
- Ch3.8: T. Nguyen, "Action Accuracy Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.8. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch4.1: T. Nguyen, "Introduction to AI Agent Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.1. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch5.1: T. Nguyen, "Chain-of-Thought (CoT) Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.1. ISBN: 9798244538229.
- Ch5.2: T. Nguyen, "Tree-of-Thought (ToT) Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.2. ISBN: 9798244538229.
- Ch5.3: T. Nguyen, "Self-Consistency Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.3. ISBN: 9798244538229.
- Ch5.8: T. Nguyen, "Semantic Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.8. ISBN: 9798244538229.
- Ch6.2A: T. Nguyen, "Vector Database Selection," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.2A. ISBN: 9798244538229.
- Ch6.4: T. Nguyen, "Data Quality Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.4. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ch7.2: T. Nguyen, "Performance Optimization and Production Monitoring," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.2. ISBN: 9798244538229.
- Ch7.3: T. Nguyen, "NeMo Agent Toolkit Profiling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.3. ISBN: 9798244538229.
- Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
- Ref3.01: NVIDIA, "Agent Evaluation in NVIDIA NeMo Agent Toolkit," NVIDIA NeMo Agent Toolkit Documentation, v1.8. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/agent-toolkit/latest/improve-workflows/evaluate.html
- Ref7.07: E. Li, V. Bellotti, R. Kraus, and R. Kao, "Build a retrieval-augmented generation (RAG) agent with NVIDIA Nemotron," NVIDIA Technical Blog, Sep. 23, 2025. [Online]. Available: https://developer.nvidia.com/blog/build-a-rag-agent-with-nvidia-nemotron/
- Ref8.01: LangChain, "LangSmith observability: AI agent observability platform," LangChain. Accessed: Sep. 27, 2026. [Online]. Available: https://www.langchain.com/langsmith/observability
- Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note