Observability & Evaluation · Software component
Statistical Comparator
Software componentObservability & EvaluationObservability & Evaluationarc:StatisticalComparator
A comparison service that tests whether metric differences between a candidate and a baseline or control group are statistically significant, reporting p-values and confidence intervals.
Responsibility. Distinguishes genuine metric changes from sampling noise.
Also known as: Significance tester, Baseline comparison, Statistical significance testing, Group disparity significance test
Relationships
is configured by structural
is invoked by dependency
reads dependency
Design guidance
- MUST NOT accept raw metric differences at face value; significance tests and confidence intervals SHOULD accompany every comparison.
- SHOULD size experiments with power analysis (baseline rate, minimum detectable effect, alpha, power) before calling a winner.
- SHOULD report confidence intervals alongside p-values to convey effect magnitude.
- MUST require both statistical significance and practical significance before recommending deployment.
- SHOULD match the test to the design (paired vs independent samples, non-normal distributions) as specified prospectively, not chosen after seeing results.
- MUST test significance from multiple runs rather than concluding from a single trial.
- MUST defer quality-based rollout decisions until each version has a minimum number of evaluated samples.
Quantitative guidance
As stated by the sources; verify before use.
- Typical alpha 0.05 and power 80%; detecting 86% -> 89% TSR needs ~1,640 queries per group (3,280 total), i.e., >=3 days at 1,000 queries/day (Ch3.1A).
- Example: TSR 87.5% vs 86.2% significant at p=0.03; +1.6 pp with 95% CI [0.4%, 2.8%] (Ch3.1A).
- Accuracy +0.3% at p=0.42 treated as noise; +2.1% at p<0.05 treated as real (Ch3.1A).
- A 78% -> 82% gain is meaningful over 1,000 examples but insignificant over 50 examples or with 6-point standard deviation (Ch3.2).
- Example deployment rule: p < 0.05 and improvement > 3% (Ch3.2).
- p < 0.05 is typically treated as significant at the 95% confidence level (Ch3.7).
- Minimum 20 evaluated conversations per version; detecting a 2% success-rate drop at 5% traffic needs ~10,000 requests (~20 minutes at 1,000 rpm) (Ch4.2).
- Example: 82% vs. 80% approval rates need significance testing; with millions of applications even 0.5-point gaps are significant (Ch9.4).
Classification
- Patterns
- Two-proportion tests for binary metricst-tests for continuous metricsConfidence intervalsStatistical power analysis / sample sizingPaired t-testMcNemar's test for paired binary outcomesWilcoxon signed-rank testBootstrap confidence intervalsDual statistical + practical significance criterionHypothesis testing
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
- Risks mitigated
- False conclusions from measurement noisePremature A/B decisions on insufficient samplesOver-interpreting noise as signal
Sources
- Ch3.1A: T. Nguyen, "Implement Evaluation Pipelines and Task Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.1A. ISBN: 9798244538229.
- Ch3.2: T. Nguyen, "Compare Agent Performance Across Tasks and Datasets," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.2. ISBN: 9798244538229.
- Ch3.7: T. Nguyen, "Tool Usage Auditing," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.7. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch9.4: T. Nguyen, "Fairness and Bias Mitigation," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.4. ISBN: 9798244538229.
- Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref10.05: "The Data Flywheel: Continuous Improvement Loop," unpublished reference note (05-Data-Flywheel-Continuous-Improvement.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note