Observability & Evaluation · Software component
Evaluation Score Aggregator
Software componentObservability & EvaluationObservability & Evaluationarc:EvaluationScoreAggregator
A component that aggregates per-case scores into stratified reports by environment, difficulty and capability, weighted composite scores, and Pareto views across competing objectives.
Responsibility. Turns per-case scores into multi-dimensional evaluation results.
Also known as: Multi-dimensional scoring, Stratified reporting
Relationships
is configured by structural
is invoked by dependency
receives data from dynamic
- Response Scorer abstract Ch3.2
sends data to dynamic
Design guidance
- SHOULD report results stratified by environment, difficulty level and capability type rather than a single aggregate.
Quantitative guidance
As stated by the sources; verify before use.
- Brittleness example: 85% DB, 70% OS, 25% KG vs a uniform ~60% generalist (Ch3.2).
Classification
- Patterns
- Stratified reportingComposite weighted scoringPareto-efficient configuration identification
- Risks mitigated
- Single aggregate metric hiding capability gaps
Sources
- Ch3.2: T. Nguyen, "Compare Agent Performance Across Tasks and Datasets," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.2. ISBN: 9798244538229.