Observability & Evaluation · Software component

Trajectory Matching Evaluator

Software componentObservability & EvaluationObservability & Evaluationarc:TrajectoryMatchingEvaluator

A task success evaluator that compares an agent's executed action sequence against a golden reference trajectory.

Responsibility. Scores runs by conformance to a reference action sequence.

Also known as: Golden path comparison, Trajectory Evaluator, Action Trajectory Evaluator, Sequence Correctness Evaluator

Variant of Task Success Evaluator abstract

When to choose. Choose only when a task has a single valid execution path; the text warns it penalizes valid alternative routes of stochastic agents.

reads; is configured byevaluatesis invoked byemits telemetry toreadsis target of alternativeTospecializesis target of alternativeTois target of alternativeToReference Trajectory: reads; is configured byReference TrajectoryAgent Controller: evaluatesAgent ControllerEvaluation Harness: is invoked byEvaluation HarnessMetrics Collector: emits telemetry toMetrics CollectorTrace Store: readsTrace StoreLLM Judge: is target of alternativeToLLM JudgeTask Success Evaluator: specializesTask Success EvaluatorState Outcome Scorer: is target of alternativeToState Outcome ScorerMilestone Evaluator: is target of alternativeToMilestone Evaluator
Direct neighbourhood (hover for relationship types)

Relationships

is configured by structural

is invoked by dependency

reads dependency

emits telemetry to dynamic

evaluates assurance

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Trajectory matchingExact matchIn-order matchAny-order matchTrajectory precision and recallStep utility scoring (task state before/after each action)Context-aware scoring with conditioned reference actionsPrerequisite/ordering safety constraint checksMulti-attempt error recovery analysis
Technologies
Google Agent Development Kit
Quality attributes
Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)Safety (ISO/IEC 25010 | NIST AI RMF: safe)
Risks mitigated
Task success masking inefficient or policy-violating action pathsExact match trap penalizing valid alternative pathsSkipped prerequisite steps in regulated workflows

Sources

  1. Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
  2. Ch3.8: T. Nguyen, "Action Accuracy Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.8. ISBN: 9798244538229.